Problem-oriented intelligent abstract method and device for enterprise service, medium and product
Through natural language processing technology, the subject words in business problems are identified and expanded, and combined with the condensation hierarchical clustering algorithm and large language model to generate abstracts, the shortcomings of existing intelligent abstract technology in expression form, readingability and content integrity are solved, and efficient processing and accurate abstract generation of highly professional text data is achieved.
Patent Information
- Application Number
- CN202510041374.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The existing intelligent summary technology has not expressed in a flexible manner, poor reading ability, insufficient content completeness and credibility, and it is difficult to effectively process text data with strong professionalism.
Natural language processing technology is used to identify the subject words in business problems, and expand the subject words based on business graph, main data information or business question responses, filter relevant documents, generate abstract results through a condensation hierarchical clustering algorithm and large language model, and perform semantic vector representation and visual display.
It improves the flexibility and readability of the expression form of the abstract, ensures the completeness and credibility of the content, can better process professional text data, and generate accurate and relevant abstract results.
Smart Images

Figure CN119938907A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing technology, and specifically relates to a problem-oriented intelligent summarization method, device, medium and product for enterprise services. Background Art
[0002] At present, the Internet has entered the era of big data, and text data has become one of the most important resources in enterprises and organizations. Existing intelligent summarization technology can refine and compress text data, which can improve the efficiency of readers' understanding of text data. Intelligent summarization technology is mainly divided into two categories, namely extractive summarization technology and generative summarization technology. Among them, extractive summarization technology splits text data into multiple components according to specific rules, scores the importance of each part or uses a binary classification method to determine whether it is important, retains the parts of the text data with higher importance scores or matching important tags, and uses them as summary results; early generative summarization technology uses a predefined template method to extract key information elements from text data and fill the key information elements into the predefined template, but with the development and application of generative artificial intelligence technology, summary results can be automatically generated through the Transformer architecture.
[0003] However, the aforementioned technologies have the following shortcomings: extractive summarization technology and early generative summarization technology are not flexible enough in expression and have poor readability; generative summarization technology based on the Transformer architecture is prone to losing text data, and the content is not complete and credible enough; all text data are analyzed and processed indiscriminately, and the generated summary results are poorly relevant to actual information needs; the summary results lack content topic structure, affecting reading and comprehension efficiency; for highly professional text data, there is a lack of understanding of business background information and knowledge, which affects importance scores or judgments of important information. Summary of the invention
[0004] The purpose of the present invention is to provide a problem-oriented intelligent summarization method, device, medium and product for enterprise services to solve the above-mentioned problems existing in the prior art.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a problem-oriented intelligent summarization method for enterprise services, comprising: Obtaining a business question, using natural language processing technology to identify a subject word in the business question, and expanding the subject word based on business information to obtain a subject word expansion term, wherein the business information includes a business graph, master data information, or a business question answer; Based on the subject words and subject word extension terms, the title and text of the document list in the database are screened to obtain a first document list, wherein the first document list includes at least one document having the subject words and / or subject word extension terms in the title or text; Filter the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; Clustering the first paragraph in at least one document in the first document list based on an agglomerative hierarchical clustering algorithm to obtain a clustering result, extracting the clustering result using a large language model to generate at least one summary result, extracting the summary result again using the large language model to generate a general summary corresponding to the summary result, representing the summary result and the general summary corresponding to the summary result with a semantic vector to generate a first semantic vector; The business questions are sorted according to the semantic vector similarity between the business question and the first semantic vector, a visualization interface is generated, and the visualization interface is displayed.
[0006] In a possible design, the subject word includes a noun or a noun phrase; correspondingly, when the subject word in the business problem is identified by using natural language processing technology, the subject word is expanded based on the business graph to obtain a subject word expansion term, the method includes: Use word segmentation technology and grammatical analysis technology to identify business problems and obtain subject terms; Use keywords to match the business graph and obtain entity nodes corresponding to the keywords; Obtaining first-order adjacent nodes in the business graph according to the subject words and entity nodes, wherein the first-order adjacent nodes include first-order forward adjacent nodes and first-order reverse adjacent nodes; According to the semantic relevance, the first-order adjacent nodes are sorted by semantic similarity to obtain a semantic similarity sequence, and the first three first-order adjacent nodes corresponding to the semantic similarities are selected from the semantic similarity sequence as the subject word expansion terms.
[0007] In a possible design, the first-order adjacent nodes are sorted by semantic similarity according to semantic relevance to obtain a semantic similarity sequence, and the first three first-order adjacent nodes corresponding to the semantic similarity are selected from the semantic similarity sequence as the subject word expansion terms, including: Extract the corresponding triples in the business graph based on the subject words, where the triples are (X, ri, Yi), where X is the subject word, ri is the relationship name of the triple, and Yi is the first-order adjacent node; The triples are concatenated into a string, the semantic similarity between the business question and the string is calculated and sorted from high to low according to the semantic similarity to obtain a semantic similarity sequence, wherein the triples are concatenated into a string as Ti=concate(X, ri, Yi), where Ti is a string and concate() is a concatenation function; The first-order adjacent nodes corresponding to the first three semantic similarities in the semantic similarity sequence are selected as the subject word expansion terms.
[0008] In a possible design, when natural language processing technology is used to identify the subject words in the business question, and the subject words are expanded based on the answer to the business question to obtain the subject word expansion terms, the method includes: Input the business question into the big language model to obtain the answer to the business question, expand the business question based on the answer to the business question, and obtain the expanded business question; Based on the word segmentation dictionary and natural language processing technology, the expanded business problems are identified to obtain the expanded terms of the subject term.
[0009] In a possible design, when natural language processing technology is used to identify the subject words in the business problem, and the subject words are expanded based on the master data information to obtain the subject word expansion terms, the method includes: Obtain master data information and input the master data information into the word segmentation dictionary; Based on the word segmentation dictionary and natural language processing technology, the master data information of business problems is identified to obtain the subject words; The subject terms are expanded using other master data information corresponding to the master data in the business graph to obtain subject term expansion terms.
[0010] In a possible design, the documents in the first document list are screened according to the relevance to the business problem to obtain at least one first paragraph, including: Extract the natural paragraph containing the subject word or the extended word of the subject word as the first paragraph; Select the natural paragraphs in which the subject words or subject word extension terms appear the most times in the document as the typical paragraphs with positive features; Based on the Embedding function, natural paragraphs in the document are transformed to obtain at least one semantic vector, and the similarity between the at least one semantic vector and the semantic vector of the typical paragraph of the positive feature is calculated; Select the natural paragraphs with the lowest semantic vector similarity in the document and those that do not contain the subject word and do not contain the extended terms of the subject word as the typical paragraphs with negative features; Based on the typical paragraphs with positive features and the typical paragraphs with negative features, the semantic distance of the natural paragraphs that do not contain the subject words or the extended terms of the subject words is compared. If the semantic distance between the natural paragraph and the typical paragraph with positive features is greater than the semantic distance between the natural paragraph and the typical paragraph with negative features, it is extracted as the first paragraph.
[0011] In a possible design, clustering the first paragraph in at least one document in the first document list based on an agglomerative hierarchical clustering algorithm to obtain a clustering result includes: Extract all first paragraphs of at least one document as subclasses and select cluster centers; If the number of subclasses is greater than a preset number threshold, the first paragraph is clustered based on the agglomerative hierarchical clustering algorithm until the number of cluster subclasses reaches the preset number threshold, and the cluster subclasses are output as clustering results; if the number of subclasses is less than or equal to the preset number threshold, the subclasses are output as clustering results.
[0012] In a second aspect, the present invention provides a problem-oriented intelligent summarization device for enterprise services, comprising: An identification and expansion unit is used to obtain business problems, identify subject words in the business problems by using natural language processing technology, expand the subject words based on business information, and obtain subject word expansion terms, wherein the business information includes business graphs, master data information or business problem answers; A first screening unit is used to screen the titles and texts of the document list in the database based on the subject words and the subject word extension terms to obtain a first document list, wherein the first document list contains at least one document having the subject words and / or the subject word extension terms in the title or text; A second screening unit, configured to screen the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; a clustering extraction unit, configured to cluster the first paragraph in at least one document in the first document list based on an agglomerative hierarchical clustering algorithm to obtain a clustering result, extract the clustering result using a large language model to generate at least one summary result, extract the summary result again using the large language model to generate a general summary corresponding to the summary result, represent the summary result and the general summary corresponding to the summary result with a semantic vector, and generate a first semantic vector; The sorting and displaying unit is used to sort the business questions according to the semantic vector similarity between the business questions and the first semantic vector, generate a visualization interface, and display the visualization interface.
[0013] In a third aspect, the present invention provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed on a computer, the problem-oriented intelligent summarization method for enterprise services as described in any one of the above is executed.
[0014] In a fourth aspect, the present invention provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to execute the problem-oriented intelligent summarization method for enterprise services as described in any one of the above.
[0015] Beneficial effects: The present invention discloses a problem-oriented intelligent summary method, device, medium and product for enterprise services, which uses natural language processing technology to segment business problems, identify keywords in business problems, and expand keywords based on business graphs, master data information or business problem answers. It is targeted at enterprise users, has rich business knowledge background information, and can process highly professional text data; all documents are analyzed and screened to ensure the integrity of the content, and only paragraphs containing keywords or keyword expansion terms are retained, the generated summary results are extracted again to generate a general summary, and the general summary and summary results are visualized, and the visualization interface is displayed, with a flexible expression form, which is convenient for users to read and understand. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flow chart of a problem-oriented intelligent summarization method for enterprise services provided by an embodiment of the present invention; Figure 2 A schematic block diagram of a problem-oriented intelligent summarization device for enterprise services provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in combination with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structure of the drawings is only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0018] It should be understood that although the terms first, second, etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another unit. For example, a first unit can be referred to as a second unit, and similarly, a second unit can be referred to as a first unit without departing from the scope of the exemplary embodiments of the present invention.
[0019] It should be understood that the term "and / or" that may appear in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" that may appear in this article describes another type of association object relationship, indicating that two relationships may exist. For example, A / and B can represent two situations: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this article generally indicates that the previous and next associated objects are in an "or" relationship.
[0020] Example: like Figure 1 As shown, this embodiment provides a problem-oriented intelligent summarization method for enterprise services, which may include but is not limited to the following steps.
[0021] S1. Obtain a business question, use natural language processing technology to identify the subject words in the business question, expand the subject words based on business information, and obtain subject word expansion terms, wherein the business information includes a business map, master data information or a business question answer; Among them, the keywords include nouns or noun phrases in business problems; the business graph is an enterprise-level knowledge graph, which records the business terms, work processes, technical standards, rules and regulations of related enterprises; the master data information includes the name, code, abbreviation, alias and English name of the business object.
[0022] Specifically, in step S1, when natural language processing technology is used to identify the subject words in the business problem, and the subject words are expanded based on the business graph to obtain the subject word expansion terms, the method includes: S101. Use word segmentation technology and grammatical analysis technology to identify business problems and obtain subject words; Preferably, word segmentation technology includes but is not limited to regular expressions, hidden Markov model (HMM), Word2Vec model and GloVe model (Global Vectors for Word Representation, word vector model); grammatical analysis technology includes but is not limited to word meaning association analysis, semantic field analysis and semantic template matching.
[0023] S102. Use the subject words to match the business graph to obtain the entity nodes corresponding to the subject words; S103. Obtain first-order adjacent nodes in the business graph according to the subject words and entity nodes, wherein the first-order adjacent nodes include first-order forward adjacent nodes and first-order reverse adjacent nodes; In an embodiment of the present application, for example, if the entity node is an energy monitoring system, there are forward and reverse relationships between the entity nodes in the business graph, such as the forward relationship is "energy monitoring system" -> "function" -> "monitoring power abnormality"; the reverse relationship is "charging pile" -> "access" -> "energy monitoring system", then the first-order forward adjacent node is the monitoring power abnormality, and the first-order reverse adjacent node is the charging pile.
[0024] S104. According to the semantic relevance, the first-order adjacent nodes are sorted by semantic similarity to obtain a semantic similarity sequence, and the first three first-order adjacent nodes corresponding to the semantic similarities are selected from the semantic similarity sequence as the subject word expansion terms.
[0025] In step S104, the first-order adjacent nodes are sorted by semantic similarity according to the semantic relevance to obtain a semantic similarity sequence, and the first three first-order adjacent nodes corresponding to the semantic similarities are selected from the semantic similarity sequence as the subject word expansion terms, including: S1041. Extract corresponding triples in the business graph based on the subject word, where the triple is (X, ri, Yi), where X is the subject word, ri is the relationship name of the triple, and Yi is the first-order adjacent node; S1042. Concatenate the triples into a string, calculate the semantic similarity between the business question and the string, and sort them from high to low according to the semantic similarity to obtain a semantic similarity sequence, wherein the triples are concatenated into a string as Ti=concate(X, ri, Yi), where Ti is a string and concate() is a concatenation function; S1043. Select the first-order adjacent nodes corresponding to the first three semantic similarities in the semantic similarity sequence as the subject word expansion terms.
[0026] In a possible design, in step S1, when natural language processing technology is used to identify the subject words in the business question, and the subject words are expanded based on the answer to the business question to obtain the subject word expansion terms, the method includes: S101. Input the business problem into the large language model, obtain a business problem answer, expand the business problem based on the business problem answer, and obtain an expanded business problem; S102. Identify the expanded business problems based on the word segmentation dictionary and natural language processing technology to obtain the expanded terms of the subject term.
[0027] In a possible design, in step S1, when natural language processing technology is used to identify the subject words in the business problem, and the subject words are expanded based on the master data information to obtain the subject word expansion terms, the method includes: S101. Obtain master data information and input the master data information into the word segmentation dictionary; S102. Identify the main data information of the business problem based on the word segmentation dictionary and natural language processing technology to obtain the subject words; S103. Expand the subject words using other master data information corresponding to the master data in the business graph to obtain subject word expansion terms.
[0028] In the embodiment of the present application, the master data includes but is not limited to supplier data, product data, equipment data and material data, and the master data information includes the name, code, abbreviation, alias and English name in the master data.
[0029] S2. Filtering the titles and texts of the document list in the database based on the subject words and subject word extension terms to obtain a first document list, wherein the first document list contains at least one document having the subject words and / or subject word extension terms in the title or text; S3. Filter the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; Specifically, in step S3, the documents in the first document list are screened according to the relevance of the business problem to obtain at least one first paragraph, including: S301. Extracting the natural paragraph containing the subject word or the subject word extension term as the first paragraph; S302. Select the natural paragraph with the most occurrences of the subject word or the subject word extension term in the document as the typical paragraph with positive features; S303. Based on the Embedding function, natural paragraphs in the document are transformed to obtain at least one semantic vector, and the similarity between the at least one semantic vector and the semantic vector of the typical paragraph of the positive feature is calculated; Embedding function, that is, embedding function, maps other types of entities, such as sentences, documents, objects or people, to another numerical vector space through a specific method.
[0030] S304. Select the natural paragraph with the lowest semantic vector similarity in the document and which does not contain the subject word and does not contain the extended terms of the subject word as the typical paragraph with negative features; S305. Based on the positive feature typical paragraphs and the negative feature typical paragraphs, a semantic distance comparison is performed on the natural paragraphs that do not contain the subject word or the subject word extension terms. If the semantic distance between the natural paragraph and the positive feature typical paragraph is greater than the semantic distance between the natural paragraph and the negative feature typical paragraph, then it is extracted as the first paragraph.
[0031] S4. clustering the first paragraph in at least one document in the first document list based on an agglomerative hierarchical clustering algorithm to obtain a clustering result, extracting the clustering result using a large language model to generate at least one summary result, extracting the summary result using the large language model again to generate a general summary corresponding to the summary result, representing the summary result and the general summary corresponding to the summary result with a semantic vector to generate a first semantic vector; Specifically, in step S4, clustering the first paragraph in at least one document in the first document list based on an agglomerative hierarchical clustering algorithm to obtain a clustering result includes: S401. Extract all first paragraphs of at least one document as subclasses, and select cluster centers; The method for selecting the cluster center is an existing technology and can be selected according to usage needs, which will not be described in detail here.
[0032] S402. If the number of subclasses is greater than a preset number threshold, clustering is performed on the first paragraph based on an agglomerative hierarchical clustering algorithm until the number of cluster subclasses reaches a preset number threshold, and the cluster subclasses are output as clustering results; if the number of subclasses is less than or equal to the preset number threshold, the subclasses are output as clustering results.
[0033] Preferably, the preset quantity threshold may be 5.
[0034] S5. Sort the business questions according to the similarity between the semantic vectors and the first semantic vector, generate a visualization interface, and display the visualization interface.
[0035] In the implementation process of this embodiment, first, a business problem is obtained, and a natural language processing technology is used to identify the subject words in the business problem. The subject words are expanded based on business information to obtain subject word expansion terms, wherein the business information includes a business graph, master data information or a business problem answer; the title and text of a document list in a database are screened based on the subject words and subject word expansion terms to obtain a first document list, wherein the first document list contains at least one document having a subject word and / or subject word expansion term in the title or text; the documents in the first document list are screened according to the relevance of the business problem to obtain at least one first paragraph; the first paragraph in at least one document in the first document list is clustered based on an agglomerative hierarchical clustering algorithm to obtain a clustering result, the clustering result is extracted using a large language model to generate at least one summary result, the summary result is extracted again using a large language model to generate a general summary corresponding to the summary result, the summary result and the general summary corresponding to the summary result are represented by a semantic vector to generate a first semantic vector; the business problem and the first semantic vector are sorted according to the semantic vector similarity, a visualization interface is generated, and the visualization interface is displayed. It has rich business knowledge background information, can process highly professional text data, is targeted at corporate users, has flexible expression forms, and is easy for users to read and understand.
[0036] like Figure 2 As shown, the second aspect of this embodiment provides a problem-oriented intelligent summarization device for enterprise services, including: An identification and expansion unit is used to obtain business problems, identify subject words in the business problems using natural language processing technology, expand the subject words based on business information, and obtain subject word expansion terms, wherein the business information includes business graphs, master data information, or business problem answers; A first screening unit is used to screen the titles and texts of the document list in the database based on the subject words and the subject word extension terms to obtain a first document list, wherein the first document list contains at least one document having the subject words and / or the subject word extension terms in the title or text; A second screening unit, configured to screen the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; a clustering extraction unit, configured to cluster the first paragraph in at least one document based on an agglomerative hierarchical clustering algorithm to obtain a clustering result, extract the clustering result using a large language model to generate at least one summary result, extract the summary result again using the large language model to generate a general summary corresponding to the summary result, represent the summary result and the general summary corresponding to the summary result with a semantic vector, and generate a first semantic vector; The sorting and displaying unit sorts the business questions according to the semantic vector similarity between the business questions and the first semantic vector, generates a visualization interface, and displays the visualization interface.
[0037] In summary, the problem-oriented intelligent summarization device for enterprise services provided in this embodiment obtains business problems through an identification and expansion unit, uses natural language processing technology to identify keywords in business problems, expands keywords based on business information, and obtains keyword expansion terms, wherein the business information includes business graphs, master data information or business problem answers; uses a first screening unit to screen titles and texts of a document list in a database based on keywords and keyword expansion terms to obtain a first document list; uses a second screening unit to screen documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; and uses an agglomerative hierarchical clustering algorithm to extract the first document list. The first paragraph in at least one document in the cluster is clustered to obtain a clustering result, the clustering result is extracted by a large language model to generate at least one summary result, the summary result is extracted by the large language model again to generate a total summary corresponding to the summary result, the summary result and the total summary corresponding to the summary result are represented by semantic vectors to generate a first semantic vector; the sorting and displaying unit sorts the business problem according to the semantic vector similarity between the business problem and the first semantic vector, generates a visualization interface, and displays the visualization interface, analyzes and screens all documents to ensure the completeness of the content, and only retains paragraphs with high similarity to the subject word or the subject word extension term, thereby improving the accuracy of summary generation.
[0038] In a third aspect of this embodiment, a computer-readable storage medium is provided, on which instructions are stored, and when the instructions are executed on a computer, the problem-oriented intelligent summarization method for enterprise services as described in the first aspect of the embodiment is executed. The computer-readable storage medium refers to a carrier for storing data, which may include but is not limited to computer-readable storage media such as floppy disks, optical disks, hard disks, flash memories, USB flash drives, and / or memory sticks, and the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0039] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the third aspect of this embodiment can be referred to the intelligent summary generation method described in the first aspect, and will not be repeated here.
[0040] A fourth aspect of the present embodiment provides a computer program product, including a computer program or instructions, which, when executed by a computer, are used to implement the problem-oriented intelligent summarization method for enterprise services as described in the first aspect of the embodiment.
[0041] The working process, working details and technical effects of the aforementioned computer program product provided in the fourth aspect of this embodiment can be referred to the intelligent distributed data crawling method described in the first aspect, and will not be repeated here.
[0042] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A problem-oriented intelligent summarization method for enterprise services, characterized in that: include: Obtaining a business question, using natural language processing technology to identify a subject word in the business question, and expanding the subject word based on business information to obtain a subject word expansion term, wherein the business information includes a business graph, master data information, or a business question answer; Based on the subject words and subject word extension terms, the title and text of the document list in the database are screened to obtain a first document list, wherein the first document list includes at least one document having the subject words and / or subject word extension terms in the title or text; Filter the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; Clustering the first paragraph in at least one document in the first document list based on an agglomerative hierarchical clustering algorithm to obtain a clustering result, extracting the clustering result using a large language model to generate at least one summary result, extracting the summary result again using the large language model to generate a general summary corresponding to the summary result, representing the summary result and the general summary corresponding to the summary result with a semantic vector to generate a first semantic vector; The business questions are sorted according to the semantic vector similarity between the business question and the first semantic vector, a visualization interface is generated, and the visualization interface is displayed.
2. According to claim 1, a problem-oriented intelligent summarization method for enterprise services is characterized in that: The subject words include nouns or noun phrases; correspondingly, when natural language processing technology is used to identify the subject words in the business problem, the subject words are expanded based on the business graph to obtain subject word expansion terms, the method includes: Use word segmentation technology and grammatical analysis technology to identify business problems and obtain subject terms; Use keywords to match the business graph and obtain entity nodes corresponding to the keywords; Obtaining first-order adjacent nodes in the business graph according to the subject words and entity nodes, wherein the first-order adjacent nodes include first-order forward adjacent nodes and first-order reverse adjacent nodes; According to the semantic relevance, the first-order adjacent nodes are sorted by semantic similarity to obtain a semantic similarity sequence, and the first three first-order adjacent nodes corresponding to the semantic similarities are selected from the semantic similarity sequence as the subject word expansion terms.
3. A problem-oriented intelligent summarization method for enterprise services according to claim 2, characterized in that: According to the semantic relevance, the first-order adjacent nodes are sorted by semantic similarity to obtain a semantic similarity sequence. The first three first-order adjacent nodes corresponding to the semantic similarity are selected from the semantic similarity sequence as the subject word expansion terms, including: Extract the corresponding triples in the business graph based on the subject words, where the triples are (X, ri, Yi), where X is the subject word, ri is the relationship name of the triple, and Yi is the first-order adjacent node; The triples are concatenated into a string, the semantic similarity between the business question and the string is calculated and sorted from high to low according to the semantic similarity to obtain a semantic similarity sequence, wherein the triples are concatenated into a string as Ti=concate(X, ri, Yi), where Ti is a string and concate() is a concatenation function; The first-order adjacent nodes corresponding to the first three semantic similarities in the semantic similarity sequence are selected as the subject word expansion terms.
4. The problem-oriented intelligent summarization method for enterprise services according to claim 1, characterized in that: When natural language processing technology is used to identify the subject words in the business question, and the subject words are expanded based on the answer to the business question to obtain the subject word expansion terms, the method includes: Input the business question into the big language model to obtain the answer to the business question, expand the business question based on the answer to the business question, and obtain the expanded business question; Based on the word segmentation dictionary and natural language processing technology, the expanded business problems are identified to obtain the expanded terms of the subject term.
5. The problem-oriented intelligent summarization method for enterprise services according to claim 1, characterized in that: When natural language processing technology is used to identify the subject words in the business problem, and the subject words are expanded based on the master data information to obtain the subject word expansion terms, the method includes: Obtain master data information and input the master data information into the word segmentation dictionary; Based on the word segmentation dictionary and natural language processing technology, the master data information of business problems is identified to obtain the subject words; The subject terms are expanded using other master data information corresponding to the master data in the business graph to obtain subject term expansion terms.
6. The problem-oriented intelligent summarization method for enterprise services according to claim 1, characterized in that: Filtering the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph includes: Extract the natural paragraph containing the subject word or the extended word of the subject word as the first paragraph; Select the natural paragraphs in which the subject words or subject word extension terms appear the most times in the document as the typical paragraphs with positive features; Based on the Embedding function, natural paragraphs in the document are transformed to obtain at least one semantic vector, and the similarity between the at least one semantic vector and the semantic vector of the typical paragraph of the positive feature is calculated; Select the natural paragraphs with the lowest semantic vector similarity in the document and those that do not contain the subject word and do not contain the extended terms of the subject word as the typical paragraphs with negative features; Based on the typical paragraphs with positive features and the typical paragraphs with negative features, the semantic distance of the natural paragraphs that do not contain the subject words or the extended terms of the subject words is compared. If the semantic distance between the natural paragraph and the typical paragraph with positive features is greater than the semantic distance between the natural paragraph and the typical paragraph with negative features, it is extracted as the first paragraph.
7. The problem-oriented intelligent summarization method for enterprise services according to claim 1, characterized in that: Clustering the first paragraph in at least one document in the first document list based on an agglomerative hierarchical clustering algorithm to obtain a clustering result, including: Extracting all first paragraphs of at least one document in the first document list as subclasses, and selecting cluster centers; If the number of subclasses is greater than a preset number threshold, the first paragraph is clustered based on the agglomerative hierarchical clustering algorithm until the number of cluster subclasses reaches the preset number threshold, and the cluster subclasses are output as clustering results; if the number of subclasses is less than or equal to the preset number threshold, the subclasses are output as clustering results.
8. A problem-oriented intelligent summarization device for enterprise services, characterized in that: include: An identification and expansion unit is used to obtain business problems, identify subject words in the business problems by using natural language processing technology, expand the subject words based on business information, and obtain subject word expansion terms, wherein the business information includes business graphs, master data information or business problem answers; A first screening unit is used to screen the titles and texts of a document list in a database based on the subject words and subject word extension terms to obtain a first document list, wherein the first document list includes at least one document having the subject words and / or subject word extension terms in the title or text; A second screening unit, configured to screen the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; a clustering extraction unit, configured to cluster the first paragraph in at least one document in the first document list based on an agglomerative hierarchical clustering algorithm to obtain a clustering result, extract the clustering result using a large language model to generate at least one summary result, extract the summary result again using the large language model to generate a general summary corresponding to the summary result, represent the summary result and the general summary corresponding to the summary result with a semantic vector, and generate a first semantic vector; The sorting and displaying unit is used to sort the business questions according to the semantic vector similarity between the business questions and the first semantic vector, generate a visualization interface, and display the visualization interface.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the problem-oriented intelligent summarization method for enterprise services according to any one of claims 1 to 7 is executed.
10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or the instruction is executed by a computer, the problem-oriented intelligent summarization method for enterprise services according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-document auto-abstracting method facing to inquiry
CN101620596A
Document processing method and system based on natural language and knowledge graph
CN116501875A
Abstract generation method and device, equipment, storage medium and product
CN116894089A
Abstract generation method, abstract generation device, electronic equipment and readable storage medium
CN118093859A
Semantic-based approach for identifying topics in a corpus of text-based items
US20130085745A1