Knowledge Structured Mapping Method, Apparatus and Readable Storage Medium
By performing text reconstruction and semantic coding analysis of the text of the sentence to be processed, structured representation is generated, which solves the limitations of traditional methods when dealing with complex sentences and contexts, and achieves a more comprehensive capture and expression of the deep semantic meaning of the sentence.
Patent Information
- Application Number
- CN202410786735.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-06-18
AI Technical Summary
Traditional text processing methods have limitations in dealing with complex sentences and contexts, resulting in the structured mapping of sentences that cannot fully express the profound semantic meanings in the text.
By reconstructing text of the to-process sentence text, extracting entity words and obtaining their synonyms, replacing entity words to generate multiple reconstructed pending sentence texts. Then, semantic encoding analysis is performed on the sentence text to be processed and the reconstructed sentence text to be processed, and a semantic encoding feature vector is generated. Based on these eigenvectors, the semantic coded eigenvectors of the pending sentence text are optimized to obtain their structured representation.
Through text reconstruction and semantic coding analysis, the deep semantic meaning of sentences can be captured and expressed more comprehensively, and more accurate structured representations can be generated, solving the limitations of traditional methods when dealing with complex sentences and contexts.
Smart Images

Figure CN118520942B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of natural language processing, and particularly to a knowledge structured mapping method, apparatus, and readable storage medium. Background Art
[0002] Knowledge structured mapping is a process of converting unstructured or semi-structured knowledge (such as text, pictures, sounds, etc.) into a structured form (such as a database, a knowledge graph, etc.). This conversion enables knowledge to be more easily processed, stored, and retrieved by computer systems.
[0003] In the field of natural language processing (NLP), knowledge structured mapping typically involves extracting information from text, identifying entities, semantic relationships, and attributes, and organizing them into a structured form.
[0004] However, traditional text processing methods often rely on keyword matching and simple semantic analysis, and these methods have limitations in processing complex sentences and contexts, resulting in the structured mapping of sentences may not be able to fully express the profound semantic meaning in the text. Therefore, an optimized solution is expected. Summary of the Invention
[0005] In view of the above problems, the present disclosure is made. An object of the present disclosure is to provide a knowledge structured mapping method, apparatus, and readable storage medium.
[0006] Embodiments of the present disclosure provide a knowledge structured mapping method, which includes:
[0007] Obtaining a sentence text to be processed;
[0008] Performing text reconstruction on the sentence text to be processed to obtain a plurality of reconstructed sentence texts to be processed;
[0009] Performing semantic encoding analysis on the sentence text to be processed and the plurality of reconstructed sentence texts to be processed to obtain a sentence semantic encoding feature vector to be processed and a plurality of reconstructed sentence text semantic encoding feature vectors;
[0010] Optimizing the sentence semantic encoding feature vector to be processed based on the plurality of reconstructed sentence text semantic encoding feature vectors to obtain a structured representation of the sentence text to be processed.
[0011] For example, according to the knowledge structured mapping method of the embodiments of the present disclosure, wherein performing text reconstruction on the sentence text to be processed to obtain a plurality of reconstructed sentence texts to be processed includes:
[0012] Extracting a plurality of entity words from the sentence text to be processed;
[0013] Obtain synonyms of the multiple entity words;
[0014] Replace the multiple entity words with their synonyms to obtain the multiple reconstructed sentence texts to be processed.
[0015] For example, in the knowledge structuring mapping method according to an embodiment of the present disclosure, wherein extracting multiple entity words from the sentence to be processed includes:
[0016] Perform named entity recognition on the sentence text to be processed to extract the multiple entity words from the sentence to be processed.
[0017] For example, in the knowledge structuring mapping method according to an embodiment of the present disclosure, wherein obtaining synonyms of the multiple entity words includes:
[0018] Obtain synonyms of the multiple entity words from a search engine respectively.
[0019] For example, in the knowledge structuring mapping method according to an embodiment of the present disclosure, wherein performing semantic encoding analysis on the sentence text to be processed and the multiple reconstructed sentence texts to be processed to obtain a sentence text semantic encoding feature vector to be processed and multiple reconstructed sentence text semantic encoding feature vectors includes:
[0020] Pass the sentence text to be processed and the multiple reconstructed sentence texts to be processed through a semantic encoder including a word embedding layer respectively to obtain the sentence text semantic encoding feature vector to be processed and the multiple reconstructed sentence text semantic encoding feature vectors.
[0021] For example, in the knowledge structuring mapping method according to an embodiment of the present disclosure, wherein passing the sentence text to be processed and the multiple reconstructed sentence texts to be processed through a semantic encoder including a word embedding layer respectively to obtain the sentence text semantic encoding feature vector to be processed and the multiple reconstructed sentence text semantic encoding feature vectors includes:
[0022] Perform word segmentation on the sentence text to be processed to convert the sentence text to be processed into a sequence of words to be processed composed of multiple words;
[0023] Use the word embedding layer of the semantic encoder including the word embedding layer to map each word in the sequence of words to be processed to a word vector to obtain a sequence of word vectors to be processed;
[0024] Use the Transformer-based Bert model of the semantic encoder including the word embedding layer to perform global context semantic encoding on the sequence of word vectors to be processed to obtain the sentence text semantic encoding feature vector to be processed.
[0025] For example, in the knowledge structured mapping method according to an embodiment of the present disclosure, based on the semantic encoding feature vectors of the multiple reconstructed sentences to be processed, optimizing the semantic encoding feature vectors of the sentences to be processed to obtain a structured representation of the text of the sentences to be processed includes:
[0026] Determine a first weighted hyperparameter and a second weighted hyperparameter;
[0027] Construct a semantic optimization bias parameter based on the first weighted hyperparameter, the semantic encoding feature vectors of the multiple reconstructed sentences to be processed, and the semantic encoding feature vectors of the sentences to be processed;
[0028] Dot-multiply the semantic optimization bias parameter with the semantic encoding feature vectors of the sentences to be processed to obtain a first semantic optimization fusion term, and dot-multiply the second weighted hyperparameter with the semantic encoding feature vectors of the sentences to be processed to obtain a second semantic optimization fusion term;
[0029] Calculate the dot-sum of the first semantic optimization fusion term and the second semantic optimization fusion term to obtain a structured representation of the text of the sentences to be processed.
[0030] For example, in the knowledge structured mapping method according to an embodiment of the present disclosure, constructing a semantic optimization bias parameter based on the first weighted hyperparameter, the semantic encoding feature vectors of the multiple reconstructed sentences to be processed, and the semantic encoding feature vectors of the sentences to be processed includes:
[0031] Calculate the cosine similarity and Euclidean distance between each semantic encoding feature vector of the multiple reconstructed sentences to be processed and the semantic encoding feature vectors of the sentences to be processed to obtain a sequence of reconstruction-sentence to be processed semantic similarity and a sequence of reconstruction-sentence to be processed semantic distance values;
[0032] Calculate the exponential function with the natural constant as the base for each value in the sequence of reconstruction-sentence to be processed semantic distance values to obtain a sequence of non-linearized reconstruction-sentence to be processed semantic distance values;
[0033] For each semantic encoding feature vector of the reconstructed sentences to be processed, calculate the product of its reconstruction-sentence to be processed semantic similarity and the non-linearized reconstruction-sentence to be processed semantic distance value to obtain a semantic fusion value, and after taking the mean of all the semantic fusion values, multiply the obtained mean by the first weighted hyperparameter to obtain the semantic optimization bias parameter.
[0034] An embodiment of the present disclosure further provides a knowledge structured mapping device, which includes:
[0035] A module for obtaining data to be processed, configured to obtain the text of the sentences to be processed;
[0036] A text reconstruction module for reconstructing the text of the to-be-processed sentence to obtain multiple reconstructed to-be-processed sentence texts;
[0037] A semantic encoding analysis module for performing semantic encoding analysis on the to-be-processed sentence text and the multiple reconstructed to-be-processed sentence texts to obtain a to-be-processed sentence semantic encoding feature vector and multiple reconstructed to-be-processed sentence text semantic encoding feature vectors;
[0038] An optimization module for optimizing the to-be-processed sentence semantic encoding feature vector based on the multiple reconstructed to-be-processed sentence text semantic encoding feature vectors to obtain a structured representation of the to-be-processed sentence text.
[0039] An embodiment of the present disclosure also provides a readable storage medium having computer-executable instructions stored thereon. When the computer-executable instructions are executed in a computer, the computer is made to execute the method as described in any one of the preceding items.
[0040] According to the knowledge structured mapping method, apparatus, and readable storage medium of the embodiments of the present disclosure, a sentence encoder is constructed using natural language processing technology and deep learning algorithms to implement the structured mapping of sentences. Specifically, in the sentence encoder of the present disclosure, text reconstruction based on entity words is performed on the to-be-processed sentence text, and the structured representation of the to-be-processed sentence text is comprehensively generated by combining the text semantic information of the reconstructed sentence with the text semantic information of the to-be-processed sentence text itself. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0042] Figure 1 Shows an application architecture diagram of the knowledge structured mapping method in the embodiments of the present disclosure;
[0043] Figure 2 Shows a flowchart of the knowledge structured mapping method in the embodiments of the present disclosure;
[0044] Figure 3 Shows a flowchart of sub-step S520 of the knowledge structured mapping method in the embodiments of the present disclosure;
[0045] Figure 4 Shows a flowchart of sub-step S540 of the knowledge structured mapping method in the embodiments of the present disclosure;
[0046] Figure 5The structural schematic diagram of the knowledge structured mapping device in the embodiments of the present disclosure is shown;
[0047] Figure 6 The application scenario diagram of the knowledge structured mapping method in the embodiments of the present disclosure is shown; and
[0048] Figure 7 The schematic diagram of the storage medium according to the embodiments of the present disclosure is shown. Detailed implementation manners
[0049] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall also fall within the protection scope of the present disclosure.
[0050] The terms used in this specification are those general terms currently widely used in the art in consideration of the functions of the present disclosure, but these terms may change according to the intentions of those of ordinary skill in the art, precedents, or new technologies in the art. In addition, specific terms may be selected by the applicant, and in this case, their detailed meanings will be described in the detailed description of the present disclosure. Therefore, the terms used in the specification should not be construed as simple names, but based on the meanings of the terms and the overall description of the present disclosure.
[0051] Although the present disclosure makes various references to certain modules in the systems according to the embodiments of the present disclosure, any number of different modules may be used and run on the user terminal and / or the server. The modules are merely illustrative, and different aspects of the systems and methods may use different modules.
[0052] Flowcharts are used in the present disclosure to illustrate the operations performed by the systems according to the embodiments of the present disclosure. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. Instead, various steps may be processed in reverse order or simultaneously as needed. At the same time, other operations may also be added to these processes, or one or several operations may be removed from these processes.
[0053] Figure 1 The application architecture schematic diagram of the knowledge structured mapping method in the embodiments of the present disclosure is shown, including a server 100 and a terminal device 200.
[0054] The terminal device 200 and the server 100 can be connected via the Internet to enable communication between them. Optionally, the above-mentioned Internet uses standard communication technologies and / or protocols. The Internet is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. In some embodiments, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0055] The server 100 can provide various network services for the terminal device 200. Among them, the server 100 can be a single server, a server cluster composed of several servers, or a cloud computing center. Specifically, the server 100 can include a processor 110 (Center Processing Unit, CPU), a memory 120, an input device 130, an output device 140, etc. The input device 130 can include a keyboard, a mouse, a touch screen, etc. The output device 140 can include a display device, such as a liquid crystal display (LCD), a cathode ray tube (CRT), etc.
[0056] The memory 120 can include a read-only memory (ROM) and a random access memory (RAM), and provide the program instructions and data stored in the memory 120 to the processor 110. In the embodiments of the present disclosure, the memory 120 can be used to store the program of the knowledge structured mapping method in the embodiments of the present disclosure.
[0057] The processor 110 is configured to execute the steps of any one of the knowledge structuring mapping methods in the embodiments of the present disclosure according to the obtained program instructions by invoking the program instructions stored in the memory 120.
[0058] In addition, the application architecture diagram in the embodiments of the present disclosure is for more clearly illustrating the technical solutions in the embodiments of the present disclosure, and does not constitute a limitation on the technical solutions provided in the embodiments of the present disclosure. Of course, for other application architectures and business applications, the technical solutions provided in the embodiments of the present disclosure are equally applicable to similar problems.
[0059] The knowledge structuring mapping method provided according to at least one embodiment of the present disclosure will be described below by way of several examples or embodiments in a non-limiting manner. As described below, different features in these specific examples or embodiments can be combined with each other without conflict, so as to obtain new examples or embodiments, and these new examples or embodiments also fall within the scope of protection of the present disclosure.
[0060] In view of the above technical problems, the technical concept of the present disclosure is: using natural language processing technology and deep learning algorithms to construct a sentence encoder to achieve the structured mapping of sentences. Specifically, in the sentence encoder of the present disclosure, the text reconstruction of the sentence text to be processed is performed based on entity words, and the structured representation of the sentence text to be processed is comprehensively generated by combining the text semantic information of the reconstructed sentence and the text semantic information of the sentence text to be processed itself.
[0061] Based on this, Figure 2 The flowchart of the knowledge structuring mapping method in the embodiments of the present disclosure is shown. For example, the knowledge structuring mapping method can be executed by a server, and the server can be Figure 1 the server 100 shown in Figure 2 As shown, according to the knowledge structuring mapping method of the embodiments of the present disclosure, the method includes the steps of: S510, obtaining the sentence text to be processed; S520, performing text reconstruction on the sentence text to be processed to obtain a plurality of reconstructed sentence texts to be processed; S530, performing semantic encoding analysis on the sentence text to be processed and the plurality of reconstructed sentence texts to be processed to obtain a sentence semantic encoding feature vector to be processed and a plurality of reconstructed sentence text semantic encoding feature vectors to be processed; S540, optimizing the sentence semantic encoding feature vector to be processed based on the plurality of reconstructed sentence text semantic encoding feature vectors to obtain the structured representation of the sentence text to be processed.
[0062] Specifically, in the technical solution of the present disclosure, first, the sentence text to be processed is obtained. Among them, the sentence text to be processed can have different sources, such as user input, documents, or online content. Then, named entity recognition is performed on the sentence text to be processed to extract multiple entity words from the sentence to be processed. Among them, entity words in the sentence text to be processed can be extracted through named entity recognition (NER), such as person names, locations, organizations, etc. Entity words are key information in the sentence and are crucial for understanding the semantics of the sentence.
[0063] Considering that when users express semantic information using text, due to their usually different word - using preferences and expression contexts, they may use different words to express the same or similar concepts, or may use the same or similar words to reveal different semantic meanings (such as puns). To explore the diversity of text semantic information, in the technical solution of the present disclosure, further, synonyms of the multiple entity words are obtained from a search engine respectively. Here, the search engine can usually provide a wide range of data sources and rich information. Obtaining synonyms through the search engine can increase the semantic diversity of the sentence and help to more comprehensively understand the intention and emotional color that the user wants to express. In addition, language is constantly developing and changing, and new words and expressions are constantly emerging. The search engine can capture these latest language usage trends and help the sentence encoder learn the semantics of specific entity words.
[0064] Subsequently, the multiple entity words are replaced with their synonyms to obtain multiple reconstructed sentence texts to be processed. In this way, by replacing entity words, sentences with different expressions can be generated, which helps to capture and express the deep meaning of the original sentence and at the same time provides more semantic variants for the sentence.
[0065] Correspondingly, as Figure 3 shown, in step S520, text reconstruction is performed on the sentence text to be processed to obtain multiple reconstructed sentence texts to be processed, including: S521, extracting multiple entity words from the sentence to be processed; S522, obtaining synonyms of the multiple entity words; S523, replacing the multiple entity words with their synonyms to obtain the multiple reconstructed sentence texts to be processed.
[0066] Among them, in step S521, extracting multiple entity words from the sentence to be processed includes: performing named entity recognition on the sentence text to be processed to extract the multiple entity words from the sentence to be processed.
[0067] Among them, in step S522, obtaining synonyms of the multiple entity words includes: obtaining synonyms of the multiple entity words from a search engine respectively.
[0068] After that, the to-be-processed sentence text and the multiple reconstructed to-be-processed sentence texts are respectively passed through a semantic encoder including a word embedding layer to obtain a to-be-processed sentence semantic encoding feature vector and multiple reconstructed to-be-processed sentence text semantic encoding feature vectors. That is, the semantic encoder is used to extract the text semantic expressions in the to-be-processed sentence text and the multiple reconstructed to-be-processed sentence texts respectively.
[0069] Correspondingly, in step S530, semantic encoding analysis is performed on the to-be-processed sentence text and the multiple reconstructed to-be-processed sentence texts to obtain a to-be-processed sentence semantic encoding feature vector and multiple reconstructed to-be-processed sentence text semantic encoding feature vectors, including: passing the to-be-processed sentence text and the multiple reconstructed to-be-processed sentence texts respectively through a semantic encoder including a word embedding layer to obtain the to-be-processed sentence semantic encoding feature vector and the multiple reconstructed to-be-processed sentence text semantic encoding feature vectors.
[0070] Specifically, passing the to-be-processed sentence text and the multiple reconstructed to-be-processed sentence texts respectively through a semantic encoder including a word embedding layer to obtain the to-be-processed sentence semantic encoding feature vector and the multiple reconstructed to-be-processed sentence text semantic encoding feature vectors includes: performing word segmentation on the to-be-processed sentence text to convert the to-be-processed sentence text into a to-be-processed word sequence composed of multiple words; using the word embedding layer of the semantic encoder including the word embedding layer to map each word in the to-be-processed word sequence to a word vector to obtain a sequence of to-be-processed word vectors; using the Transformer-based Bert model of the semantic encoder including the word embedding layer to perform global context semantic encoding on the sequence of to-be-processed word vectors to obtain the to-be-processed sentence semantic encoding feature vector.
[0071] It should be understood that the text semantic expressions in the multiple reconstructed to-be-processed sentence texts are supplements to the to-be-processed sentence text. Therefore, in the technical solution of the present disclosure, it is expected to optimize the to-be-processed sentence semantic encoding feature vector based on the multiple reconstructed to-be-processed sentence text semantic encoding feature vectors to utilize the semantic diversity and comprehensiveness between different words to form a comprehensive text semantic expression of the to-be-processed sentence text, thereby obtaining a structured representation of the to-be-processed sentence text.
[0072] In a specific example of the present disclosure, the encoding process for optimizing the semantic encoding feature vector of the to-be-processed sentence to obtain the structured representation of the to-be-processed sentence text based on the semantic encoding feature vectors of the multiple reconstructed to-be-processed sentence texts includes: first, determining a first weighting hyperparameter and a second weighting hyperparameter; then, constructing a semantic optimization bias parameter based on the first weighting hyperparameter, the semantic encoding feature vectors of the multiple reconstructed to-be-processed sentence texts, and the semantic encoding feature vector of the to-be-processed sentence; next, performing a dot product of the semantic optimization bias parameter and the semantic encoding feature vector of the to-be-processed sentence to obtain a first semantic optimization fusion term, and performing a dot product of the second weighting hyperparameter and the semantic encoding feature vector of the to-be-processed sentence to obtain a second semantic optimization fusion term; then, calculating the sum of the dot products of the first semantic optimization fusion term and the second semantic optimization fusion term to obtain the structured representation of the to-be-processed sentence text.
[0073] Correspondingly, as Figure 4 shown, in step S540, optimizing the semantic encoding feature vector of the to-be-processed sentence based on the semantic encoding feature vectors of the multiple reconstructed to-be-processed sentence texts to obtain the structured representation of the to-be-processed sentence text includes: S541, determining a first weighting hyperparameter and a second weighting hyperparameter; S542, constructing a semantic optimization bias parameter based on the first weighting hyperparameter, the semantic encoding feature vectors of the multiple reconstructed to-be-processed sentence texts, and the semantic encoding feature vector of the to-be-processed sentence; S543, performing a dot product of the semantic optimization bias parameter and the semantic encoding feature vector of the to-be-processed sentence to obtain a first semantic optimization fusion term, and performing a dot product of the second weighting hyperparameter and the semantic encoding feature vector of the to-be-processed sentence to obtain a second semantic optimization fusion term; S544, calculating the sum of the dot products of the first semantic optimization fusion term and the second semantic optimization fusion term to obtain the structured representation of the to-be-processed sentence text.
[0074] In step S542, constructing a semantic optimization bias parameter based on the first weighted hyperparameter, the semantic encoding feature vectors of the multiple sentences to be reconstructed, and the semantic encoding feature vector of the sentence to be processed includes: calculating the cosine similarity and Euclidean distance between each semantic encoding feature vector of the multiple sentences to be reconstructed and the semantic encoding feature vector of the sentence to be processed to obtain a sequence of reconstruction-sentence to be processed semantic similarity and a sequence of reconstruction-sentence to be processed semantic distance values; calculating the exponential function with the natural constant as the base for each value in the sequence of reconstruction-sentence to be processed semantic distance values to obtain a sequence of non-linearized reconstruction-sentence to be processed semantic distance values; for each semantic encoding feature vector of the sentence to be reconstructed, calculating the product of its reconstruction-sentence to be processed semantic similarity and the non-linearized reconstruction-sentence to be processed semantic distance value to obtain a semantic fusion value, and after taking the mean of all the semantic fusion values, multiplying the obtained mean by the first weighted hyperparameter to obtain the semantic optimization bias parameter.
[0075] Specifically, in one example, optimizing the semantic encoding feature vector of the sentence to be processed based on the semantic encoding feature vectors of the multiple sentences to be reconstructed to obtain a structured representation of the sentence text to be processed includes: optimizing the semantic encoding feature vector of the sentence to be processed based on the semantic encoding feature vectors of the multiple sentences to be reconstructed with the following optimization formula to obtain a structured representation of the sentence text to be processed; where the optimization formula is: ; where is the semantic encoding feature vector of the sentence to be processed, is the th semantic encoding feature vector of the multiple sentences to be reconstructed, represents the cosine similarity, the value of is the number of vectors of the semantic encoding feature vectors of the multiple sentences to be reconstructed, represents the exponential function with the natural constant as the base, and are the first weighted hyperparameter and the second weighted hyperparameter respectively, represents the Euclidean distance.
[0076] The determination of the first weighted hyperparameter and the second weighted hyperparameter can control the influence degrees of the first semantic optimization fusion term and the second semantic optimization fusion term on semantic expression, making this optimization process adjustable. The calculation of cosine similarity and Euclidean distance can measure the semantic similarity and difference between each semantic encoding feature vector of the reconstructed sentence to be processed and the semantic encoding feature vector of the sentence to be processed. The introduction of non-linearity can reduce the excessive influence on the optimization process when the distance between the semantic encoding feature vector of the reconstructed sentence to be processed and the semantic encoding feature vector of the sentence to be processed is large, and enhance the model's ability to distinguish different distance levels. In this way, different perspectives provided by each semantic encoding feature vector of the reconstructed sentence to be processed on the possible representation of the sentence text to be processed are used to correct or refine the semantic meaning in the original sentence, generating a richer and more comprehensive structured semantic representation.
[0077] Preferably, in step S540, optimizing the semantic encoding feature vector of the sentence to be processed based on the multiple semantic encoding feature vectors of the reconstructed sentences to be processed to obtain the structured representation of the sentence text to be processed includes the following steps: concatenating the multiple semantic encoding feature vectors of the reconstructed sentences to be processed to obtain a joint semantic encoding feature vector of the reconstructed sentences to be processed; calculating the sum of the dot products of the joint semantic encoding feature vector of the reconstructed sentences to be processed with the square root of its length and the reciprocal of the square root of its second norm to obtain a first intermediate joint semantic encoding feature vector of the reconstructed sentences to be processed; calculating the exponential function with the natural constant as the base of the first intermediate joint semantic encoding feature vector of the reconstructed sentences to be processed to obtain a second intermediate joint semantic encoding feature vector of the reconstructed sentences to be processed; calculating the product of the dot product of the joint semantic encoding feature vector of the reconstructed sentences to be processed with its first norm and the weight hyperparameter to obtain a third intermediate joint semantic encoding feature vector of the reconstructed sentences to be processed; calculating the sum of the dot products of the second intermediate joint semantic encoding feature vector of the reconstructed sentences to be processed and the third intermediate joint semantic encoding feature vector of the reconstructed sentences to be processed to obtain an optimized joint semantic encoding feature vector of the reconstructed sentences to be processed; converting the optimized joint semantic encoding feature vector of the reconstructed sentences to be processed into an optimized multiple semantic encoding feature vectors of the reconstructed sentences to be processed based on the concatenation of the multiple semantic encoding feature vectors of the reconstructed sentences to be processed; and, optimizing the semantic encoding feature vector of the sentence to be processed based on the optimized multiple semantic encoding feature vectors of the reconstructed sentences to be processed to obtain the structured representation of the sentence text to be processed.
[0078] Specifically, considering the differences in semantic feature expressions among the synonyms of the multiple sentences to be reconstructed under word embedding encoding, the semantic encoding feature vectors of the multiple sentences to be reconstructed will also have feature deviation differences relative to the encoding text semantic features of the sentences to be processed, thereby affecting the mapping regression constraint of the semantic encoding feature vectors of the multiple sentences to be reconstructed to the semantic encoding feature vector of the sentence to be processed as the semantic feature deviation source in the semantic reconstruction domain.
[0079] Based on this, in the above preferred example, the structured norm representation of the semantic joint encoding feature vector of the sentences to be reconstructed, which is the joint semantic representation of the semantic encoding feature vectors of the multiple sentences to be reconstructed, is used as the local canonical coordinate for each eigenvalue of the semantic encoding feature vectors of the multiple sentences to be reconstructed, to determine the offset prediction direction of the overall feature distribution representation of the semantic encoding feature vectors of the multiple sentences to be reconstructed relative to the eigenvalue rotation with respect to each eigenvalue of the semantic encoding feature vectors of the multiple sentences to be reconstructed as the center, and the boundary box of the eigenvalue distribution of the semantic encoding feature vectors of the multiple sentences to be reconstructed is used for eigenvalue constraint, to enhance the mapping regression constraint of the semantic encoding feature vectors of the multiple sentences to be reconstructed to the semantic encoding feature vector of the sentence to be processed as the semantic feature deviation source, thereby improving the training speed of the model and the optimization effect on the semantic encoding feature vector of the sentence to be processed.
[0080] It is worth mentioning that the Knowledge Structured Mapping Algorithm (KSMA) is an algorithm designed specifically for knowledge base construction and optimization. By converting a large amount of unstructured text data into a structured knowledge graph form, it realizes the efficient extraction, organization, and application of knowledge. KSMA adopts advanced data processing and natural language understanding technologies, significantly improving the accuracy of knowledge retrieval and the system's response ability to complex user queries, and perfectly solving the contradiction between the high-frequency update requirement of the knowledge base in the fiscal and tax fields and the high overlap degree of source material content.
[0081] It should be understood that the Knowledge Structured Mapping Algorithm (KSMA) is an algorithm that converts unstructured text data into a structured knowledge graph. This method is particularly useful in the field of finance and taxation because it can solve the problems of high-frequency updates of the knowledge base and handling a large amount of repetitive content. The following is the general step-by-step expansion of KSMA: 1. Data collection: Collect unstructured text data from various sources in the field of finance and taxation, such as legal documents, policy statements, academic papers, news reports, etc. 2. Preprocessing: Clean the collected text to remove irrelevant content, such as advertisements, headers and footers, etc.; perform word segmentation to break the text into individual words or phrases. 3. Entity recognition: Use natural language processing (NLP) techniques to identify entities in the text (such as person names, locations, organizations, events, etc.). 4. Relationship extraction: Determine the relationships between entities in the text, such as "belong to", "cause", "affect", etc. 5. Knowledge graph construction: Use the identified entities and relationships to construct a knowledge graph, which is a graphical representation method for representing entities and the relationships between them. 6. Knowledge graph optimization: Optimize the constructed knowledge graph to improve its quality and usability, which can include eliminating redundancy, resolving conflicts, and enhancing the connectivity of the graph. 7. Knowledge verification and update: Regularly verify the accuracy of the knowledge graph and update the graph according to new data and information. 8. Knowledge retrieval and application: Develop an efficient retrieval system that allows users to query the knowledge graph and obtain relevant information; apply the knowledge graph to decision support systems, intelligent question answering systems, etc. to improve the ability to respond to complex queries. 9. User interaction and feedback: Design a user interface that allows users to interact with the knowledge graph and provide feedback; further optimize the knowledge graph and retrieval algorithm based on user feedback. 10. Maintenance and iteration: Continuously monitor the usage of the knowledge graph and perform regular maintenance and iterative updates.
[0082] Through these steps, KSMA can convert a large amount of unstructured data into a useful and queryable knowledge structure, thereby improving the update frequency and retrieval efficiency of the knowledge base in the field of finance and taxation. In addition, KSMA can also be adapted to other fields as long as there is a large amount of text data in these fields that needs to be structured.
[0083] Based on the above embodiments, refer to Figure 5As shown in the figure, it is a schematic structural diagram of a knowledge structured mapping device 800 in an embodiment of the present disclosure. The knowledge structured mapping device 800 includes: a to-be-processed data acquisition module 810, configured to acquire to-be-processed sentence text; a text reconstruction module 820, configured to perform text reconstruction on the to-be-processed sentence text to obtain a plurality of reconstructed to-be-processed sentence texts; a semantic encoding analysis module 830, configured to perform semantic encoding analysis on the to-be-processed sentence text and the plurality of reconstructed to-be-processed sentence texts to obtain a to-be-processed sentence semantic encoding feature vector and a plurality of reconstructed to-be-processed sentence text semantic encoding feature vectors; and an optimization module 840, configured to optimize the to-be-processed sentence semantic encoding feature vector based on the plurality of reconstructed to-be-processed sentence text semantic encoding feature vectors to obtain a structured representation of the to-be-processed sentence text.
[0084] Here, those skilled in the art can understand that the specific functions and operations of each module in the above knowledge structured mapping device 800 have been introduced in detail in the description of the knowledge structured mapping method above with reference to Figures 2 to 4 and therefore, the repeated description thereof will be omitted.
[0085] Figure 6 It is an application scenario diagram of the knowledge structured mapping method according to an embodiment of the present disclosure. As Figure 6 shown, in this application scenario, first, to-be-processed sentence text is acquired (for example, Figure 6 D shown in the figure), and then, synonyms of the plurality of entity words are input into a server (for example, Figure 6 S shown in the figure) deployed with a knowledge structured mapping algorithm, where the server can use the knowledge structured mapping algorithm to process the synonyms of the plurality of entity words to obtain a structured representation of the to-be-processed sentence text.
[0086] Based on the above embodiments, an electronic device according to another exemplary embodiment is further provided in an embodiment of the present disclosure. In some possible implementation manners, the electronic device in the embodiment of the present disclosure may include a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the knowledge structured mapping method in the above embodiments can be implemented.
[0087] For example, taking the server 100 in the present disclosure Figure 1 as an example of the electronic device for illustration, the processor in the electronic device is the processor 110 in the server 100, and the memory in the electronic device is the memory 120 in the server 100.
[0088] An embodiment of the present disclosure also provides a readable storage medium. Figure 7 The figure shows a schematic diagram of a readable storage medium 1000 according to an embodiment of the present disclosure. AsFigure 7 As shown, computer-executable instructions 1001 are stored on the readable storage medium 1000. When the computer-executable instructions 1001 are run by a processor, the knowledge structuring mapping method according to the embodiments of the present disclosure described with reference to the above drawings can be executed. The readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0089] Embodiments of the present disclosure also provide a readable storage medium, on which computer-executable instructions are stored. When the computer-executable instructions are executed in a computer, the computer is made to execute the method as described in any of the previous items.
[0090] The processor of the computer device reads the computer-executable instructions from the readable storage medium, and the processor executes the computer-executable instructions, so that the computer device executes the knowledge structuring mapping method according to the embodiments of the present disclosure.
[0091] Those skilled in the art can understand that the content disclosed in the present disclosure can have various variations and improvements. For example, the various devices or components described above can be implemented by hardware, or can be implemented by software, firmware, or some or all combinations of the three.
[0092] In addition, although the present disclosure makes various references to certain units in the system according to the embodiments of the present disclosure, however, any number of different units can be used and run on the client and / or server. The units are only illustrative, and different aspects of the system and method can use different units.
[0093] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a readable storage medium, such as read-only memory, disk, or optical disc, etc. Optionally, all or part of the steps in the above embodiments can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiments can be implemented in the form of hardware, or can be implemented in the form of a software function module. The present disclosure is not limited to any specific form of combination of hardware and software.
[0094] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0095] The foregoing is a description of the present disclosure and should not be considered as limiting thereof. Although several exemplary embodiments of the present disclosure have been described, those skilled in the art will readily appreciate that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the present disclosure. Accordingly, all such modifications are intended to be included within the scope of the present disclosure as defined by the claims.
[0096] It should be understood that the foregoing is a description of the present disclosure and should not be considered as limited to the specific embodiments disclosed, and modifications to the disclosed embodiments as well as other embodiments are intended to be included within the scope of the appended claims. The present disclosure is defined by the claims and their equivalents.
Claims
1. A knowledge structured mapping method, characterized in that: include: Get the sentence text to be processed; Performing text reconstruction on the sentence text to be processed to obtain a plurality of reconstructed sentence texts to be processed; Performing semantic coding analysis on the sentence text to be processed and the multiple reconstructed sentence texts to be processed to obtain a semantic coding feature vector of the sentence to be processed and multiple reconstructed semantic coding feature vectors of the sentence text to be processed; Based on the multiple reconstructed semantic coding feature vectors of the sentence text to be processed, optimizing the semantic coding feature vector of the sentence text to be processed to obtain a structured representation of the sentence text to be processed; The method of optimizing the semantic coding feature vectors of the sentence text to be processed based on the multiple reconstructed semantic coding feature vectors of the sentence text to be processed to obtain a structured representation of the sentence text to be processed includes: determining a first weighted hyperparameter and a second weighted hyperparameter; Constructing a semantic optimization bias parameter based on the first weighted hyperparameter, the multiple reconstructed semantic encoding feature vectors of the sentence text to be processed, and the semantic encoding feature vector of the sentence to be processed; Performing a dot multiplication of the semantic optimization bias parameter and the semantic coding feature vector of the sentence to be processed to obtain a first semantic optimization fusion item, and performing a dot multiplication of the second weighted hyperparameter and the semantic coding feature vector of the sentence to be processed to obtain a second semantic optimization fusion item; The dot-wise sum of the first semantic optimization fusion item and the second semantic optimization fusion item is calculated to obtain a structured representation of the sentence text to be processed.
2. The knowledge structured mapping method according to claim 1, characterized in that: Performing text reconstruction on the sentence text to be processed to obtain a plurality of reconstructed sentence texts to be processed, including: Extracting multiple entity words from the sentence to be processed; Obtaining synonyms of the multiple entity words; The multiple entity words are replaced with synonyms of the multiple entity words to obtain the multiple reconstructed sentence texts to be processed.
3. The knowledge structured mapping method according to claim 2, characterized in that: Extracting multiple entity words from the sentence to be processed includes: Named entity recognition is performed on the sentence text to be processed to extract the multiple entity words from the sentence to be processed.
4. The knowledge structured mapping method according to claim 3, characterized in that: Obtaining synonyms of the multiple entity words includes: The synonyms of the multiple entity words are respectively obtained from the search engine.
5. The knowledge structured mapping method according to claim 4, characterized in that: Performing semantic coding analysis on the sentence text to be processed and the multiple reconstructed sentence texts to be processed to obtain a semantic coding feature vector of the sentence to be processed and multiple reconstructed semantic coding feature vectors of the sentence text to be processed, including: The sentence text to be processed and the multiple reconstructed sentence texts to be processed are respectively passed through a semantic encoder including a word embedding layer to obtain a semantic encoding feature vector of the sentence to be processed and a semantic encoding feature vector of the multiple reconstructed sentence texts to be processed.
6. The knowledge structured mapping method according to claim 5, characterized in that: The sentence text to be processed and the multiple reconstructed sentence texts to be processed are respectively passed through a semantic encoder including a word embedding layer to obtain a semantic encoding feature vector of the sentence to be processed and a semantic encoding feature vector of the multiple reconstructed sentence texts to be processed, including: Performing word segmentation processing on the sentence text to be processed to convert the sentence text to be processed into a word sequence to be processed consisting of multiple words; Using the word embedding layer of the semantic encoder including the word embedding layer, each word in the to-be-processed word sequence is mapped to a word vector to obtain a sequence of to-be-processed word vectors; The sequence of word vectors to be processed is subjected to global contextual semantic encoding using the converter-based Bert model of the semantic encoder including the word embedding layer to obtain the semantic encoding feature vector of the sentence to be processed.
7. The knowledge structured mapping method according to claim 6, characterized in that: Constructing a semantic optimization bias parameter based on the first weighted hyperparameter, the multiple reconstructed semantic encoding feature vectors of the sentence text to be processed, and the semantic encoding feature vector of the sentence to be processed, including: Calculating the cosine similarity and Euclidean distance between each reconstructed semantic coding feature vector of the sentence text to be processed in the multiple reconstructed semantic coding feature vectors of the sentence text to be processed and the semantic coding feature vector of the sentence to be processed to obtain a sequence of reconstruction-sentence semantic similarities and a sequence of reconstruction-sentence semantic distance values to be processed; Calculating an exponential function of each value in the sequence of semantic distance values of the reconstruction-to-be-processed sentence with a natural constant as the base to obtain a sequence of nonlinear semantic distance values of the reconstruction-to-be-processed sentence; For each of the reconstructed semantic encoding feature vectors of the sentence text to be processed, the product of the reconstruction-sentence semantic similarity and the nonlinear reconstruction-sentence semantic distance value to be processed is calculated to obtain a semantic fusion value, and after taking the mean of all the semantic fusion values, the obtained mean is multiplied by the first weighted hyperparameter to obtain the semantic optimization bias parameter.
8. A knowledge structured mapping device, characterized in that: include: A module for obtaining data to be processed, used to obtain sentence texts to be processed; A text reconstruction module, used for performing text reconstruction on the sentence text to be processed to obtain a plurality of reconstructed sentence texts to be processed; A semantic coding analysis module, used for performing semantic coding analysis on the sentence text to be processed and the multiple reconstructed sentence texts to be processed to obtain a semantic coding feature vector of the sentence to be processed and multiple reconstructed semantic coding feature vectors of the sentence text to be processed; An optimization module, configured to optimize the semantic encoding feature vector of the sentence to be processed based on the multiple reconstructed semantic encoding feature vectors of the sentence text to be processed to obtain a structured representation of the sentence text to be processed; The method of optimizing the semantic coding feature vectors of the sentence text to be processed based on the multiple reconstructed semantic coding feature vectors of the sentence text to be processed to obtain a structured representation of the sentence text to be processed includes: determining a first weighted hyperparameter and a second weighted hyperparameter; Constructing a semantic optimization bias parameter based on the first weighted hyperparameter, the multiple reconstructed semantic encoding feature vectors of the sentence text to be processed, and the semantic encoding feature vector of the sentence to be processed; Performing a dot multiplication of the semantic optimization bias parameter and the semantic coding feature vector of the sentence to be processed to obtain a first semantic optimization fusion item, and performing a dot multiplication of the second weighted hyperparameter and the semantic coding feature vector of the sentence to be processed to obtain a second semantic optimization fusion item; The dot-wise sum of the first semantic optimization fusion item and the second semantic optimization fusion item is calculated to obtain a structured representation of the sentence text to be processed.
9. A readable storage medium having computer executable instructions stored thereon, wherein when the computer executable instructions are executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Text similarity calculation method and device, equipment and storage medium
CN113987115A