Language processing method and system based on context-related word components

By combining syntactic and semantic analysis models, the problem of grammatical differences in word vectors in different contexts is solved, the precise correction of word meaning is achieved, and the accuracy of language processing is improved.

CN115526181BActive Publication Date: 2025-07-22北京国瑞数智技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111636209.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-22
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

Existing word vector methods ignore grammatical differences in vocabulary in different contexts, resulting in poor language processing.

Method used

By entering the sentence into the syntax model for breaking sentences, obtaining word components, and correlating them with the context statements, the semantic analysis model is used to correct word meanings in combination with the context content.

Benefits of technology

It realizes accurate correction of the meaning of word parts and quantifiers, and improves the accuracy and effectiveness of language processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115526181B_ABST
    Figure CN115526181B_ABST
Patent Text Reader

Abstract

The present invention provides a language processing method and system based on context-related word components. By inputting a sentence into a syntactic model for sentence segmentation to obtain a first word component, and performing correlation calculations with the previous and next sentences respectively to determine the context content. Then, inputting the first word component into a semantic analysis model one by one, and using its context content as the reference input of the semantic analysis model to obtain the word meaning corresponding to the first word component. By making predictions on the word meanings of the context content, the correction of the word meaning corresponding to the first word component can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network multimedia, and in particular, to a language processing method and system based on context-related word components. Background Art

[0002] Word vectors are an essential foundation for machine learning in natural language processing. However, existing word vectors are static assumptions about a vocabulary. Although they can reduce complexity in modeling, they ignore the syntactic differences of a vocabulary in different contexts.

[0003] Therefore, there is an urgent need for a targeted language processing method and system based on context-related word components. Summary of the Invention

[0004] The object of the present invention is to provide a language processing method and system based on context-related word components. By inputting a sentence into a syntactic model for sentence segmentation to obtain a first word component, and performing correlation calculations with the previous sentence and the next sentence respectively to determine the context content, inputting the first word component into a semantic analysis model one by one, and using its context content as the reference input of the semantic analysis model to obtain the word meaning corresponding to the first word component. By predicting the word meaning of the context content, the word meaning corresponding to the first word component can be corrected.

[0005] In a first aspect, this application provides a language processing method based on context-related word components, the method comprising:

[0006] Obtain a network data stream, extract a sentence therefrom, input the sentence into a syntactic model for sentence segmentation, obtain and cache a first word component, perform a correlation calculation on the first word component with the previous sentence, and determine the previous context content of the first word component according to the magnitude of the correlation degree;

[0007] At the next sentence segmentation, extract the cached first word component and perform a correlation calculation with the next sentence, and determine the next context content of the first word component according to the magnitude of the correlation degree;

[0008] Input the first word component into a semantic analysis model one by one, and use the context content of the first word component as the reference input of the semantic analysis model to obtain the word meaning corresponding to the first word component;

[0009] Wherein, the semantic analysis model includes: obtaining the corresponding context word meaning based on the context content of the semantic analysis reference input, predicting the second word meaning corresponding to the context word meaning, comparing the second word meaning with the word meaning directly obtained from the first word component, calculating the variance value, and using the variance value as a coefficient to correct the word meaning corresponding to the first word component;

[0010] Output the word meaning corresponding to the first word component to the server for keyword detection.

[0011] Combined with the first aspect, in the first possible implementation manner of the first aspect, the extracted statement includes: setting extraction windows with different widths according to each word type, including updating the type of the word and establishing a corresponding relationship between the new word type and the extraction window width.

[0012] Combined with the first aspect, in the second possible implementation manner of the first aspect, the semantic analysis model performs semantic analysis according to the requirements of sentence grammar.

[0013] Combined with the first aspect, in the third possible implementation manner of the first aspect, the kernels of both the semantic analysis model and the syntactic model use neural network models.

[0014] In a second aspect, the present application provides a language processing system based on context-related word components, and the system includes a processor and a memory:

[0015] The memory is used to store program code and transmit the program code to the processor;

[0016] The processor is used to execute the method described in any one of the four possibilities of the first aspect according to the instructions in the program code.

[0017] In a third aspect, the present application provides a computer-readable storage medium, and the computer-readable storage medium is used to store program code, and the program code is used to execute the method described in any one of the four possibilities of the first aspect.

[0018] The present invention provides a language processing method and system based on context-related word components. By inputting a statement into a syntactic model for sentence segmentation to obtain a first word component, and performing correlation calculations with the previous and next sentences respectively to determine the context content, inputting the first word component into a semantic analysis model one by one, and using its context content as the reference input of the semantic analysis model to obtain the word meaning corresponding to the first word component. By making predictions on the word meanings of the context content, the correction of the word meaning corresponding to the first word component can be achieved. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0021] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the protection scope of the present invention more clearly defined.

[0022] Figure 1 A flowchart of a language processing method based on context-related word components provided for this application includes:

[0023] Obtain network data streams, extract sentences therefrom, input the sentences into a syntactic model for sentence segmentation, obtain and cache the first word components, perform an association calculation between the first word components and the previous sentence, and determine the context content of the first word components according to the degree of association.

[0024] During the next sentence segmentation, extract the cached first word components and perform an association calculation with the next sentence, and determine the context content of the first word components according to the degree of association.

[0025] Input the first word components into a semantic analysis model one by one, and use the context content of the first word components as the reference input of the semantic analysis model to obtain the word meaning corresponding to the first word components.

[0026] Among them, the semantic analysis model includes: obtaining the corresponding context word meaning based on the context content of the semantic analysis reference input, predicting the second word meaning corresponding to the context word meaning, comparing the second word meaning with the word meaning directly obtained from the first word components to calculate the variance value, and using the variance value as a coefficient to correct the word meaning corresponding to the first word components.

[0027] Output the word meaning corresponding to the first word components to the server for keyword detection.

[0028] In some preferred embodiments, the extracting of the sentences includes: setting extraction windows with different widths according to each word type, including updating the word type and establishing a corresponding relationship between the new word type and the extraction window width.

[0029] In some preferred embodiments, the semantic analysis model performs semantic analysis according to the requirements of sentence grammar.

[0030] In some preferred embodiments, the kernels of both the semantic analysis model and the syntactic model use neural network models.

[0031] This application provides a language processing system based on context-related word components. The system includes: a processor and a memory:

[0032] The memory is used to store program codes and transmit the program codes to the processor;

[0033] The processor is configured to execute the method according to any one of all the embodiments of the first aspect based on the instructions in the program code.

[0034] This application provides a computer-readable storage medium, which is configured to store program code, and the program code is configured to execute the method according to any one of all the embodiments of the first aspect.

[0035] In a specific implementation, the present invention further provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, it may include some or all of the steps in various embodiments of the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (abbreviation: ROM), or a random access memory (abbreviation: RAM), etc.

[0036] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods according to various embodiments or some parts of the embodiments of the present invention.

[0037] For the same or similar parts among the various embodiments of this specification, reference can be made to each other. In particular, for the embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0038] The above-described embodiments of the present invention do not constitute a limitation on the protection scope of the present invention.

Claims

1. A language processing method based on context-related word components, characterized in that, The method includes: Obtain network data streams, extract statements therefrom, input the statements into a syntactic model for sentence segmentation to obtain and cache first word components, perform an association calculation between the first word components and the previous sentence statement, and determine the context of the first word components according to the degree of association; At the next sentence segmentation, extract the cached first word components and perform an association calculation with the next sentence statement, and determine the context of the first word components according to the degree of association; Input the first word components into a semantic analysis model one by one, and use the context of the first word components as the reference input of the semantic analysis model to obtain the word meaning corresponding to the first word components; Wherein, the semantic analysis model includes: analyzing the context of the semantic analysis reference input to obtain the corresponding context word meaning, predicting the second word meaning corresponding to the context word meaning, comparing the second word meaning with the word meaning directly obtained from the first word components to calculate the variance value, and using the variance value as a coefficient to correct the word meaning corresponding to the first word components; Output the word meaning corresponding to the first word components to the server for keyword detection.

2. The method according to claim 1, characterized in that: The extraction of the statements includes: setting extraction windows with different widths according to each word type, including updating the word type and establishing a corresponding relationship between the new word type and the extraction window width.

3. The method according to any one of claims 1-2, characterized in that: The semantic analysis model performs semantic analysis according to the requirements of sentence grammar.

4. The method according to any one of claims 1-2, characterized in that: The kernels of both the semantic analysis model and the syntactic model use neural network models.

5. A language processing system based on context-related word components, characterized in that, The system includes a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute the method according to any one of claims 1-4 according to the instructions in the program codes.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program codes, and the program codes are used to execute the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Word vector acquisition method and apparatus

    CN106372086A

  • Word vector correction method based on semantic relation constraint and computing system

    CN112966523A