Language structure analysis method using language expression
By dividing the sentence structure into four quadrants and using a language expression computation unit to generate detailed language expressions, the problem of inaccurate sentence structure analysis in existing technologies is solved, and efficient analysis and understanding of specific languages is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 崔朝植
- Filing Date
- 2024-08-02
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, language structure analysis methods fail to perform detailed hierarchical subdivision of sentence structure, resulting in inconsistent and unclear representation of grammatical information, which affects the accuracy and efficiency of sentence structure analysis.
Using the language expression method, the sentence structure is divided into four quadrants, representing the subject, predicate, and non-predicate parts respectively. The corresponding language expression is generated through the language expression calculation unit, which includes various markers and symbols to represent syntactic relations in detail.
It enables precise and detailed analysis of sentences in specific languages, improving the efficiency of sentence recognition, understanding, and translation, and allowing for the assessment of language development levels and similarity.
Smart Images

Figure CN121970060A_ABST
Abstract
Description
Language structure analysis method using language expressions Technical Field
[0001] This invention relates to a method for language structure analysis using linguistic expressions. More specifically, it relates to a method for hierarchically subdividing and representing sentences of a specific language in terms of both functional and morphological components through linguistic expressions, thereby performing syntactic analysis of language structure based on the syntactic principles of word-based sentence construction. Background Technology
[0002] With the development of artificial intelligence (AI) and machine learning technologies, the analysis of natural language (including Esperanto) used by humans for communication in daily life, natural language understanding that enables electronic devices to operate based on natural language input, and natural language generation that converts content such as videos and tables into human-understandable forms have become important technological challenges.
[0003] Linguistics, the study of the origin, evolution, and usage principles of language, has been widely applied in various fields such as anthropology, psychology, semiotics, and cryptography, and has had a significant impact on foreign language learning. In recent years, with the development of natural language processing technology, its application in areas such as question-answering systems, information retrieval, voice input, information analysis, foreign language translation, speech recognition, and text representation within the field of artificial intelligence has become increasingly active as a key technology for improving the ease of use of electronic devices.
[0004] There are over 7,000 natural languages on Earth. To process these languages, it is necessary to accurately analyze sentence structure, and the accuracy of the analysis increases with more detailed hierarchical subdivision of sentence structure.
[0005] As a method for representing sentence structure, as shown in Figure 1, a hierarchical classification of sentence structure can be adopted and represented in the form of a tree diagram.
[0006] However, in the tree diagrams described above, as the complexity of the sentences themselves increases, the structure of the tree diagram also becomes more complex.
[0007] Furthermore, despite the complexity of the above structure, grammatical information such as person, nominative, possessive, gender, singular, and plural is not consistently and explicitly represented, thus limiting its application in sentence structure analysis.
[0008] On the other hand, in the prior art related to language analysis, Korean Patent Publication No. 10-2018-0057922, cited in this application, discloses a post-processing-based speech recognition error correction device, as shown in Figure 2. This device includes: a speech recognition unit for recognizing input speech signals and outputting the recognition results in text form; a language recognition unit for performing morphological analysis on the output text sentence, and performing syntactic analysis based on the morphological analysis results, while simultaneously performing semantic analysis according to dependency structure units based on the syntactic analysis results; an error candidate search unit for searching for error candidates using pre-defined co-occurrence information based on the analysis results obtained through the morphological analysis, syntactic analysis, and semantic analysis; a correct candidate search unit for determining possible correct candidates among the error candidates and generating a final correct result; and a text generation unit for generating a final text with corrected errors for the input speech based on the generated correct candidates. With the above device, errors generated during speech recognition can be corrected through post-processing.
[0009] However, as shown in Figure 3, the syntactic analysis results of the post-processing speech recognition error correction device based on language analysis described in the prior art only list the syntactic components of the sentence and do not provide detailed classification information on the functional and morphological components of each component.
[0010] For example, in the prior art, although the syntactic structure of a sentence is decomposed and displayed, specific information related to each constituent element and detailed information about the overall sentence structure are not provided.
[0011] In other words, the following information has not been systematically categorized: information indicating whether each constituent element is a subject, predicate, or modifier; person, gender, and grammatical case information such as nominative and possessive when it is a subject; verb type information determining sentence structure when it is a predicate; and information indicating whether it is a word, derived phrase, subordinate clause, or a negative or interrogative sentence when it is a non-finite verb. Because this diverse information has not been systematically categorized, there are limitations in language structure analysis. Summary of the Invention
[0012] Technical Problem: This invention is proposed to solve the above-mentioned problems. Its purpose is to: target a specific language among the more than 7,000 languages existing on Earth, and through a three-dimensional and hierarchical subdivision of the sentence structure of that language, from the smallest structural unit to the structural unit containing the most complex content, and represent all of them in the form of linguistic expressions, thereby enabling a precise and detailed analysis of the sentence structure of that language, and thus providing a method for language structure analysis using linguistic expressions.
[0013] Although the objectives of the present invention have been described in detail, the present invention is not limited to the objectives described above, and any incidental objectives arising in the process of achieving the objectives described above should also be included within the scope of the technical problems to be solved by the present invention.
[0014] In a preferred embodiment of the present invention, to solve the aforementioned technical problem, a method for analyzing language structure using language expressions is provided. The method comprises: a step of analyzing the morphology and syntactic structure of a target sentence by a morphology and syntactic analysis unit; and a step of generating a corresponding language expression from the morphology and syntactic structure of the sentence analyzed by the morphology and syntactic analysis unit by a language expression calculation unit based on a pre-trained language expression learning algorithm provided by a language expression learning algorithm unit. The language expression calculation unit: divides the sentence into four quadrants—upper left, lower left, upper right, and lower right—centered on a sentence center symbol; assigns a subject symbol to the upper left quadrant; assigns a core verb symbol to the lower left quadrant; assigns a verb type symbol to the lower right quadrant; and assigns non-finite verb symbols representing the remaining phrases other than the subject and verb to the upper right quadrant.
[0015] In another preferred embodiment of the present invention, in order to solve the above-mentioned technical problems, based on the above-mentioned embodiments, the sentence center symbol, subject symbol, core verb symbol, verb type symbol, and non-finite verb symbol representing the remaining phrase all include four subscript positions, namely left superscript, left subscript, right superscript, and right subscript, thereby providing a method for analyzing language structure using language expressions.
[0016] In another preferred embodiment of the present invention, in order to solve the above-mentioned technical problem, based on the above-described embodiments, the verb type symbols include: a first verb (V1) symbol representing all intransitive verbs except the copula "be"; a second verb (V2) symbol representing the copula "be"; a third verb (V3) symbol representing all transitive verbs except the fourth verb (V4); and a fourth verb (V4) symbol representing transitive verbs that require both an object and an object complement. Thus, a method for language structure analysis using linguistic expressions is provided.
[0017] In another preferred embodiment of the present invention, to solve the above-mentioned technical problem, based on the above-described embodiments, for each of the first verb (V1) to fourth verb (V4) representation symbols among the superscript, subscript, superscript, and subscript of the verb type symbols, at least one of the following is assigned: a person marker indicating grammatical person; a quantity marker indicating singular or plural; and a tense marker indicating any one of the following: infinitive, simple present, simple past, present perfect, or past perfect tense. Thus, a method for language structure analysis using linguistic expressions is provided.
[0018] In another preferred embodiment of the present invention, in order to solve the above-mentioned technical problems, based on the above embodiments, the subject symbol and non-finite verb symbol include at least one of the following: a noun symbol representing a noun; an adjective symbol representing an adjective; an adverbial symbol representing an adverb or prepositional phrase; a verb infinitive symbol representing "to + verb infinitive" or an infinitive without "to"; a gerund or participle symbol representing "verb + -ing"; an -ed symbol representing "verb + -ed"; a participle symbol representing "being + past participle"; and a clause symbol representing a subordinate clause. Thus, a method for analyzing the language structure using linguistic expressions is provided.
[0019] In another preferred embodiment of the present invention, in order to solve the above-mentioned technical problems, based on the above-mentioned embodiments, the subject symbol and non-finite verb symbol may be further marked with one or more marks at any of the four subscript positions of the superscript, subscript, superscript and subscript centered on the symbol, including at least one of the following: a person mark indicating the first person, second person, third person or impersonal form (1-4); a character count mark (.) combined with a number to indicate the number of characters of the corresponding noun; a quantity mark indicating singular, even or plural (1, 2, 4); a property mark indicating the semantic attributes of the noun (a, b, c, d), including living organism, inanimate object, mental entity or abstract entity; a case mark indicating the nominative, possessive, accusative, dative, vocative, virtuosic, abductor, instrumental or reflexive case (1-9); a gender mark indicating masculine, feminine or neuter (1, 2, 3); and a biological mark indicating living or non-living things (1, 2). Furthermore, at least one indicator may be added above the subject symbol or non-predicate symbol, which is selected from symbols indicating "this," "that," "there," or movement relationships (., ). , 、→).
[0020] In another preferred embodiment of the present invention, to solve the above-mentioned technical problem, based on the above-described embodiments, parentheses are added to the right of the subject symbol and the non-finite verb symbol, and within the parentheses, at least one symbol representing the remaining phrases other than the subject and the predicate is represented, including: a noun symbol representing a noun; an adjective symbol representing an adjective; an adverbial symbol representing an adverb or prepositional phrase; a verb infinitive symbol representing "to + verb infinitive" or an infinitive without "to"; a gerund or participle symbol representing "verb + -ing"; an -ed symbol representing "verb + -ed"; a participle symbol representing "being + past participle"; and a clause symbol representing a subordinate clause. In this way, the linguistic expression corresponding to the modifying phrase used to modify the subject symbol or non-finite verb symbol preceding the parentheses can be represented.
[0021] In another preferred embodiment of the present invention, in order to solve the above-mentioned technical problems, based on the above-described embodiments, the superscript of the sentence center symbol includes any of the following marks: a mark (0) indicating speech in chronological order; a mark (1) indicating text written from left to right; a mark (2) indicating text written from right to left; a mark (3) indicating text written from top to bottom; a mark (4) indicating text written from bottom to top; a mark (5) indicating drumming; and a mark (6) indicating other writing or expression methods. Thus, a method for analyzing language structure using linguistic expressions is provided.
[0022] In another preferred embodiment of the present invention, in order to solve the above-mentioned technical problem, based on the above-described embodiments, the superscript of the sentence center symbol includes: a marker (a) for representing the literal expression of the sentence; and a marker (e) for representing the semantic content of the sentence. Thus, a method for language structure analysis using linguistic expressions is provided.
[0023] In another preferred embodiment of the present invention, in order to solve the above-mentioned technical problem, based on the above-described embodiments, a numerical or alphabetic marker for indicating the language to which the sentence belongs may be included below the sentence center symbol. Thus, a method for language structure analysis using linguistic expressions is provided.
[0024] In another preferred embodiment of the present invention, in order to solve the above-mentioned technical problems, based on the above embodiments, the markers in the language expression further include: a connector marker representing the conjunctions used in natural language, the connector markers being appended in a parallel manner; and a punctuation marker representing the punctuation marks used in natural language, thereby enabling the representation of the sentence type. Thus, a method for language structure analysis using language expressions is provided.
[0025] According to a preferred embodiment of the present invention, by subdividing the sentence structure of a specific language and representing it in the form of linguistic expressions, it is possible to achieve precise and detailed structural analysis of specific sentences in that language, thereby improving the efficiency of language processing such as sentence recognition, understanding, translation, and retrieval.
[0026] According to a preferred embodiment of the present invention, by subdividing and representing the sentence structure of a particular language in the form of linguistic expressions, a systematic understanding of the structure of that language can be achieved.
[0027] According to a preferred embodiment of the present invention, by subdividing sentence structure and representing it in the form of linguistic expressions, and using the expressiveness of the subdivided linguistic expressions as an indicator, the development level of the language can be evaluated.
[0028] According to a preferred embodiment of the present invention, by comparing linguistic expressions representing subdivided sentence structures, the similarity between a particular language and other languages can be assessed, thereby contributing to linguistic research.
[0029] The effects of the present invention have been described in detail above, but the present invention is not limited to the above effects. Any incidental effects derived in the process of achieving the above effects should also be included within the scope of protection of the present invention. Attached Figure Description
[0030] Figure 1 is a schematic diagram showing an example of how sentence structure is hierarchically classified and represented in the form of a tree diagram in traditional linguistics.
[0031] Figure 2 is a flowchart illustrating the operation of a post-processing speech recognition error correction method based on language analysis in the prior art.
[0032] Figure 3 is an example diagram showing the syntactic analysis results of a post-processing speech recognition error correction device based on language analysis in the prior art.
[0033] Figure 4 is a schematic diagram of the system configuration, illustrating the system implemented by the language structure analysis method using language expressions according to a preferred embodiment of the present invention.
[0034] Figure 5 is a schematic diagram illustrating the language expression structure in a language structure analysis method utilizing language expressions according to a preferred embodiment of the present invention.
[0035] Figure 6 is an example diagram illustrating the configuration of subscript information (or auxiliary information) for subject symbols, verb symbols, and non-finite verb symbols in a language structure analysis method utilizing language expressions according to a preferred embodiment of the present invention.
[0036] Throughout this specification (including the accompanying drawings and detailed descriptions), identical components, as well as components having the same function and / or the same technical or physical effect, are represented by the same reference numerals or the same names. The components shown or described in different embodiments and their functional descriptions can be substituted for each other or applied between different embodiments.
[0037] Furthermore, to facilitate clear and concise identification of the drawings, reference numerals shown in one drawing may be omitted in other drawings. Additionally, parts unrelated to the invention or irrelevant to the description of a particular part may be omitted from the drawings; however, such omission does not imply the absence of these parts.
[0038] Furthermore, it should be noted that the reference numerals used in the prior art drawings are independent of those used in the drawings of this invention. Even if the same reference numerals are used, they should not be construed as representing the same constituent elements. Detailed Implementation
[0039] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that those skilled in the art can easily implement the present invention.
[0040] It should be noted that the embodiments described below are not intended to limit the present invention, and the present invention can be implemented in many different forms, including various modifications, equivalents, or alternatives.
[0041] The features and effects of the present invention can be more clearly understood by referring to the following detailed description of the embodiments in conjunction with the accompanying drawings.
[0042] In describing embodiments of the present invention, well-known functions or configurations will be omitted from detailed description unless they are substantially necessary for understanding the embodiments of the present invention. Furthermore, the terminology used herein is defined based on the functions in the embodiments of the present invention and may vary according to the intention or convention of the user or operator. Therefore, the definitions of these terms should be interpreted based on the overall content of this specification.
[0043] Throughout this specification, when a part is described as being “connected” to another part, this includes not only “direct connection” but also “electrical connection” achieved through one or more intermediate elements.
[0044] Furthermore, when a part is described as "including" a constituent element, unless otherwise specified, the presence of other constituent elements is not excluded, but rather means that it may further include one or more other features, values, steps, operations, constituent elements, components or combinations thereof.
[0045] The following, in conjunction with Figures 4 to 6, provides a more detailed description of a preferred embodiment of the language structure analysis method based on language expressions according to the present invention through specific embodiments.
[0046] First, the system implemented by the language structure analysis method of language expression according to a preferred embodiment of the present invention includes: a morphological and syntactic analysis unit (100), a language expression calculation unit (200), a language expression database (300), and a language expression learning algorithm unit (400).
[0047] The morphological and syntactic analysis unit (100) is used to perform morphological analysis on sentences input in text form, divide the sentences into words of morphological units, determine the part of speech of each word, and identify the syntactic relationship between each word based on the morphological analysis results.
[0048] The language expression calculation unit (200) generates corresponding language expressions based on the parts of speech and syntactic relationships of words divided into morphological units obtained by the morphology and syntactic analysis unit (100), and according to the language expression learning algorithm provided by the language expression learning algorithm unit (400). The process of generating language expressions will be further explained later.
[0049] The language expression generated by the language expression calculation unit (200) is stored in the language expression database (300) along with the corresponding sentence, the morphological analysis result of the sentence, and the syntactic analysis result.
[0050] The language expression learning algorithm unit (400) learns from the generated language expressions and their corresponding sentences, morphological analysis results, and syntactic analysis results stored in the language expression database (300) based on a pre-set algorithm for implementing a preferred embodiment of the present invention, and updates the language expression generation algorithm accordingly. The updated language expression generation algorithm is then provided to the language expression calculation unit (200).
[0051] The following will describe in detail the process of generating language expressions based on the language expression calculation unit (200) analyzing the words divided by morphological units and their syntactic relationships based on the morphological units and the syntactic analysis unit (100), and the language expression learning algorithm provided by the language expression learning algorithm unit (400).
[0052] First, as shown in Figure 5, according to a preferred embodiment of the present invention, the language expression is divided into four quadrants centered on the sentence center symbol (S) (hereinafter referred to as "integral S"): the upper left region (SLU), the lower left region (SLD), the upper right region (SRU), and the lower right region (SRD).
[0053] In this embodiment, the elements assigned to each quadrant around the sentence center symbol (S) are merely examples, and their specific positions can be adjusted and reassigned as needed.
[0054] In a language structure analysis method using language expressions according to a preferred embodiment of the present invention, based on the sentence center symbol (S), the lower left region (SLD) and lower right region (SRD) correspond to the predicate part of the sentence and are assigned analysis symbols to represent the predicate part; the upper left region (SLU) and upper right region (SRU) correspond to the non-predicate part of the sentence, such as the subject and modifier parts.
[0055] Based on the sentence center symbol (S), the upper left region (SLU) is allocated symbols to represent the first non-finite verb components in the sentence. The non-finite verb symbols allocated in the SLU are used as subject symbols.
[0056] For example, as shown in the table below, any of the non-finite verb symbols that contain subject symbols represented by numbers can be selected for assignment.
[0057] Table 1
[0058] In Table 1 above, the symbols can be divided into groups M1, M2 and M3.
[0059] For example, "1" is a noun symbol representing a noun. A noun is a word that represents an entity or a part of an entity. Based on their mode of reference, they can be divided into proper nouns represented by names, demonstrative nouns represented by indicators, and personal nouns represented by persons.
[0060] Furthermore, detailed information related to the noun symbol "1" can be provided by adding parentheses to the right of "1" and specifying the relevant content within the parentheses.
[0061] The symbol “2” represents an adjective. Adjectives are words used to answer the question “what is it like?” and are used to describe or limit nouns.
[0062] Adjectives modify nouns by limiting or specifying their meaning, and can be placed before or after the noun they modify. The adjective symbol "2" can also include determiners such as "a," "an," "the," "this," and "that." Furthermore, adjectives can be modified by adding a dot (·), a hyphen (—), or a wavy line above the adjective symbol "2." Distinguishing markers such as ) are used to differentiate between different types of adjectives.
[0063] The symbol “3” indicates an adverb or prepositional phrase. Adverbs are words that modify nouns and limit their meaning; they can modify verbs, adjectives, adverbs, or entire sentences. When an adverb consists of a single word, it can be represented as “3.1”, or simply “3” with the character count symbol “.1” omitted. When multiple adverbs exist, they can be represented as “3” using the character count symbol (.”) according to the number of words (n). n ".
[0064] On the other hand, when "3" represents a preposition or prepositional phrase, a circular mark can be added above it to distinguish it from "3" representing an adverb. The colon "" indicates a prepositional phrase. When a prepositional phrase consists of two or more words, it can be represented by a period (.) combined with the corresponding number, based on the number of words (n). This can be represented by "". Additionally, it can also be used in this symbol " The parentheses are added to the right of the prepositional phrase, and the specific content of the prepositional phrase is described in detail within the parentheses.
[0065] “4” is the symbol for the infinitive “to + verb infinitive” or the infinitive without “to”.
[0066] “5” is a symbol that indicates a gerund or participle in the form of a verb with “+ing”.
[0067] “6” is the ed form symbol for the past participle of a verb in the “+ed” form.
[0068] “7” is the participle form symbol for the verb form of the structure “being + past participle (pp)”.
[0069] The symbols “4” to “7” belonging to group M2 are symbols indicating the derived forms of verbs.
[0070] "8" belongs to group M3 and is a clause symbol representing a derived clause (subordinate clause).
[0071] As shown in Figure 6, the above-mentioned non-finite verb symbols can also have one or more marks added at one or more of the four subscript positions of superscript, subscript, superscript and subscript centered on the subject symbol or non-finite verb symbol.
[0072] At least one of the following markers may be assigned to at least one of the four subscript positions: a person marker indicating first person, second person, third person, or impersonal form (1-4); a character number marker (.) indicating the number of characters in the corresponding noun combined with a number; a quantity marker indicating singular, even, or plural (1, 2, 4); a property marker indicating the semantic attributes of the noun (a, b, c, d), including living organisms, inanimate objects, mental entities, or abstract entities; a case marker indicating nominative, possessive, accusative, dative, vocative, virtuosic, abductor, instrumental, or reflexive cases (1-9); a gender marker indicating masculine, feminine, or neuter (1, 2, 3); and a biological marker indicating whether the noun is alive or inanimate (1, 2).
[0073] Furthermore, at least one indicator may be added above the subject symbol or non-predicate symbol, the indicator being selected from symbols representing "this," "that," "here," or movement relationships (., ). , 、→).
[0074] The above non-finite verb symbols apply not only to the subject part located in the upper left region (SLU) of the sentence center symbol (S), but also to the modifier part located in the upper right region (SRU).
[0075] Furthermore, when there are multiple non-finite verb components, language expressions can be generated by repeatedly arranging the corresponding symbols according to their quantity.
[0076] On the other hand, the core verb (ε) is set in the lower left (SLD) or lower right (SRD) region of the sentence center symbol (S).
[0077] When a compound verb consists of two or more words, the first word is the core verb (ε), and the remaining words are the residual verbs (ω).
[0078] When a verb is composed of a single word, that word itself serves as the core verb (ε).
[0079] The core verb (ε) configured in the lower left region (SLD) or the lower right region (SRD) has the following functions: (1) The core verb (ε) is used to determine the tense of the sentence.
[0080] (2) The core verb (ε) requires that there be a corresponding subject.
[0081] (3) When the core verb (ε) is located before the subject in the upper left region (SLU), the sentence is formed as an interrogative sentence.
[0082] (4) When “not” appears after the core verb (ε), the sentence is a negative sentence.
[0083] (5) When the core verb (ε) is placed before the subject along with “not”, the sentence is a negative interrogative sentence.
[0084] (6) In a derived clause, the core verb (ε) requires a derivation identifier at the beginning of the clause.
[0085] (7) In a derived clause, when the core verb (ε) is removed from the clause, the subject preceding the core verb (ε) and the derived identifier are omitted.
[0086] For example, let's take the sentence "Do fishes swim?" as an example. Since "fishes" is a noun and serves as the subject, the noun symbol "1" is assigned to the subject position in the upper left area (SLU). Because the core verb "Do" precedes the subject "fishes", the core verb symbol (ε) is assigned to the lower left area (SLD) of the sentence center symbol (S), and its position is set before the subject symbol "1", thus indicating a question structure. The remaining verb (ω), "swim", is an intransitive verb other than the verb "be", and is therefore represented in the lower right area (SRD) as the Class 1 verb (V1) symbol "1". Therefore, we can obtain a structure similar to " It is a language expression similar to a mathematical expression. In addition, punctuation marks such as "?" can be added to indicate the question type.
[0087] In the lower right area (SRD) of the sentence center symbol (S), the type of verb (or predicate symbol) is indicated by numbers or symbols. For example, in English: Class 1 verbs (V1), which represent all non-be verbs that are intransitive, are represented by the number "1"; Class 2 verbs (V2), which represent be verbs, are represented by the number "2"; Class 3 verbs (V3), which represent all transitive verbs except Class 4 verbs (V4), are represented by the number "3"; and Class 4 verbs (V4), which represent transitive verbs that require an object and an object complement, are represented by the number "4".
[0088] For example, Class 1 verbs (V1) can be represented as follows.
[0089] Example: In the sentence "Water flows.", according to a preferred embodiment of the language structure analysis method using linguistic expressions according to the present invention, since "water" is a noun and serves as the subject, it is represented by the noun symbol "1" in the upper left region (SLU). Since "flows" is an intransitive verb that is not a "be" verb, it is represented by the first type of verb (V1) symbol "1" in the lower right region (SRD). Therefore, it is possible to generate a sentence similar to "Water flows." 1 A language expression similar to a mathematical expression for "S1".
[0090] Furthermore, the aforementioned symbols may further include connectors representing conjunctions in natural language, and these connectors may be appended side-by-side to existing symbols. Additionally, punctuation marks from natural language may be introduced to indicate sentence type.
[0091] Furthermore, in the language structure analysis method using language expressions according to a preferred embodiment of the present invention, all symbols (e.g., numbers) can be separated by spaces, and spaces can also be used for numbers within parentheses. However, no spaces are used between the symbols and the parentheses immediately following them; they should be represented continuously.
[0092] On the other hand, in the language structure analysis method using language expressions according to a preferred embodiment of the present invention, the language expression can be represented in a simplified form (simple form) or in a basic form (extended form) that reflects more detailed analysis results.
[0093] For example, the sentence "She was a student" can be represented as a simple expression in a simple language as follows.
[0094]
[0095] Conversely, in the basic (extended) form, since the subject "She" is a noun, it is represented by the noun symbol "1". Furthermore, because it is a third person pronoun, a person marker "3" can be added to the lower right of the noun symbol "1", forming a more specific representation such as "13". Additionally, since "She" is a personal pronoun, a circular marker (o) can be added above the noun symbol "1" to indicate... "Further refinement yields a noun symbol representation that includes the circular marker and the third-person subscript."
[0096] Furthermore, since "She" is singular, a singular quantity marker "1" can be added to the superscript position of the noun punctuation mark to indicate the singularity. The text indicates that "She" indicates a female, and since "She" represents a female, a gender marker "2" can be appended to the left subscript position to indicate female. The sign indicates that, because it is in the nominative case in the sentence, a case marker "1" can be added at the superscript position to indicate the nominative case. This indicates that, since "She" is a single word, a character count marker can be appended at the subscript position. The symbol “.1” indicates the number of words; however, since the symbol “.1”, which indicates a word count of 1, is usually omitted, it can also be omitted in actual usage. Based on the above information, the subject “She” is represented as “1” in the simple form, and specifically as “…” in the basic form. Alternatively, the ".1" can be omitted and replaced with "". "to indicate".
[0097] Furthermore, since the verb "was" is a Class 2 verb (V2), i.e., the verb "be," it is assigned the Class 2 verb symbol "2" as "S2" in the lower right region (SRD) relative to the sentence center symbol (S). Because "was" is in the past tense, a past tense marker "2" can be added to the upper left of the verb type symbol "2," thus forming " 2 The verb "was" can be represented as "2"; since it is singular, a quantity marker "1" can be added to its superscript position, forming the representation "2¹"; furthermore, since the verb consists of a single word, the character count marker ".1" can be omitted, so it can also be represented simply as "2". Based on the above information, "was" is a past tense singular verb of type 2 (V2), and its verb part can be represented according to the above-described linguistic expression form.
[0098]
[0099] Furthermore, since "a student" is a noun phrase consisting of two words, it can be represented by the symbol ".", for example, "1". .2 ".
[0100] Specifically, since the noun phrase is modified by the article "a", it can first be indicated by adding a dot (·) above the adjective symbol "2" which represents the determiner, and then the noun symbol "1" representing the noun "student" can be configured.
[0101] Therefore, this noun phrase can be represented as "1 .2 , and can be enclosed in parentheses to its right to include more detailed information, for example, expressed as " ".
[0102] Furthermore, since "student" refers to a human being and is a biological entity, a biological marker "a" can be added to the right of the noun symbol "1" representing the noun "student," thus obtaining a more specific representation. ".
[0103] In addition, when the sentence is a declarative sentence, a period (.) can be added after the parentheses.
[0104] Therefore, the basic form of "She was a student" can be represented as follows.
[0105]
[0106] Therefore, in the language structure analysis method using language expressions according to a preferred embodiment of the present invention, the hierarchical depth of the language expressions can be selectively adjusted according to analysis needs or design requirements.
[0107] The language structure analysis method using language expressions according to a preferred embodiment of the present invention can also be implemented in the form of a recording medium storing computer-executable instructions, such as application programs or program modules. The computer-readable medium can be any available medium accessible to a computer, and can include volatile and non-volatile media, as well as removable and non-removable media. Furthermore, the computer-readable medium can also include a computer storage medium. The computer storage medium includes volatile or non-volatile media, as well as removable or non-removable media, implemented by any method or technology for storing information (e.g., computer-readable instructions, data structures, program modules, or other data), implemented in this way.
[0108] According to one embodiment of the present invention, the language structure analysis method using language expressions can be executed by an application pre-installed on a terminal device, which may include a program embedded in the platform or operating system of the terminal device. Alternatively, the method can also be executed by an application (i.e., a program) installed by a user on the main terminal device via an application providing server (e.g., an application store server, an application server, or a web server associated with a related service). In this sense, the language structure analysis method using language expressions according to one embodiment of the present invention can be implemented in the form of an application (i.e., a program), whether the application is pre-installed on the terminal device or installed by the user, and can be stored on a computer-readable recording medium for execution by a computer such as a terminal device.
[0109] The foregoing description of the present invention is for illustrative purposes only. Those skilled in the art should understand that various modifications and variations can be made to the present invention without departing from its technical concept or essential characteristics. Therefore, the above embodiments should be considered exemplary in all respects, and not restrictive. For example, constituent elements described in a unified form may be implemented in a distributed form; conversely, constituent elements described in a distributed form may also be implemented in a unified form.
[0110] The scope of protection of this invention should be determined by the appended claims, rather than by the foregoing detailed description. Therefore, all modifications, variations, or equivalent substitutions made within the meaning and scope of the claims, and based on their equivalents, should be interpreted as being included within the scope of protection of this invention.
Claims
1. A method for analyzing language structure using linguistic expressions, characterized in that, include: The morphological and syntactic structure of the target sentence is analyzed by morphological and syntactic analysis units; The language expression calculation unit generates a language expression based on the pre-trained language expression learning algorithm provided by the language expression learning algorithm unit, based on the morphology and syntactic structure of the sentence analyzed by the morphology and syntactic analysis unit; wherein, the language expression calculation unit: divides the sentence into four regions—upper left, lower left, upper right, and lower right—based on the sentence center symbol; and assigns a subject symbol in the upper left region; The core verb symbol is assigned in the lower left region; the verb type symbol is assigned in the lower right region; and the non-finite verb symbol representing the remaining phrases other than the subject and the verb is assigned in the upper right region.
2. The method according to claim 1, characterized in that: Each of the sentence center symbol, subject symbol, core verb symbol, verb type symbol, and non-finite verb symbol includes: left superscript, left subscript, right superscript, and right subscript.
3. The method according to claim 1, characterized in that: The verb type symbols include: a first-class verb (V1) symbol representing all non-be verbs as transitive verbs; a second-class verb (V2) symbol representing be verbs; a third-class verb (V3) symbol representing all transitive verbs except for the fourth-class verb (V4); and a fourth-class verb (V4) symbol representing transitive verbs with objects and object complements.
4. The method according to claim 3, characterized in that: At at least one of the superscript, subscript, superscript and subscript of the verb type symbol, for each of the first type verb (V1) to the fourth type verb (V4) symbols, at least one of the following marks is assigned: a person mark indicating person; a quantity mark indicating singular or plural; and a tense mark indicating any of the tenses of the infinitive, present tense, past tense, present perfect tense or past perfect tense.
5. The method according to claim 1, characterized in that: The subject symbols and non-finite verb symbols include at least one of the following: noun symbols representing nouns; adjective symbols representing adjectives; adverb symbols representing adverbs or prepositional phrases; verb infinitive symbols representing to-infinitives or infinitives without to; gerund or participle symbols representing verbs in the "ing" form; past participle symbols representing verbs in the "ed" form; being participle symbols representing "being + past participle"; and derived clause symbols representing subordinate clauses.
6. The method according to claim 5, characterized in that: The subject and non-finite verb symbols shall have at least one of the following markings attached to at least one of the four subscript positions (superscript, subscript, superscript, and subscript) centered on their respective symbols: a person marker indicating first person, second person, third person, or impersonal form (1-4); a character number marker (.) indicating the number of characters in the corresponding noun; a quantity marker indicating singular, plural, or even number (1, 2, 4); a property marker indicating the semantic attributes of the noun (a, b, c, d), including living organisms, inanimate objects, mental entities, or abstract entities; a case marker indicating nominative, possessive, accusative, dative, vocative, virtuosic, abductor, instrumental, or reflexive cases (1-9); a gender marker indicating masculine, feminine, or neuter (1, 2, 3); and a biological marker indicating living or non-living things (1, 2); and above the subject or non-finite verb symbol, at least one indicator mark shall be attached, the indicator mark being selected from symbols indicating "this," "that," "here," "there," or movement relationships (., ., .). 、 、→)。 7. The method according to claim 5, characterized in that: Parentheses are placed to the right of the subject symbol and the non-finite verb symbol, and at least one of the following symbols is displayed within the parentheses: a noun symbol representing a noun; an adjective symbol representing an adjective; an adverb symbol representing an adverb or prepositional phrase; a symbol representing an infinitive; a gerund or participle symbol representing the "ing" form of a verb; a past participle symbol representing the "ed" form of a verb; a being participle symbol representing the form of "being + past participle"; and a derived clause symbol representing a subordinate clause; thereby indicating the linguistic expression corresponding to the phrase used to modify the subject symbol or non-finite verb symbol located before the parentheses.
8. The method according to claim 2, characterized in that: The superscript of the sentence center symbol includes one of the following: a symbol indicating chronological order of speech (0); a symbol indicating writing from left to right (1); a symbol indicating writing from right to left (2); a symbol indicating writing from top to bottom (3); a symbol indicating writing from bottom to top (4); a symbol indicating drum language (5); and a symbol indicating other writing systems (6).
9. The method according to claim 2, characterized in that: The superscript of the sentence center symbol includes: a symbol representing the literal expression of the sentence (a); a symbol representing the semantic content of the sentence (e); and a numerical symbol (n) representing the number of the smallest unit of sentence classification.
10. The method according to claim 2, characterized in that: Below the sentence center symbol, a number or character marker is provided to indicate the language to which the sentence belongs.
11. The method according to any one of claims 1 to 10, characterized in that: The symbols also include: connectors representing conjunctions in natural language, which can be added in parallel; and punctuation marks representing sentence types in natural language; thereby enabling the representation of sentence types in the language expression.
Citation Information
Patent Citations
Apparatus and method of speech-recognition error correction based on linguistic post-processing
KR1020180057922A