Corpus processing method and device, and storage medium
By splitting the corpus participle and meaning position, constructing the meaning original collection and target meaning original dictionary, the problem of inaccurate semantic description in the vertical field of the existing semantic knowledge base is solved, and a higher accuracy of corpus labeling is achieved.
Patent Information
- Application Number
- CN202210302051.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-03-25
AI Technical Summary
The existing semantic knowledge base lacks the construction of superior and subordinate relationships in vertical domain segmentation, which leads to inaccurate semantic descriptions of words in the domain, which in turn affects the accuracy and accuracy of corpus labeling.
By splitting the sample corpus participle and meaning, a sample corpus is constructed, and the target corpus original dictionary is determined using these corpus originals to achieve more accurate corpus annotation.
Through word segmentation and meaning splitting technology, the semantics of words in the vertical field can be described more accurately, the accuracy of corpus labeling is improved, and the problem of low accuracy of labeling based on word meaning knowledge base is solved.
Smart Images

Figure CN114756649B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of corpus processing, and in particular to a corpus processing method and device, and a storage medium. Background Art
[0002] Both deep learning and machine learning require a large amount of annotated data for model training, and the existing automatic annotation technology is to select keywords with the mouse and click preset tags to label the keywords in the text accordingly. The existing tags are usually determined based on the existing semantic knowledge base, and the existing semantic knowledge base usually describes words by combining semantic primitives and dynamic roles.
[0003] However, the existing semantic knowledge base does not take into account the hierarchical relationship of words in the knowledge base, that is, it does not subdivide in vertical fields to build the knowledge base, but only considers the semantics of the word itself, which cannot accurately describe the semantics of the word in the vertical field. Therefore, when the existing semantic knowledge base only has coarse-grained ontological semantics, the accuracy of the label is also limited, and the accuracy of corpus labeling is also limited, and accurate labeling in vertical fields cannot be achieved.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present invention provide a corpus processing method and device, and a storage medium, so as to at least solve the technical problem of low accuracy of corpus annotation based on a word meaning knowledge base.
[0006] According to one aspect of an embodiment of the present invention, a corpus processing method is provided, comprising: performing word segmentation on a sample corpus in a sample corpus set to obtain a sample field; performing semantic segmentation on the sample field to obtain at least one sample semantic original character; determining a sample semantic original set corresponding to the sample field according to the sample field and the sample semantic original character contained in the sample field, wherein the sample semantic original set includes at least one sample semantic original, each of the sample semantic originals includes a semantic original identifier and a semantic original feature, and the semantic original feature is used to indicate a feature corresponding to the sample field or the sample semantic original character and the semantic original identifier; and determining a target semantic original dictionary according to the sample field and the sample semantic original set corresponding to the sample field.
[0007] According to another aspect of an embodiment of the present invention, a corpus processing device is provided, comprising: a word segmentation unit, configured to perform word segmentation on a sample corpus in a sample corpus set to obtain a sample field; a splitting unit, configured to perform semantic segmentation on the sample field to obtain at least one sample semantic original character; a semantic original unit, configured to determine a sample semantic original set corresponding to the sample field according to the sample field and the sample semantic original characters contained in the sample field, wherein the sample semantic original set includes at least one sample semantic original, each of the sample semantic originals includes a semantic original identifier and a semantic original feature, and the semantic original feature is used to indicate a feature corresponding to the sample field or the sample semantic original character and the semantic original identifier; and a determination unit, configured to determine a target semantic original dictionary according to the sample field and the sample semantic original set corresponding to the sample field.
[0008] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned corpus processing method when running.
[0009] According to another aspect of an embodiment of the present invention, there is provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the corpus processing method through the computer program.
[0010] In an embodiment of the present invention, a sample corpus in a sample corpus set is segmented to obtain a sample field, and the sample field is semantically split to obtain at least one sample semantic primitive word. A sample semantic primitive set corresponding to the sample field is determined according to the sample field and the sample semantic primitive words contained in the sample field, wherein the sample semantic primitive set includes at least one sample semantic primitive, each sample semantic primitive includes a semantic primitive identifier and a semantic primitive feature, and the semantic primitive feature is used to indicate the feature corresponding to the sample field or the sample semantic primitive word and the semantic primitive identifier. According to the sample field and the sample semantic primitive set corresponding to the sample field, a target semantic primitive dictionary is determined. The sample fields are obtained by performing word segmentation on the sample corpus in the sample corpus set, and then the sense word splitting is performed to obtain the sample sense original characters, and the corresponding sample sense original sets are determined based on the sample sense original characters, thereby determining the target sense original dictionary for the sample fields and the sample sense original sets, achieving the purpose of determining the corresponding sample sense original sets based on the word segmentation and sense word splitting of the sample corpus, and forming the target sense original dictionary, thereby using the target sense original dictionary to annotate the corpus, thereby achieving the technical effect of accurately annotating the corpus based on the target sense original dictionary, and further solving the technical problem of low accuracy in corpus annotation based on a word meaning knowledge base. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0012] Figure 1 is a schematic diagram of an application environment of an optional corpus processing method according to an embodiment of the present invention;
[0013] Figure 2 is a flowchart of an optional corpus processing method according to an embodiment of the present invention;
[0014] Figure 3 is a flowchart of an optional corpus processing method according to an embodiment of the present invention;
[0015] Figure 4 is a flowchart of an optional corpus processing method according to an embodiment of the present invention;
[0016] Figure 5 is a flowchart of an optional corpus processing method according to an embodiment of the present invention;
[0017] Figure 6 is a schematic structural diagram of an optional corpus processing device according to an embodiment of the present invention;
[0018] Figure 7 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0021] According to one aspect of an embodiment of the present invention, a corpus processing method is provided. Optionally, the corpus processing method can be applied to, but is not limited to, Figure 1 In the environment shown, the terminal device 102 is not limited to data interaction with the server 112 through the network 110 , the terminal device 102 sends a corpus annotation request to the server 112 through the network 110 , and the server 112 annotates the corpus based on semantics and returns the annotated corpus to the terminal device 102 through the network 110 .
[0022] The server 112 is not limited to running a database 114 and a processing engine 116. The database 114 is used for data storage. The processing engine 116 is not limited to executing S102 to S106 in sequence to realize corpus processing. S102, obtain a sample field. Perform word segmentation on the sample corpus in the sample corpus set to obtain a sample field. S104, obtain a sample semantic primitive, perform semantic segmentation on the sample field, and obtain at least one sample semantic primitive. S106, determine a sample semantic primitive set. Determine a sample semantic primitive set corresponding to the sample field based on the sample field and the sample semantic primitive words contained in the sample field. The sample semantic primitive set includes at least one sample semantic primitive. Each sample semantic primitive includes a semantic primitive identifier and a semantic primitive feature. The semantic primitive feature is used to indicate the feature corresponding to the sample field or the sample semantic primitive word and the semantic primitive identifier. S108, determine a target semantic primitive dictionary. Determine a target semantic primitive dictionary using the sample field and the corresponding sample semantic primitive set.
[0023] Optionally, in this embodiment, the terminal device 102 may be a terminal device configured with a target client, which may include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an IOS phone, etc.), a laptop, a tablet computer, a PDA, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client is not limited to a client that needs to annotate corpus, and is specifically not limited to an audio client, a video client, an instant messaging client, a browser client, an educational client, etc. The network 110 may include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network, and a wide area network, and the wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The server 112 may be a single server, or a server cluster consisting of multiple servers, or a cloud server. The above is only an example, and no limitation is made to this in this embodiment.
[0024] As an optional implementation, Figure 2 As shown, the above corpus processing method includes:
[0025] S202, performing word segmentation on the sample corpus in the sample corpus set to obtain sample fields;
[0026] S204, performing semantic segmentation on the sample field to obtain at least one sample semantic original word;
[0027] S206, determining a sample semantic original set corresponding to the sample field according to the sample field and the sample semantic original characters contained in the sample field;
[0028] In the above S206, the sample sememe set includes at least one sample sememe, each sample sememe includes a sememe identifier and a sememe feature, and the sememe feature is used to indicate a feature corresponding to the sample field or sample sememe word and the sememe identifier;
[0029] S208: Determine a target semantic primitive dictionary using the sample fields and the corresponding sample semantic primitive sets.
[0030] The sample corpus set includes multiple sample corpora, each sample corpus is segmented to obtain a sample field, and the sample field is segmented to obtain a sample semantic original word. The sample corpus is segmented not limited to the stammering word segmentation, and the sample field obtained by the stammering word segmentation is segmented.
[0031] For example, "adjust up" and "adjust down" are one word in the field dimension, but within the field, they can be split into two meanings based on semantics: adjust and high, adjust and low. Washing machines are split into three meanings: wash, clothes, and machine according to their meanings, and washing powder can be split into three meanings: wash, clothes, and powder. Sample semantic primitives are constructed based on the meanings. For example, the four fields of "adjust up", "adjust down", "washing machine", and "washing powder" are split into seven sample semantic primitives: adjust, high, low, wash, clothes, machine, and powder according to their meanings.
[0032] After the sample fields are split into meanings, they are not limited to classification. For example, according to the part of speech, the sample fields are divided into nouns, verbs, adjectives, adverbs, prepositions, etc. It is not limited to further division in each part of speech. For example, verbs are divided into transitive verbs and intransitive verbs. Still taking the target field of household appliances as an example, verbs are not limited to being subdivided into switch categories, which are defined as: 1. Indicates the operation and stop of control equipment and modes; 2. Indicates the opening and closing of control objects. Such as: open, open, open, close, turn off, etc. Adjustment category, defined as indicating the change of the dimensional value of the control equipment. Such as adjust, adjust, adjust, debug, etc. Increase and decrease category, defined as indicating the increase and decrease of the dimensional value of the device. Such as increase, add, and subtract. Lift and lower category, defined as indicating 1. The height of the control equipment and object position; 2. The height of the dimensional value of the control equipment. Such as rise and fall, etc.
[0033] For example, nouns are divided into common nouns and proper nouns. Common nouns are further divided into equipment categories, which are defined as electrical devices, equipment and components used to connect and disconnect circuits and transform circuit parameters to achieve control, adjustment, switching, detection and protection of circuits or electrical equipment. Such as refrigerators, air conditioners, wine cabinets, dehumidifiers, etc. Component categories are defined as a component of a device, such as refrigerator doors, washing machine lids, chair backs, etc. Dimension categories are defined as built-in controllable, adjustable and switchable parameters of the device. Such as temperature, mode, speed, strength, etc. Dimension values are defined as parameter values that can be controlled, adjusted and switched by the device, such as: 35 degrees, 200 watts, 1 hour, etc.
[0034] For example, adjectives are defined as describing the range of dimensions, such as: big, small, high, low, bright, dark, etc. Prepositions are defined as: in, for, to, to, into, indicating the specific dimensional value of the equipment adjustment.
[0035] As an optional implementation, Figure 3 As shown, the above-mentioned determination of the sample semantic original set corresponding to the sample field according to the sample field and the sample semantic original word contained in the sample field includes:
[0036] S302, determining a description format corresponding to the sample field according to the part of speech of the sample field;
[0037] In the above S302, the description format includes at least one description semantic primitive, and the at least one description semantic primitive includes a part-of-speech semantic primitive.
[0038] S304, determining at least one sample meaning original symbol corresponding to each sample meaning original word according to the description format corresponding to the sample field and the sample meaning original word contained in the sample field;
[0039] S306: Determine a sample semantic primitive set corresponding to the sample field according to at least one sample semantic primitive symbol corresponding to each sample semantic primitive character.
[0040] The description format is not limited to the description semantic primitives at least included in the sample field determined according to the part of speech, and different parts of speech are not limited to corresponding to different description semantic primitives. The description semantic primitives corresponding to different parts of speech all include part of speech semantic primitives for identifying the part of speech of the sample field.
[0041] As an optional implementation, the above-mentioned determination of at least one sample semantic original symbol corresponding to each sample semantic original word according to the description format corresponding to the sample field and the sample semantic original words contained in the sample field includes: when the part of speech of the sample field is a noun, determining at least one of the attribute semantic original, characteristic semantic original, and feature semantic original corresponding to each sample semantic original word in the sample field.
[0042] Semantics are not limited to being expressed using the JSON string format, and include semantic identifiers and semantic features. Semantic features are used to indicate the features of the sample field corresponding to the semantic identifier. Taking the part-of-speech semantics as an example, the semantic identifier is "part-of-speech", and the semantic features are the part-of-speech of the sample field, which is not limited to verbs, nouns, adjectives, etc. In the case where the part-of-speech of the sample field is a noun, the description format corresponding to the noun is not limited to sequentially determining the part-of-speech semantics of the sample field and at least one of the attribute semantics, characteristic semantics, and feature semantics corresponding to each sample semantic character. Each sample semantic character corresponds to at least one attribute semantic, characteristic semantic, or feature semantic. The number of attribute semantics, characteristic semantics, and feature semantics corresponding to the sample field is not limited to one or more.
[0043] Taking the sample field "refrigerator" as an example, the corresponding sample semantic primitive set is not limited to being expressed as: {"part of speech": "N"}, {"attribute": "equipment"}, {"container": "square"}, {"characteristic": "electricity"}, {"action": "ice"}, where {"part of speech": "N"} is a part of speech semantic primitive, used to indicate that refrigerator is a noun field, {"attribute": "equipment"} is an attribute semantic primitive, used to indicate that the attribute of the refrigerator is equipment, {"characteristic": "electricity"} is a characteristic semantic primitive, used to indicate that the refrigerator is an equipment that consumes electricity, {"container": "square"} and {"action": "ice"} are characteristic semantic primitives, {"action": "ice"} is a characteristic semantic primitive determined by the word "ice", and {"container": "square"} is a characteristic semantic primitive determined by the word "box".
[0044] Taking the sample field "dryer" as an example, the corresponding sample semantic primitive set is not limited to being expressed as: {"part of speech": "N"}, {"attribute": "equipment"}, {"attribute 2": "machine"}, {"characteristic": "electricity"}, {"action": "dry"}, {"state": "dry"}, {"container": "square"}. Among them, {"attribute 2": "machine"}, {"action": "dry"}, {"state": "dry"}, {"container": "square"} are all characteristic semantic primitives, and {"attribute 2": "machine"} is a characteristic semantic primitive determined based on "machine".
[0045] Taking the sample field "washing machine" as an example, the corresponding sample semantic primitive set is not limited to being expressed as: {"part of speech": "N"}, {"attribute": "equipment"}, {"attribute 2": "machine"}, {"characteristic": "electricity"}, {"action": "wash"}, {"object": "clothes"}, {"container": "square"}. Among them, {"attribute 2": "machine"}, {"action": "wash"}, {"object": "clothes"}, {"container": "square"} are all characteristic semantic primitives, and {"object": "clothes"} is a characteristic semantic primitive determined based on "clothes".
[0046] In an embodiment of the present application, a sample corpus in a sample corpus set is segmented to obtain a sample field, and the sample field is segmented to obtain at least one sample semantic original word. A sample semantic original set corresponding to the sample field is determined according to the sample field and the sample semantic original word contained in the sample field, wherein the sample semantic original set includes at least one sample semantic original, each sample semantic original includes a semantic original identifier and a semantic original feature, and the semantic original feature is used to indicate the feature corresponding to the sample field or the sample semantic original word and the semantic original identifier. According to the sample field and the sample semantic original set corresponding to the sample field, a target semantic original dictionary is determined. By segmenting the sample corpus, a sample semantic original dictionary is generated. The sample corpus in the combination is segmented to obtain a sample field, and then the sense word splitting is performed to obtain the sample semantic original word, and the corresponding sample semantic original set is determined based on the sample semantic original word, so as to determine the target semantic original dictionary for the sample field and the sample semantic original set, thereby achieving the purpose of determining the corresponding sample semantic original set based on the sample semantic original word determined by the segmentation and sense word splitting of the sample corpus, forming the target semantic original dictionary, and then using the target semantic original dictionary to annotate the corpus, thereby achieving the technical effect of accurately annotating the corpus based on the target semantic original dictionary, and further solving the technical problem of low accuracy of corpus annotation based on the word meaning knowledge base.
[0047] As an optional implementation, Figure 4 As shown, after determining the target semantic dictionary, it also includes:
[0048] S402, obtaining a target annotation tag for annotating the corpus to be annotated.
[0049] In the above S402, the target annotation tag includes at least one target semantic original tag and the annotation matching degree corresponding to the target semantic original tag.
[0050] The target annotation label is not limited to include one or more target sememe labels, each target sememe label is not limited to include at least one target sememe, and each target sememe is not limited to be any sample sememe including a sememe identifier and a sememe feature. Taking the JSON string format as an example, each target sememe is not limited to be a JSON string.
[0051] The tag matching degree is used to indicate the matching degree of the field to be annotated, including but not limited to complete matching and inclusion matching. Complete matching is used to indicate that it needs to be completely consistent with the target semantic original label, and inclusion matching is not limited to indicating that it only needs to include the target semantic original in the target semantic original label.
[0052] S404: in the target semantic dictionary, determine candidate fields that match the annotation matching degree of the target semantic tag in sequence.
[0053] The target semantic dictionary is not limited to being a semantic dictionary corresponding to the target domain, so that the semantic interpretation of each target semantic tag in the target annotation tag is realized through the target semantic dictionary to determine the candidate field that matches the annotation matching degree of the target semantic tag.
[0054] In the target sememe dictionary, when the sample sememe set matches the target sememe label, the sample field is determined as a candidate field matching the target sememe label.
[0055] S406: If there is a target field matching the candidate field in the corpus to be annotated, annotate the target field.
[0056] In the above S406, the target field is a field obtained by segmenting the corpus to be annotated. Annotating the corpus to be annotated is not limited to annotating based on fields in the corpus, and the field is not limited to being obtained by segmenting the corpus.
[0057] As an optional implementation, labeling the target field with the target semantic original label includes: adjusting the display parameters of the target field to highlight the target field in the corpus to be labeled; or labeling the target field with the target semantic original label. The highlighting is not limited to any one or a combination of adjusting the color, font, font size, etc. of the field.
[0058] The target fields are labeled using the target semantic original labels, which is not limited to adding the target semantic original labels to the target fields in the form of annotations. The fields in the corpus are labeled according to the definitions in the target semantic original dictionary, thereby achieving the technical effect of interpreting the corpus according to the semantic original dictionary.
[0059] In an embodiment of the present application, a target annotation tag is obtained for annotating a corpus to be annotated, wherein the target annotation tag includes at least one target semantic primitive tag and an annotation matching degree corresponding to the target semantic primitive tag. In a target semantic primitive dictionary, candidate fields matching the annotation matching degree of the target semantic primitive tag are determined in sequence. The target semantic primitive dictionary records sample fields and sample semantic primitive sets corresponding to the sample fields. When a target field obtained by word segmentation in the corpus to be annotated hits a candidate field, the target field is annotated using the target semantic primitive tag. By searching for a candidate field matching the target semantic primitive tag of the target annotation tag in the target semantic primitive dictionary corresponding to the field, the target field is annotated using the target semantic primitive tag when the target field obtained by word segmentation of the corpus to be annotated hits a candidate field. This achieves the purpose of determining an annotation tag based on the sample field in the semantic primitive dictionary and the corresponding sample semantic primitive set to annotate the target field matching the semantic primitive set in the corpus, thereby achieving the technical effect of accurately annotating fields in the corpus based on semantics.
[0060] As an optional implementation, the candidate fields that are sequentially determined to match the target semantic original label with the target semantic original label in the target semantic original dictionary include:
[0061] S1, construct an initial field set corresponding to the target semantic original label, where the initial field set is an empty set;
[0062] S2, when the annotation matching degree indicates a complete match, in the target semantic original dictionary, search for candidate fields whose semantic original set is consistent with the target semantic original label, and add the candidate fields to the initial field set to form a candidate field set that matches the target semantic original label.
[0063] When determining candidate fields that match a target semantic label, we are not limited to building a set of fields that are paired with each target semantic label.
[0064] The constructed initial field set is an empty set, which does not contain any fields. When candidate fields are identified in the target semantic dictionary, the candidate fields are added to the target initial field set to form a candidate field set. The number of candidate fields in each candidate field set is not limited, but each candidate field is matched with the target semantic label in terms of label matching.
[0065] When the labeled match indicates a complete match, it is determined to search the target semantic dictionary for a candidate field that is consistent with the semantic set of the target semantic tag. For example, loc=[semetic tag] is used to indicate a complete match, and loc≈[semetic tag] is used to indicate a contained match.
[0066] For example, if the target annotation label is loc=[{"part of speech":"N"},{"attribute":"equipment"},{"container":"square"},{"feature":"electricity"},{"action":"bake"}], and the target semantic original label is {"part of speech":"N"},{"attribute":"equipment"},{"container":"square"},{"feature":"electricity"},{"action":"bake"}, and the annotation matching degree is a complete match, then the target semantic original dictionary is searched for fields that completely match, and is not limited to finding: oven, electric oven.
[0067] As an optional implementation, after constructing the initial field set corresponding to the target semantic primitive label, the above also includes searching the target semantic primitive dictionary for candidate fields whose semantic primitive set contains all the target semantic primitives in the target semantic primitive label when the annotation matching degree indicates a contained match, and adding the candidate fields to the initial field set to form a candidate field set that matches the target semantic primitive label.
[0068] Taking the target annotation label loc≈[{"attribute":"equipment"}] as an example, the target semantic primitive label is {"attribute":"equipment"}, and the annotation matching degree is inclusion match. Then, the field containing {"attribute":"equipment"} is searched in the target semantic primitive dictionary. Then, the above-mentioned oven, electric oven, washing machine, dryer and refrigerator are all determined as candidate fields. The candidate fields also include fields whose semantic primitive features of attribute semantic primitives such as dehumidifiers are equipment.
[0069] As an optional implementation, Figure 5 As shown, after determining the candidate fields that match the annotation matching degree of the target semantic original label in the target semantic original dictionary, the following is also included:
[0070] S502, performing word segmentation on the corpus to be annotated to obtain at least one current field;
[0071] S504, searching for the current field in the candidate field set corresponding to each target semantic original label in turn;
[0072] S506: When the current field is found in the candidate field set, determine that the current field is a target field that hits the target semantic original label.
[0073] After determining the candidate field set corresponding to the annotation matching degree of each target semantic original label, the annotated corpus is segmented to obtain the current field. Each current field is searched in turn, and a field search is performed in each candidate field set. If the current field is found in the candidate field set, the current field is determined to be the target field that hits the target semantic original label.
[0074] For example, if the corpus to be annotated is "adjust the temperature of the refrigerator to 30 degrees" and the target annotation label is loc≈[{"attribute":"equipment"}], then after determining the candidate field set that fully matches {"attribute":"equipment"}, the fields obtained by performing word segmentation on the corpus to be annotated are: adjust, refrigerator, of, temperature, to, 30 degrees, and the fields are searched in the candidate field set in sequence. Based on finding refrigerator in the candidate field set, refrigerator is determined to be the target field in the corpus to be annotated with the target annotation label of loc≈[{"attribute":"equipment"}], and refrigerator is annotated.
[0075] Taking the target annotation label [loc≈{"part of speech":"V"}, loc≈{"attribute":"equipment"}, loc≈{"attribute":"dimension"}, loc≈{"attribute":"dimension value"}] as an example, determine the candidate field sets corresponding to the target annotation labels in the target semantic dictionary in turn, and when the current field is obtained by process segmentation of the corpus to be annotated, search the process fields in each candidate field set in turn, so as to determine the target fields that hit the target semantic label in turn. Still taking the corpus to be annotated as "adjust the temperature of the refrigerator to 30 degrees" as an example, annotate "adjust", "refrigerator", "to", and "30 degrees".
[0076] In the embodiment of the present application, by including multiple target semantic tags in the target annotation tag, flexible and accurate annotation of the corpus based on semantics is achieved, such as accurate annotation of operation instructions in the target field.
[0077] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0078] According to another aspect of the embodiments of the present invention, a corpus processing device for implementing the above corpus processing method is also provided. Figure 6 As shown, the device comprises:
[0079] A word segmentation unit 602, used to segment the sample corpus in the sample corpus set to obtain a sample field;
[0080] A splitting unit 604 is used to perform semantic segmentation on the sample field to obtain at least one sample semantic original word;
[0081] A semantic unit 606, configured to determine a sample semantic primitive set corresponding to the sample field according to the sample field and the sample semantic primitive word contained in the sample field, wherein the sample semantic primitive set includes at least one sample semantic primitive, each sample semantic primitive includes a semantic primitive identifier and a semantic primitive feature, and the semantic primitive feature is used to indicate a feature corresponding to the sample field or the sample semantic primitive word and the semantic primitive identifier;
[0082] The determining unit 608 is configured to determine a target semantic primitive dictionary according to the sample field and the sample semantic primitive set corresponding to the sample field.
[0083] In an embodiment of the present application, a sample corpus in a sample corpus set is segmented to obtain a sample field, and the sample field is segmented to obtain at least one sample semantic original word. A sample semantic original set corresponding to the sample field is determined according to the sample field and the sample semantic original word contained in the sample field, wherein the sample semantic original set includes at least one sample semantic original, each sample semantic original includes a semantic original identifier and a semantic original feature, and the semantic original feature is used to indicate the feature corresponding to the sample field or the sample semantic original word and the semantic original identifier. According to the sample field and the sample semantic original set corresponding to the sample field, a target semantic original dictionary is determined. By segmenting the sample corpus, a sample semantic original dictionary is generated. The sample corpus in the combination is segmented to obtain a sample field, and then the sense word splitting is performed to obtain the sample semantic original word, and the corresponding sample semantic original set is determined based on the sample semantic original word, so as to determine the target semantic original dictionary for the sample field and the sample semantic original set, thereby achieving the purpose of determining the corresponding sample semantic original set based on the sample semantic original word determined by the segmentation and sense word splitting of the sample corpus, forming the target semantic original dictionary, and then using the target semantic original dictionary to annotate the corpus, thereby achieving the technical effect of accurately annotating the corpus based on the target semantic original dictionary, and further solving the technical problem of low accuracy of corpus annotation based on the word meaning knowledge base.
[0084] Optionally, the semantic primitive unit is further used to determine a description format corresponding to the sample field according to the part of speech of the sample field, wherein the description format includes at least one description semantic primitive, and at least one description semantic primitive includes a part-of-speech semantic primitive; determine at least one sample semantic primitive symbol corresponding to each sample semantic primitive word according to the description format corresponding to the sample field and the sample semantic primitive words contained in the sample field; determine the sample semantic primitive set corresponding to the sample field according to at least one sample semantic primitive symbol corresponding to each sample semantic primitive word.
[0085] Optionally, the semantic unit is further used to determine at least one of an attribute semantic primitive, a characteristic semantic primitive, and a feature semantic primitive corresponding to each sample semantic primitive character in the sample field when the part of speech of the sample field is a noun.
[0086] Optionally, the above-mentioned corpus processing device further includes:
[0087] An acquisition unit is used to acquire a target annotation tag for annotating the corpus to be annotated after determining the target semantic original dictionary, wherein the target annotation tag includes at least one target semantic original tag and an annotation matching degree corresponding to the target semantic original tag;
[0088] A candidate unit is used to determine, in the target semantic dictionary, candidate fields that match the annotation matching degree of the target semantic tag in sequence, wherein the target semantic dictionary records the candidate fields and the semantic set corresponding to the candidate fields, and the semantic set includes at least one semantic identifier indicating a semantic of a target domain, and the target domain is the domain to which the corpus to be annotated belongs;
[0089] The annotation unit is used to annotate the target field if there is a target field matching the candidate field in the corpus to be annotated, where the target field is the field obtained by segmenting the corpus to be annotated.
[0090] Optionally, the above-mentioned candidate unit is also used to construct an initial field set corresponding to the target semantic original label, wherein the initial field set is an empty set; when the annotation matching degree indicates a complete match, in the target semantic original dictionary, search for candidate fields whose semantic original set is consistent with the target semantic original label, and add the candidate fields to the initial field set to form a candidate field set that matches the target semantic original label.
[0091] Optionally, the above-mentioned candidate unit is also used to, after constructing an initial field set corresponding to the target semantic primitive label, search in the target semantic primitive dictionary for candidate fields whose semantic primitive set contains all the target semantic primitives in the target semantic primitive label when the annotation matching degree indicates a contained match, and add the candidate fields to the initial field set to form a candidate field set that matches the target semantic primitive label.
[0092] Optionally, the above-mentioned corpus processing device also includes a processing unit, which is used to determine the candidate fields that match the annotation matching degree of the target semantic original label in the target semantic original dictionary in turn, and then perform word segmentation on the annotated corpus to obtain at least one current field; search for the current field in the candidate field set corresponding to each target semantic original label in turn; when the current field is found from the candidate field set, determine that the current field is the target field that hits the target semantic original label.
[0093] Optionally, the annotation unit is further used to adjust display parameters of the target field to highlight the target field in the corpus to be annotated; or, annotate the target field using a target semantic original label.
[0094] In an embodiment of the present application, a target annotation tag is obtained for annotating a corpus to be annotated, wherein the target annotation tag includes at least one target semantic primitive tag and an annotation matching degree corresponding to the target semantic primitive tag. In a target semantic primitive dictionary, candidate fields matching the annotation matching degree of the target semantic primitive tag are determined in sequence. The target semantic primitive dictionary records sample fields and sample semantic primitive sets corresponding to the sample fields. When a target field obtained by word segmentation in the corpus to be annotated hits a candidate field, the target field is annotated using the target semantic primitive tag. By searching for a candidate field matching the target semantic primitive tag of the target annotation tag in the target semantic primitive dictionary corresponding to the field, the target field is annotated using the target semantic primitive tag when the target field obtained by word segmentation of the corpus to be annotated hits a candidate field. This achieves the purpose of determining an annotation tag based on the sample field in the semantic primitive dictionary and the corresponding sample semantic primitive set to annotate the target field matching the semantic primitive set in the corpus, thereby achieving the technical effect of accurately annotating fields in the corpus based on semantics.
[0095] According to another aspect of the embodiments of the present invention, there is also provided an electronic device for implementing the above-mentioned corpus processing method. The electronic device may be Figure 1 The terminal device or server shown in the figure. This embodiment is described by taking the electronic device as a server as an example. Figure 7 As shown, the electronic device includes a memory 702 and a processor 704. The memory 702 stores a computer program, and the processor 704 is configured to execute the steps in any of the above method embodiments through the computer program.
[0096] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0097] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0098] S1, segment the sample corpus in the sample corpus set to obtain sample fields;
[0099] S2, performing semantic segmentation on the sample field to obtain at least one sample semantic original word;
[0100] S3, determining a sample semantic primitive set corresponding to the sample field according to the sample field and the sample semantic primitive word contained in the sample field, wherein the sample semantic primitive set includes at least one sample semantic primitive, each sample semantic primitive includes a semantic primitive identifier and a semantic primitive feature, and the semantic primitive feature is used to indicate a feature corresponding to the sample field or the sample semantic primitive word and the semantic primitive identifier;
[0101] S4, determining a target semantic primitive dictionary using the sample fields and the corresponding sample semantic primitive sets.
[0102] Alternatively, a person skilled in the art may understand that: Figure 7 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an IOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, and other terminal devices. Figure 7 The structure of the electronic device is not limited. Figure 7 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 7 Different configurations are shown.
[0103] Among them, the memory 702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the corpus processing method and device in the embodiment of the present invention. The processor 704 executes various functional applications and data processing by running the software programs and modules stored in the memory 702, that is, realizing the above-mentioned corpus processing method. The memory 702 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 702 may further include a memory remotely arranged relative to the processor 704, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 702 can be specifically, but not limited to, used to store information such as a target semantic dictionary, sample fields, sample semantic characters, and sample semantic sets. As an example, if Figure 7 As shown, the memory 702 may include, but is not limited to, the word segmentation unit 602, the splitting unit 604, the semantic unit 606, and the determination unit 608 in the corpus processing device. In addition, it may also include, but is not limited to, other module units in the device, which will not be described in detail in this example.
[0104] Optionally, the transmission device 706 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one example, the transmission device 706 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 706 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0105] In addition, the electronic device further comprises: a display 708 for displaying the target semantic primitive dictionary, sample fields and sample semantic primitive sets; and a connection bus 710 for connecting various module components in the electronic device.
[0106] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes may form a peer-to-peer (P2P, Peer To Peer) network, and any form of computing device, such as a server, terminal and other electronic devices, may become a node in the blockchain system by joining the peer-to-peer network.
[0107] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional implementations of the above-mentioned corpus processing. The computer program is configured to execute the steps of any of the above-mentioned method embodiments when it is run.
[0108] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0109] S1, segment the sample corpus in the sample corpus set to obtain sample fields;
[0110] S2, performing semantic segmentation on the sample field to obtain at least one sample semantic original word;
[0111] S3, determining a sample semantic primitive set corresponding to the sample field according to the sample field and the sample semantic primitive word contained in the sample field, wherein the sample semantic primitive set includes at least one sample semantic primitive, each sample semantic primitive includes a semantic primitive identifier and a semantic primitive feature, and the semantic primitive feature is used to indicate a feature corresponding to the sample field or the sample semantic primitive word and the semantic primitive identifier;
[0112] S4, determining a target semantic primitive dictionary using the sample fields and the corresponding sample semantic primitive sets.
[0113] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0114] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0115] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention.
[0116] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0117] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0118] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0119] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0120] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A corpus processing method, It is characterized in that include: Segment the sample corpus in the sample corpus set to obtain sample fields; Performing semantic segmentation on the sample field to obtain at least one sample semantic original word; Determine a sample semantic primitive set corresponding to the sample field according to the sample field and the sample semantic primitive word contained in the sample field, wherein the sample semantic primitive set includes at least one sample semantic primitive symbol, each of the sample semantic primitive symbols includes a semantic primitive identifier and a semantic primitive feature, and the semantic primitive feature is used to indicate a feature corresponding to the sample field or the sample semantic primitive word and the semantic primitive identifier; Determining a target semantic primitive dictionary according to the sample field and the sample semantic primitive set corresponding to the sample field; Obtaining a target annotation tag for annotating the corpus to be annotated, wherein the target annotation tag includes at least one target semantic original tag and an annotation matching degree corresponding to the target semantic original tag; In the target semantic dictionary, candidate fields that match the annotation matching degree of the target semantic tag are determined in sequence; In the case that there is a target field matching the candidate field in the corpus to be annotated, the target field is annotated, wherein the target field is a field obtained by segmenting the corpus to be annotated.
2. The method according to claim 1, It is characterized in that The determining, according to the sample field and the sample semantic primitive words contained in the sample field, a sample semantic primitive set corresponding to the sample field comprises: Determining a description format corresponding to the sample field according to the part of speech of the sample field, wherein the description format includes at least one description semantic primitive, and the at least one description semantic primitive includes a part of speech semantic primitive; Determine at least one sample meaning original symbol corresponding to each sample meaning original word according to the description format corresponding to the sample field and the sample meaning original word contained in the sample field; The sample semantic primitive set corresponding to the sample field is determined according to the at least one sample semantic primitive symbol corresponding to each of the sample semantic primitive characters.
3. The method according to claim 2, It is characterized in that The step of determining at least one sample meaning original symbol corresponding to each sample meaning original word according to the description format corresponding to the sample field and the sample meaning original word contained in the sample field comprises: In the case that the part of speech of the sample field is a noun, at least one of an attribute semantic primitive, a characteristic semantic primitive, and a feature semantic primitive corresponding to each of the sample semantic primitive characters in the sample field is determined respectively.
4. The method according to claim 1, It is characterized in that The step of sequentially determining, in the target semantic dictionary, candidate fields that match the annotation matching degree of the target semantic label comprises: Constructing an initial field set corresponding to the target semantic original label, wherein the initial field set is an empty set; When the annotation matching degree indicates a complete match, in the target semantic primitive dictionary, a candidate field whose semantic primitive set is consistent with the target semantic primitive label is searched, and the candidate field is added to the initial field set to form a candidate field set that matches the target semantic primitive label.
5. The method according to claim 4, It is characterized in that After constructing the initial field set corresponding to the target semantic original label, the method further includes: When the annotated matching degree indicates a match, in the target semantic primitive dictionary, the semantic primitive set is searched for candidate fields whose target semantic primitives include all target semantic primitives in the target semantic primitive label, and the candidate fields are added to the initial field set to form a candidate field set that matches the target semantic primitive label.
6. The method according to claim 5, It is characterized in that After sequentially determining the candidate fields that match the annotation matching degree of the target semantic original label in the target semantic original dictionary, the method further includes: Segmenting the corpus to be annotated to obtain at least one current field; Searching for the current field in the candidate field set corresponding to each of the target semantic labels in turn; When the current field is found from the candidate field set, the current field is determined to be the target field that hits the target semantic original label.
7. The method according to claim 1, It is characterized in that The marking of the target field comprises: Adjusting the display parameters of the target field to highlight the target field in the corpus to be annotated; Or, the target field is labeled using the target semantic tag.
8. A corpus processing device, It is characterized in that include: The word segmentation unit is used to segment the sample corpus in the sample corpus set to obtain sample fields; A splitting unit, used for performing semantic bit splitting on the sample field to obtain at least one sample semantic original word; A semantic unit, used to determine a sample semantic unit set corresponding to the sample field according to the sample field and the sample semantic unit word contained in the sample field, wherein the sample semantic unit set includes at least one sample semantic unit, each of the sample semantic units includes a semantic unit identifier and a semantic unit feature, and the semantic unit feature is used to indicate a feature corresponding to the sample field or the sample semantic unit word and the semantic unit identifier; a determining unit, configured to determine a target semantic primitive dictionary according to the sample field and the sample semantic primitive set corresponding to the sample field; The device is also used for, after determining a target semantic primitive dictionary according to the sample field and the sample semantic primitive set corresponding to the sample field, obtaining a target annotation tag for annotating the corpus to be annotated, wherein the target annotation tag includes at least one target semantic primitive tag and an annotation matching degree corresponding to the target semantic primitive tag; in the target semantic primitive dictionary, sequentially determining candidate fields that match the annotation matching degree of the target semantic primitive tag; and if there is a target field matching the candidate field in the corpus to be annotated, annotating the target field, wherein the target field is a field obtained by segmenting the corpus to be annotated.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 7 when executed.
Citation Information
Patent Citations
A lexical semantic predicting method and device
CN108984533A
Word meaning processing method and device, electronic equipment and storage medium
CN112528670A