Classification Method of Coal Industry Thesaurus
Through the coal industry lexicon classification method, the initial lexicon and text data are processed using the text classification model, which solves the problem of different lexicon quality in the coal industry, and achieves the improvement of accurate lexicon classification and natural language processing accuracy.
Patent Information
- Application Number
- CN202510066573.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The high threshold, non-universality and professionalism of the coal industry make it difficult for ordinary industry vocabulary to identify key terms in the coal industry, affecting the accuracy of natural language processing, and the existing vocabulary quality is different, making it difficult to determine a high-quality industry vocabulary.
A coal industry lexicon classification method is proposed. By obtaining the initial lexicon, long sequence text and short sequence text, using the text classification model to obtain predictive classification information, and determining the target lexicon from the initial lexicon based on the label classification information and predictive classification information.
It realizes the accurate determination of the target lexicon from multiple initial lexicons, improves the classification effect of the coal industry lexicon, and improves the accuracy of natural language processing.
Smart Images

Figure CN119513321B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of natural language processing, and in particular to a coal industry vocabulary classification method. Background Art
[0002] Given the high threshold, non-universality and professionalism of the coal industry, some key terms proposed by coal industry researchers cannot be recognized by ordinary industry lexicons, which seriously affects the accuracy of natural language processing in the coal industry. In order to improve the accuracy of natural language processing in the coal industry, it is necessary to build a coal industry lexicon.
[0003] In the related technologies, the quality of coal industry related vocabulary varies. Therefore, how to determine high-quality industry vocabulary from multiple vocabulary has become a key research direction. Summary of the invention
[0004] The present disclosure aims to solve one of the technical problems in the related art at least to some extent.
[0005] To this end, the purpose of the present disclosure is to propose a coal industry vocabulary classification method, device, electronic equipment and storage medium.
[0006] The coal industry vocabulary classification method proposed in the first aspect of the present disclosure includes:
[0007] Acquire an initial vocabulary, a long sequence text, and a short sequence text, wherein the long sequence text has corresponding first annotation classification information, and the short sequence text has corresponding second annotation classification information;
[0008] Inputting the initial word library and the long sequence text into the first text classification model to obtain first predicted classification information output by the first text classification model;
[0009] Inputting the long sequence text into the first text classification model to obtain second predicted classification information output by the first text classification model;
[0010] Inputting the initial word library and the short sequence text into the second text classification model to obtain third predicted classification information output by the second text classification model;
[0011] Inputting the short sequence text into the second text classification model to obtain fourth prediction classification information output by the second text classification model;
[0012] A target vocabulary is determined from the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information.
[0013] The coal industry vocabulary classification device proposed in the second aspect of the present disclosure includes:
[0014] An acquisition module, used to acquire an initial vocabulary, a long sequence text and a short sequence text, wherein the long sequence text has corresponding first annotation classification information and the short sequence text has corresponding second annotation classification information;
[0015] A first input module, used for inputting the initial word library and the long sequence text into the first text classification model to obtain first predicted classification information output by the first text classification model;
[0016] A second input module, used for inputting the long sequence text into the first text classification model to obtain second predicted classification information output by the first text classification model;
[0017] A third input module, used for inputting the initial word library and the short sequence text into the second text classification model to obtain third predicted classification information output by the second text classification model;
[0018] A fourth input module, used for inputting the short sequence text into the second text classification model to obtain fourth prediction classification information output by the second text classification model;
[0019] The determination module is used to determine a target vocabulary from an initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information.
[0020] The electronic device proposed in the third aspect of the present disclosure includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the coal industry vocabulary classification method proposed in the first aspect of the present disclosure is implemented.
[0021] The fourth aspect of the present disclosure proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, it implements the coal industry vocabulary classification method proposed in the first aspect of the present disclosure.
[0022] The fifth embodiment of the present disclosure proposes a computer program product. When the instructions in the computer program product are executed by a processor, the coal industry vocabulary classification method proposed in the first embodiment of the present disclosure is executed.
[0023] The coal industry vocabulary classification method, device, electronic device and storage medium provided by the present disclosure have at least the following beneficial effects: by obtaining an initial vocabulary, a long sequence text and a short sequence text, wherein the long sequence text has corresponding first annotation classification information and the short sequence text has corresponding second annotation classification information, the initial vocabulary and the long sequence text are input into a first text classification model together to obtain first prediction classification information output by the first text classification model, the long sequence text is input into the first text classification model to obtain second prediction classification information output by the first text classification model, the initial vocabulary and the short sequence text are input into a second text classification model together to obtain third prediction classification information output by the second text classification model, the short sequence text is input into the second text classification model to obtain fourth prediction classification information output by the second text classification model, and a target vocabulary is determined from the initial vocabulary according to the first annotation classification information, the second annotation classification information, the first prediction classification information, the second prediction classification information, the third prediction classification information and the fourth prediction classification information, thereby being able to accurately determine the target vocabulary from multiple initial vocabularys and improve the classification effect of the coal industry vocabulary.
[0024] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description or learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and / or additional aspects and advantages of the present disclosure will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0026] Figure 1 It is a flow chart of a coal industry vocabulary classification method proposed in an embodiment of the present disclosure;
[0027] Figure 2 It is a flow chart of a coal industry vocabulary classification method proposed in another embodiment of the present disclosure;
[0028] Figure 3 It is a structural schematic diagram of a coal industry vocabulary classification device proposed in an embodiment of the present disclosure;
[0029] Figure 4 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0030] Embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present disclosure, and are not to be construed as limitations of the present disclosure. On the contrary, the embodiments of the present disclosure include all changes, modifications, and equivalents that fall within the spirit and connotation of the appended claims.
[0031] Figure 1 It is a flowchart of a coal industry vocabulary classification method proposed in an embodiment of the present disclosure.
[0032] Among them, it should be noted that the executor of the coal industry vocabulary classification method of this embodiment is the coal industry vocabulary classification device, which can be implemented by software and / or hardware. The device can be configured in an electronic device, and the electronic device can include but is not limited to a terminal, a server, etc. For example, the terminal can be a mobile phone, a handheld computer, etc.
[0033] like Figure 1 As shown, the coal industry vocabulary classification method includes:
[0034] S101: Acquire an initial vocabulary, a long sequence text, and a short sequence text, wherein the long sequence text has corresponding first annotation classification information, and the short sequence text has corresponding second annotation classification information.
[0035] Among them, the initial vocabulary refers to a pre-built vocabulary for coal mining industry scenarios (for example, coal mine underground operation scenarios, coal mine power transformation scenarios, etc.).
[0036] Among them, long sequence texts are longer in length and contain more information and details. They usually involve complex topics and structures, such as novels, academic papers, news reports, etc., and there are no restrictions on this.
[0037] Among them, the long sequence text is shorter in length and contains less information and details, such as article titles, entry introductions, text messages, etc., and there is no restriction on this.
[0038] Among them, the long sequence text has corresponding first annotation classification information, which is text classification information predetermined for the long sequence text. The text classification information can be used to describe the text type to which the long sequence text belongs. The text type can be, for example: security text, production text, basic introductory text, etc., without any restriction.
[0039] Among them, the short sequence text has corresponding second annotation classification information, which is text classification information predetermined for the short sequence text. The text classification information can be used to describe the text type to which the short sequence text belongs. The text type can be, for example: security text, production text, basic introductory text, etc., without any restriction.
[0040] In the disclosed embodiment, an algorithm may be written based on high-level industry journal papers, and 5,000 paper abstracts may be extracted using the Python language as long sequence texts. At the same time, each long sequence text may be classified to determine the first annotation classification information corresponding to the long sequence text.
[0041] In the disclosed embodiment, the short sequence text is obtained by using a large model after obtaining the long sequence text to perform relationship extraction on the long sequence text, simplifying it to a text with 30 words or less as the short sequence text, and adjusting and updating the original first annotation classification information of the long sequence text to determine the second annotation classification information corresponding to the short sequence text, without limitation.
[0042] Optionally, in some embodiments, obtaining the initial vocabulary may be obtaining an initial text, extracting initial keywords from the initial text, deleting function words in the initial keywords to obtain candidate keywords, determining the number of each candidate keyword, and determining the candidate keyword as a target keyword when the number is greater than a quantity threshold, wherein the target keyword constitutes the initial vocabulary.
[0043] The initial text may be, for example, historical journals and papers related to the coal industry, and there is no limitation to this.
[0044] Among them, the function words may be, for example, "keywords", "abstract", "document mark code", "document identification code", "article number", "article number", etc., without limitation.
[0045] In the disclosed embodiment, the initial text may be obtained, and then combined with the scene description information of the coal industry scene, regular expressions may be used to match the "keywords", and regular expressions may be used to determine the initial keywords from the birth text, and then the functional words in the initial keywords may be deleted to obtain candidate keywords.
[0046] In the embodiment of the present disclosure, after determining the candidate keywords, the number of each candidate keyword in the candidate keywords can be determined, and then the number of each candidate keyword can be determined. When the number is greater than the quantity threshold, the candidate keyword is determined as the target keyword, thereby obtaining an initial vocabulary consisting of the target keywords.
[0047] For example, in the embodiment of the present disclosure, multiple candidate keywords and the quantity of each candidate keyword can be determined: gas layer - quantity 184, numerical simulation - quantity 143, rock burst - quantity 118, coal mining machine - quantity - 47, wave position information - quantity 1, and then, the candidate keywords (gas layer, numerical simulation, rock burst and coal mining machine) whose quantity is greater than the quantity threshold (2) can be used as target keywords, without any restriction.
[0048] S102: Inputting the initial vocabulary and the long sequence text into the first text classification model to obtain first predicted classification information output by the first text classification model.
[0049] Among them, the first text classification model can be used to classify long sequence texts. The first text classification model can be, for example, a deep learning model such as a (Bidirectional Encoder Representations from Transformers, BERT) model (specifically, such as Bert, Bert_cnn, etc.), and there is no limitation to this.
[0050] In the disclosed embodiment, the first text classification model can be pre-trained by a sample long sequence text, that is, the BERT model can be trained based on a training set in the sample long sequence text, and the BERT model in the training phase can be tested based on a test set in the sample long sequence text. When it is determined that the BERT model training meets the standards, the trained BERT model is used as the first text classification model.
[0051] In the disclosed embodiment, the initial vocabulary and the long sequence text may be input into the first text classification model together to obtain the first predicted classification information output by the first text classification model.
[0052] S103: Inputting the long sequence text into the first text classification model to obtain second predicted classification information output by the first text classification model.
[0053] In the disclosed embodiment, the long sequence text may be input into the first text classification model to obtain the second predicted classification information output by the first text classification model.
[0054] S104: Inputting the initial vocabulary and the short sequence text into the second text classification model to obtain third predicted classification information output by the second text classification model.
[0055] The second text classification model may be used to classify short sequence texts. The second text classification model may be, for example, a support vector machine (Support Vector Machine, SVM) or other model, and there is no limitation to this.
[0056] That is to say, in the embodiment of the present disclosure, the initial vocabulary and the short sequence text may be input into the second text classification model together to obtain the third predicted classification information output by the second text classification model.
[0057] S105: Inputting the short sequence text into the second text classification model to obtain fourth predicted classification information output by the second text classification model.
[0058] In the embodiment of the present disclosure, the short sequence text may be input into the second text classification model to obtain the fourth predicted classification information output by the second text classification model.
[0059] S106: Determine a target vocabulary from the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information, and the fourth predicted classification information.
[0060] In an embodiment of the present disclosure, after determining the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information, the target vocabulary can be determined from the initial vocabulary based on the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information.
[0061] Optionally, in some embodiments, a target vocabulary is determined from an initial vocabulary based on the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information. The first classification accuracy of the first predicted classification information is determined based on the first labeled classification information and the first predicted classification information, the second classification accuracy of the second predicted classification information is determined based on the first labeled classification information and the second predicted classification information, the third classification accuracy of the third predicted classification information is determined based on the second labeled classification information and the third predicted classification information, the fourth classification accuracy of the fourth predicted classification information is determined based on the second labeled classification information and the third predicted classification information, and the target vocabulary is determined from the initial vocabulary based on the first classification accuracy, the second classification accuracy, the third classification accuracy and the fourth classification accuracy.
[0062] Among them, classification accuracy can be used to describe the classification accuracy of the text classification model.
[0063] That is, in the embodiment of the present disclosure, the first labeled classification information and the first predicted classification information may be compared to determine the first classification accuracy of the first predicted classification information ( P aL ), compare the first labeled classification information with the second predicted classification information to determine the second classification accuracy of the second predicted classification information (P aS ), compare the second labeled classification information with the third predicted classification information to determine the third classification accuracy of the third predicted classification information ( P L ), compare the second labeled classification information and the third predicted classification information to determine the fourth classification accuracy of the fourth predicted classification information ( P S ), and then, according to the second labeled classification information and the third predicted classification information, determine the fourth classification accuracy of the fourth predicted classification information, and according to the first classification accuracy, the second classification accuracy, the third classification accuracy and the fourth classification accuracy, determine the target vocabulary from the initial vocabulary.
[0064] Optionally, in some embodiments, after determining the target vocabulary from the initial vocabulary based on the first labeled classification information, the second labeled classification information, the first predicted classification information, and the second predicted classification information, the initial vocabulary corresponding to the first classification accuracy and / or the initial vocabulary corresponding to the third classification accuracy may be used as the target vocabulary when the first classification accuracy is greater than or equal to the second classification accuracy and / or the third classification accuracy is greater than or equal to the fourth classification accuracy.
[0065] That is to say, in the embodiment of the present disclosure, it can be P aL ≥P L , and / or P aS ≥P S In this case, it is determined that under the effect of the initial vocabulary, the model classification accuracy of the text classification model is improved, so that the initial vocabulary corresponding to the first classification accuracy and / or the initial vocabulary corresponding to the third classification accuracy can be used as the target vocabulary, thereby ensuring the availability of the target vocabulary.
[0066] In the disclosed embodiment, an initial vocabulary, a long sequence text and a short sequence text are obtained, wherein the long sequence text has corresponding first annotation classification information and the short sequence text has corresponding second annotation classification information, the initial vocabulary and the long sequence text are input into a first text classification model together to obtain first prediction classification information output by the first text classification model, the long sequence text is input into the first text classification model to obtain second prediction classification information output by the first text classification model, the initial vocabulary and the short sequence text are input into a second text classification model together to obtain third prediction classification information output by the second text classification model, the short sequence text is input into the second text classification model to obtain fourth prediction classification information output by the second text classification model, and a target vocabulary is determined from the initial vocabulary according to the first annotation classification information, the second annotation classification information, the first prediction classification information, the second prediction classification information, the third prediction classification information and the fourth prediction classification information. Thus, the target vocabulary can be accurately determined from multiple initial vocabularys, thereby improving the classification effect of the coal industry vocabulary.
[0067] Figure 2 It is a flowchart of a coal industry vocabulary classification method proposed in another embodiment of the present disclosure.
[0068] like Figure 2 As shown, the coal industry vocabulary classification method includes:
[0069] S201: Acquire an initial vocabulary, a long sequence text, and a short sequence text, wherein the long sequence text has corresponding first annotation classification information, and the short sequence text has corresponding second annotation classification information.
[0070] S202: Inputting the initial vocabulary and the long sequence text into the first text classification model to obtain first predicted classification information output by the first text classification model.
[0071] S203: Inputting the long sequence text into the first text classification model to obtain second predicted classification information output by the first text classification model.
[0072] S204: Inputting the initial vocabulary and the short sequence text into the second text classification model to obtain third predicted classification information output by the second text classification model.
[0073] S205: Inputting the short sequence text into the second text classification model to obtain fourth predicted classification information output by the second text classification model.
[0074] S206: Determine a target vocabulary from the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information, and the fourth predicted classification information.
[0075] The specific description of S201-S206 can be found in the above embodiment, which will not be repeated here.
[0076] S207: Determine a vocabulary type corresponding to the target vocabulary, wherein target vocabulary of different vocabulary types are applied to different scenarios.
[0077] The scene type may be, for example, a long text scene type, a short text scene type, etc., and there is no limitation on this.
[0078] In the disclosed embodiment, by determining the vocabulary type corresponding to the target vocabulary, a reference basis can be provided for vocabulary selection in different coal scenarios, that is, the target vocabulary of the corresponding vocabulary type can be accurately pushed for different scenario types, thereby effectively meeting the vocabulary requirements of different scenario types.
[0079] Optionally, in some embodiments, the word type corresponding to the target word library is determined by, when a first classification accuracy related to the target word library is greater than or equal to a second classification accuracy, and a third classification accuracy related to the target word library is greater than or equal to a fourth classification accuracy, then the word type of the target word library is determined to be the first type, wherein the first type of word library is applied to long text scenarios and short text scenarios, and when the first classification accuracy related to the target word library is greater than or equal to the second classification accuracy, and the third classification accuracy related to the target word library is less than the fourth classification accuracy, then the word type of the target word library is determined to be the second type, wherein the second type of word library is applied to long text scenarios, and when the first classification accuracy related to the target word library is less than the second classification accuracy, and the third classification accuracy related to the target word library is greater than or equal to the fourth classification accuracy, then the word type of the target word library is determined to be the third type, wherein the third type of word library is applied to short text scenarios.
[0080] That is to say, in the embodiment of the present disclosure, it can be P aL ≥P L , and P aS ≥P S In the case of, determine the word library type of the target word library is the first type, in P aL ≥P L , and P aS <P S In the case of, determine the vocabulary type of the target vocabulary is the second type, in P aS ≥P S ,andP aL <P L In this case, the vocabulary type of the target vocabulary is determined to be the third type.
[0081] In the embodiment of the present disclosure, an initial vocabulary, a long sequence text and a short sequence text are obtained, wherein the long sequence text has corresponding first annotation classification information and the short sequence text has corresponding second annotation classification information, the initial vocabulary and the long sequence text are input into a first text classification model together to obtain first prediction classification information output by the first text classification model, the long sequence text is input into the first text classification model to obtain second prediction classification information output by the first text classification model, the initial vocabulary and the short sequence text are input into a second text classification model together to obtain third prediction classification information output by the second text classification model, and the short sequence text is input into the second text classification model. In the classification model, the fourth predicted classification information output by the second text classification model is obtained, and the target vocabulary is determined from the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information. Thus, the target vocabulary can be accurately determined from multiple initial vocabularys, the construction effect of the coal industry vocabulary can be improved, and the vocabulary type corresponding to the target vocabulary can be determined. This can provide a reference for vocabulary selection in different coal scenarios, that is, it can accurately push target vocabulary of corresponding vocabulary types for different scenario types, thereby effectively meeting the vocabulary requirements of different scenario types.
[0082] Figure 3 It is a structural schematic diagram of a coal industry vocabulary classification device proposed in an embodiment of the present disclosure.
[0083] like Figure 3 As shown, the coal industry vocabulary classification device 30 includes:
[0084] An acquisition module 301 is used to acquire an initial vocabulary, a long sequence text and a short sequence text, wherein the long sequence text has corresponding first annotation classification information and the short sequence text has corresponding second annotation classification information;
[0085] A first input module 302, used to input the initial word library and the long sequence text into the first text classification model to obtain the first predicted classification information output by the first text classification model;
[0086] The second input module 303 is used to input the long sequence text into the first text classification model to obtain the second predicted classification information output by the first text classification model;
[0087] A third input module 304 is used to input the initial word library and the short sequence text into the second text classification model to obtain third predicted classification information output by the second text classification model;
[0088] A fourth input module 305, used to input the short sequence text into the second text classification model to obtain fourth prediction classification information output by the second text classification model;
[0089] The determination module 306 is used to determine a target vocabulary from the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information.
[0090] In some embodiments of the present disclosure, the acquisition module 301 is further used to:
[0091] Get the initial text;
[0092] Extract initial keywords from the initial text;
[0093] Delete function words from the initial keywords to obtain candidate keywords;
[0094] Determine the quantity of each candidate keyword;
[0095] When the quantity is greater than the quantity threshold, the candidate keyword is determined as the target keyword, wherein the target keyword constitutes the initial vocabulary.
[0096] In some embodiments of the present disclosure, the determination module 306 is further configured to:
[0097] Determining a first classification accuracy of the first predicted classification information according to the first labeled classification information and the first predicted classification information;
[0098] Determining a second classification accuracy of the second predicted classification information according to the first labeled classification information and the second predicted classification information;
[0099] Determining a third classification accuracy of the third predicted classification information according to the second labeled classification information and the third predicted classification information;
[0100] Determining a fourth classification accuracy of fourth predicted classification information according to the second labeled classification information and the third predicted classification information;
[0101] A target vocabulary is determined from the initial vocabulary according to the first classification accuracy, the second classification accuracy, the third classification accuracy, and the fourth classification accuracy.
[0102] In some embodiments of the present disclosure, the determination module 306 is further configured to:
[0103] When the first classification accuracy is greater than or equal to the second classification accuracy, and / or the third classification accuracy is greater than or equal to the fourth classification accuracy, the initial vocabulary corresponding to the first classification accuracy and / or the initial vocabulary corresponding to the third classification accuracy is used as the target vocabulary.
[0104] In some embodiments of the present disclosure, the determination module 306 is further configured to:
[0105] After determining a target vocabulary from the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, and the second predicted classification information, a vocabulary type corresponding to the target vocabulary is determined, wherein target vocabulary of different vocabulary types are applied to different scenarios.
[0106] In some embodiments of the present disclosure, the determination module 306 is further configured to:
[0107] When the first classification accuracy related to the target vocabulary is greater than or equal to the second classification accuracy, and the third classification accuracy related to the target vocabulary is greater than or equal to the fourth classification accuracy, the vocabulary type of the target vocabulary is determined to be the first type, wherein the vocabulary of the first type is applied to long text scenarios and short text scenarios;
[0108] When the first classification accuracy related to the target vocabulary is greater than or equal to the second classification accuracy, and the third classification accuracy related to the target vocabulary is less than the fourth classification accuracy, the vocabulary type of the target vocabulary is determined to be the second type, wherein the vocabulary type of the second type is applied to the long text scenario;
[0109] When the first classification accuracy related to the target vocabulary is less than the second classification accuracy, and the third classification accuracy related to the target vocabulary is greater than or equal to the fourth classification accuracy, the vocabulary type of the target vocabulary is determined to be the third type, wherein the third type of vocabulary is applied to short text scenarios.
[0110] It should be noted that the above explanation of the coal industry vocabulary classification method is also applicable to the coal industry vocabulary classification device of this embodiment, and will not be repeated here.
[0111] In the disclosed embodiment, an initial vocabulary, a long sequence text and a short sequence text are obtained, wherein the long sequence text has corresponding first annotation classification information and the short sequence text has corresponding second annotation classification information, the initial vocabulary and the long sequence text are input into a first text classification model together to obtain first prediction classification information output by the first text classification model, the long sequence text is input into the first text classification model to obtain second prediction classification information output by the first text classification model, the initial vocabulary and the short sequence text are input into a second text classification model together to obtain third prediction classification information output by the second text classification model, the short sequence text is input into the second text classification model to obtain fourth prediction classification information output by the second text classification model, and a target vocabulary is determined from the initial vocabulary according to the first annotation classification information, the second annotation classification information, the first prediction classification information, the second prediction classification information, the third prediction classification information and the fourth prediction classification information. Thus, the target vocabulary can be accurately determined from multiple initial vocabularys, thereby improving the classification effect of the coal industry vocabulary.
[0112] Figure 4 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0113] like Figure 4 As shown, the electronic device is in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors or processing units 16, a memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0114] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of a variety of bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (hereinafter referred to as: ISA) bus, Micro Channel Architecture (hereinafter referred to as: MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (hereinafter referred to as: VESA) local bus and Peripheral Component Interconnection (hereinafter referred to as: PCI) bus.
[0115] Electronic devices typically include a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, removable and non-removable media.
[0116] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 4 Not shown, usually called a "hard drive").
[0117] although Figure 4 Not shown in the figure, a disk drive for reading and writing a removable non-volatile disk (such as a "floppy disk"), and an optical disk drive for reading and writing a removable non-volatile optical disk (such as a compact disc read only memory (hereinafter referred to as: CD-ROM), a digital versatile disc read only memory (hereinafter referred to as: DVD-ROM) or other optical media) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data medium interfaces. The memory 28 may include at least one program product, which has a set (for example, at least one) of program modules, which are configured to perform the functions of the various embodiments of the present disclosure.
[0118] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including but not limited to an operating system, one or more application programs, other program modules, and program data, each of which or some combination thereof may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described in the present disclosure.
[0119] The electronic device may also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), may also communicate with one or more devices that enable the human body to interact with the electronic device, and / or communicate with any device that enables the electronic device to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed through an input / output (I / O) interface 22. In addition, the electronic device may also communicate with one or more networks (e.g., a local area network (Local Area Network; hereinafter referred to as: LAN), a wide area network (Wide Area Network; hereinafter referred to as: WAN) and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device through a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0120] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the coal industry vocabulary classification method mentioned in the above embodiment.
[0121] In order to implement the above embodiments, the present disclosure further proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the coal industry vocabulary classification method proposed in the above embodiments of the present disclosure is implemented.
[0122] In order to implement the above embodiments, the present disclosure further proposes a computer program product. When the instruction processor in the computer program product is executed, the coal industry vocabulary classification method proposed in the above embodiments of the present disclosure is executed.
[0123] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0124] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
[0125] It should be noted that, in the description of the present disclosure, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present disclosure, unless otherwise specified, the meaning of "plurality" is two or more.
[0126] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0127] It should be understood that the various parts of the present disclosure may be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0128] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0129] In addition, each functional unit in each embodiment of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0130] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0131] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0132] Although the embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.
Claims
1. A coal industry vocabulary classification method, characterized in that: The method comprises: Acquire multiple initial word libraries and long sequence texts. After acquiring the long sequence texts, use a large model to perform relationship extraction on the long sequence texts to obtain texts with a word count of no more than 30, and use the texts with a word count of no more than 30 as short sequence texts, wherein the initial word library includes a coal mine underground operation scene word library and / or a coal mine substation scene word library, the long sequence text is a high-level industry journal paper, the long sequence text has corresponding first annotation classification information, and the short sequence text has corresponding second annotation classification information. The method of constructing the initial word library includes: acquiring initial text, the initial text includes historical journals and papers related to the coal industry, combining scene description information of coal industry scenes, using regular expressions to extract initial keywords from the initial text, deleting function words in the initial keywords to obtain multiple candidate keywords, and when the number of the candidate keywords is greater than a quantity threshold, determining the candidate keywords as target keywords, and constructing the initial word library based on the multiple target keywords; Inputting the initial word library and the long sequence text into a first text classification model to obtain first predicted classification information output by the first text classification model; Inputting the long sequence text into the first text classification model to obtain second predicted classification information output by the first text classification model; Inputting the initial word library and the short sequence text into a second text classification model to obtain third predicted classification information output by the second text classification model; Inputting the short sequence text into the second text classification model to obtain fourth prediction classification information output by the second text classification model; A target vocabulary is determined from a plurality of the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information; wherein a first classification accuracy of the first predicted classification information is determined according to the first labeled classification information and the first predicted classification information, a second classification accuracy of the second predicted classification information is determined according to the first labeled classification information and the second predicted classification information, a third classification accuracy of the third predicted classification information is determined according to the second labeled classification information and the third predicted classification information, a fourth classification accuracy of the fourth predicted classification information is determined according to the second labeled classification information and the third predicted classification information, and the target vocabulary is determined from a plurality of the initial vocabulary according to the first classification accuracy, the second classification accuracy, the third classification accuracy and the fourth classification accuracy.
2. The method according to claim 1, characterized in that The step of determining a target vocabulary from the initial vocabulary according to the first classification accuracy, the second classification accuracy, the third classification accuracy, and the fourth classification accuracy includes: When the first classification accuracy is greater than or equal to the second classification accuracy, and / or the third classification accuracy is greater than or equal to the fourth classification accuracy, an initial vocabulary corresponding to the first classification accuracy and / or an initial vocabulary corresponding to the third classification accuracy is used as the target vocabulary.
3. The method according to claim 2, characterized in that After determining a target vocabulary from the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, and the second predicted classification information, the method further includes: A word library type corresponding to the target word library is determined, wherein target word libraries of different word library types are applied to different scenarios.
4. The method according to claim 3, characterized in that The determining of the vocabulary type corresponding to the target vocabulary includes: When the first classification accuracy associated with the target vocabulary is greater than or equal to the second classification accuracy, and the third classification accuracy associated with the target vocabulary is greater than or equal to the fourth classification accuracy, it is determined that the vocabulary type of the target vocabulary is the first type, wherein the vocabulary of the first type is applied to long text scenarios and short text scenarios; When the first classification accuracy related to the target vocabulary is greater than or equal to the second classification accuracy, and the third classification accuracy related to the target vocabulary is less than the fourth classification accuracy, it is determined that the vocabulary type of the target vocabulary is the second type, wherein the vocabulary of the second type is applied to the long text scenario; When the first classification accuracy associated with the target vocabulary is less than the second classification accuracy, and the third classification accuracy associated with the target vocabulary is greater than or equal to the fourth classification accuracy, it is determined that the vocabulary type of the target vocabulary is the third type, wherein the third type of vocabulary is applied to short text scenarios.
5. A coal industry vocabulary classification device, characterized in that: The device comprises: An acquisition module is used to acquire multiple initial word libraries and long sequence texts. After acquiring the long sequence texts, a large model is used to perform relationship extraction on the long sequence texts to obtain texts with a word count of no more than 30, and the texts with a word count of no more than 30 are used as short sequence texts, wherein the initial word library includes a coal mine underground operation scene word library and / or a coal mine substation scene word library, the long sequence text includes a paper, the long sequence text has corresponding first annotation classification information, and the short sequence text has corresponding second annotation classification information. The method of constructing the initial word library includes: acquiring initial text, the initial text includes historical journals and papers related to the coal industry, combining scene description information of coal industry scenes, extracting initial keywords from the initial text using regular expressions, deleting function words in the initial keywords to obtain multiple candidate keywords, and when the number of the candidate keywords is greater than a quantity threshold, determining the candidate keywords as target keywords, and constructing the initial word library based on the multiple target keywords; A first input module, used for inputting the initial word library and the long sequence text into a first text classification model to obtain first predicted classification information output by the first text classification model; A second input module, used for inputting the long sequence text into the first text classification model to obtain second predicted classification information output by the first text classification model; A third input module, used for inputting the initial word library and the short sequence text into the second text classification model to obtain third prediction classification information output by the second text classification model; a fourth input module, used for inputting the short sequence text into the second text classification model to obtain fourth prediction classification information output by the second text classification model; A determination module is used to determine a target vocabulary from a plurality of the initial vocabulary according to the first labeled classification information, the second labeled classification information, the first predicted classification information, the second predicted classification information, the third predicted classification information and the fourth predicted classification information; wherein a first classification accuracy of the first predicted classification information is determined according to the first labeled classification information and the first predicted classification information, a second classification accuracy of the second predicted classification information is determined according to the first labeled classification information and the second predicted classification information, a third classification accuracy of the third predicted classification information is determined according to the second labeled classification information and the third predicted classification information, a fourth classification accuracy of the fourth predicted classification information is determined according to the second labeled classification information and the third predicted classification information, and the target vocabulary is determined from a plurality of the initial vocabulary according to the first classification accuracy, the second classification accuracy, the third classification accuracy and the fourth classification accuracy.
6. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: in, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Voice data intention determination method and device, computer equipment and storage medium
CN110162633A
Text recognition method, device and equipment and storage medium
CN110909725A