Method and device for constructing retrieval model, electronic equipment and storage medium
By utilizing a standard terminology glossary and relational attributes to construct a retrieval model, the problem of poor retrieval results in existing technologies has been solved, enabling more professional, accurate, and rapid fault data retrieval.
Patent Information
- Application Number
- CN202310259984.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-03-16
AI Technical Summary
Existing fault database retrieval models based on natural language algorithms suffer from poor retrieval performance, resulting in a large amount of unnecessary search information.
By acquiring fault record text, matching and processing using a pre-established standard terminology vocabulary, determining the words to be used, and determining the association relationship of content words based on the relational attributes of relational words, a target retrieval model is constructed.
It improves the professionalism, accuracy, and speed of the retrieval model, reduces the search for redundant and invalid terms, and meets users' retrieval needs.
Smart Images

Figure CN116361339B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer processing, and in particular to a construction method and device of a retrieval model, an electronic device and a storage medium. BACKGROUND
[0002] At present, a large number of power distribution Internet of Things devices are usually deployed in the construction of new power systems. In order to ensure the stable operation of the power grid, fault data of the power distribution Internet of Things devices are usually collected, and unified modeling is performed based on the fault data to improve the processing speed of power grid faults.
[0003] The existing method for processing fault data is usually based on natural language algorithm to identify and process the fault data, determine the definition of entities in the fault data and the association relationship between the entities, and construct a database applied to fault retrieval. However, this language recognition-based method will cause a large amount of unnecessary search information in the constructed fault database, and there is a problem of poor retrieval effect. SUMMARY
[0004] The present application provides a construction method and device of a retrieval model, an electronic device and a storage medium to improve the professionalism, accuracy and effectiveness of the retrieval model, and achieve the technical effect of improving user retrieval requirements.
[0005] According to an aspect of the present application, a construction method of a retrieval model is provided, which comprises:
[0006] acquiring at least one fault record text and determining at least one to-be-processed word corresponding to the fault record text;
[0007] matching and processing each to-be-processed word based on a pre-established standard term vocabulary table to determine at least one to-be-used word; wherein the to-be-used word includes a relationship word and a content word;
[0008] determining the association relationship of at least two content words associated with the relationship word based on the relationship attribute of the relationship word;
[0009] determining a target retrieval model based on the at least two content words and the corresponding association relationship to determine a fault association word corresponding to to-be-retrieved information based on the target retrieval model.
[0010] According to another aspect of the present application, a construction device of a retrieval model is provided, which comprises:
[0011] a to-be-processed word determination module configured to acquire at least one fault record text and determine at least one to-be-processed word corresponding to the fault record text;
[0012] The to-be-used word determining module is configured to determine at least one to-be-used word by matching each of the to-be-processed words with a pre-established standard term vocabulary, wherein the to-be-used words include relationship words and content words;
[0013] The association relationship determining module is configured to determine an association relationship between the at least two content words associated with the relationship word based on a relationship attribute of the relationship word.
[0014] The search model determining module is configured to determine a target search model based on the at least two content words and the corresponding association relationship, so as to determine a fault association word corresponding to the to-be-searched information based on the target search model.
[0015] According to another aspect of the present application, an electronic device is provided, which comprises:
[0016] at least one processor; and
[0017] a memory connected to the at least one processor in communication; wherein
[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the search model construction method according to any one of the embodiments of the present application.
[0019] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the search model construction method according to any one of the embodiments of the present application when executed by the processor.
[0020] The technical scheme of the embodiment of the present application comprises the following steps: obtaining at least one fault record text, and determining at least one to-be-processed word corresponding to the fault record text; performing matching processing on each to-be-processed word based on a standard term vocabulary table, and determining at least one to-be-used word; determining the association relationship between at least two content words associated with the relationship word based on the relationship attribute of the relationship word; and determining a target retrieval model based on the at least two content words and the corresponding association relationship, so as to determine the fault association word corresponding to the to-be-retrieved information based on the target retrieval model. The problem that the retrieval model has poor retrieval effect due to the analysis of fault data by language recognition in the prior art is solved. The to-be-used word related to the professional field is extracted from the fault record text based on the standard term vocabulary table. Then, the association relationship between at least two content words associated with the relationship word is determined based on the relationship attribute of the relationship word in the to-be-used word. Then, the target retrieval model is constructed based on the association relationship between the content words. The fault association word retrieved based on the target retrieval model is related to the professional fault record term. The search of redundant and invalid words is reduced. The professional nature, accuracy and speed of the retrieval model are improved. The technical effect of improving the user retrieval demand is achieved.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. Those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0023] Figure 1 is a flow chart of a retrieval model construction method provided by the first embodiment of the present application;
[0024] Figure 2 is a flow chart of a retrieval model construction method provided by the second embodiment of the present application;
[0025] Figure 3 is a schematic diagram of a retrieval model construction method provided by the third embodiment of the present application;
[0026] Figure 4 is a schematic diagram for representing the association relationship of content words provided by the third embodiment of the present application;
[0027] Figure 5is a structural schematic diagram of a retrieval model construction device provided according to an embodiment four of the present application;
[0028] Figure 6 is a structural schematic diagram of an electronic device for implementing a retrieval model construction method of an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to enable personnel in the technical field to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the scope of protection of the present application.
[0030] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Embodiment one
[0032] Figure 1 A flowchart of a retrieval model construction method is provided for the embodiment one of the present application. The present embodiment can be applicable to the case of device fault record ontology model construction. The method can be performed by a retrieval model construction device, which can be realized in the form of hardware and / or software, and can be configured in a computing device. As shown in the figure, the method comprises: Figure 1
[0033] S110, at least one fault record text is acquired, and at least one to-be-processed word corresponding to the fault record text is determined.
[0034] The fault record text is associated with the fault information of the device and records the process information of the device from the occurrence of the fault to the repair of the fault. The device can be a power distribution Internet of Things device, such as a transformer, a high-voltage cabinet, a low-voltage cabinet, a bus bridge, a direct current screen, a high-voltage cable, and the like, which will not be described in detail. The fault record text can include device structure, fault mode, fault reason, fault time, fault influence, maintenance measure, fault diagnosis, and the like, and the text type can be Excel, Word, or txt.
[0035] In the embodiment, the text data in which the fault record is stored can be called from a preset position through an interface, that is, it is considered that the fault record text is obtained, and the number of texts can be one or multiple. Alternatively, it can also be considered that the fault record text is received when it is detected that a user triggers a text upload control of a system page. Further, the recorded data in the fault record text can be segmented to obtain multiple words, and each segmented word can be used as a to-be-processed word. The to-be-processed words can also be extracted from each segmented word, and used as to-be-processed words to construct a retrieval model based on the to-be-processed words. For example, the fault record text can be segmented by using a bidirectional matching method based on a Chinese dictionary, the fault record text is split into words, the length of the words can be two characters, three characters, six characters, and the like, and correspondingly, multiple to-be-processed words are obtained.
[0036] In S120, each to-be-processed word is matched based on a pre-established standard term vocabulary table to determine at least one to-be-used word.
[0037] The standard term vocabulary table can include a large number of professional power distribution Internet of Things device fault record text words, such as voltage value, serial number, shutdown, spark, warranty time, display fault, and the like, and the format can be Excel. The to-be-used words include relationship words and content words. The relationship words can be used to represent the relationship between objects, for example, the relationship words corresponding to parallel relationship can be: one side, and, also, and; the relationship words corresponding to the transition relationship can be: then, then, only, and; the relationship words corresponding to the transition relationship can be: although, but, although, however, and the like. The content words can be used to represent objects, such as device ownership, line, station building, terminal disconnection, and the like.
[0038] In order to ensure the standardization of the domain terms of the ontology model, the to-be-processed words and the standard term words in the standard term vocabulary table can be matched, and the to-be-processed words that are successfully matched can be used as to-be-used words. For example, a similarity calculation algorithm (such as cosine similarity, Euclidean distance, Pearson correlation coefficient, and the like) can be used to analyze the similarity between the to-be-processed words and each standard term word, and the to-be-processed words with a similarity higher than a preset threshold can be used as to-be-used words.
[0039] Exemplarily, the standard term vocabulary can include standard terms of different word numbers, and the standard term vocabulary is shown in Table 1.
[0040] Table 1
[0041] Word count Standard terminology glossary 1 And, also, again, moreover, and 2 Amplitude, shutdown, spark, error 3 Voltage value, current value, serial number 4 Version difference, warranty period, display failure … …
[0042] The to-be-processed words can be compared with the standard term vocabulary, and the related information of the fault equipment in the to-be-processed words is extracted, including the words of line, station building, type, equipment ID, and equipment ownership, and the related information of the fault processing in the to-be-processed words is extracted, including the words of time, defect phenomenon, defect level, defect cause, and defect elimination scheme, and these words are all used as to-be-used words.
[0043] In order to improve the coverage rate of the field-related words and improve the retrieval accuracy of the retrieval model, the to-be-processed words in the fault record text can be extracted by evaluating the importance of the words for the fault record text, so that the to-be-processed words have good fault distinguishing ability. In this embodiment, the implementation manner of determining at least one to-be-used word based on the matching processing of each to-be-processed word based on the pre-established standard term vocabulary can be: determining at least one to-be-used word from each to-be-processed word based on the similarity between each standard term vocabulary in the standard term vocabulary and each to-be-processed word; determining the inverse document frequency corresponding to the at least one to-be-used word, and determining at least one to-be-used word from the at least one to-be-used word based on the inverse document frequency and the preset word frequency threshold.
[0044] Wherein, the inverse document frequency refers to Term Frequency-Inverse Document Frequency (TF-IDF), which can be used to represent the importance and frequency of the words for the fault record text. For example, the importance of the words will increase in direct proportion to the number of times it appears in a certain fault record text, but at the same time, it will decrease in inverse proportion to the frequency of its appearance in all fault record texts.
[0045] In this embodiment, the similarity between each to-be-processed word and each standard term vocabulary in the standard term vocabulary can be calculated by using an algorithm, the to-be-processed word with a similarity higher than a preset threshold can be used as a to-be-used word, and further, the TF-IDF, i.e., the inverse document frequency, of the to-be-used word can be calculated based on the number of times the to-be-used word appears in a certain fault record text and the frequency of its appearance in all fault record texts, and the to-be-used word with an inverse document frequency higher than a preset word frequency threshold can be used as a to-be-used word, so as to extract the terms in the fault record text.
[0046] Specifically, the implementation manner of determining the inverse document term frequency corresponding to the at least one to-be-applied term can be: for each to-be-applied term, determining the occurrence frequency of the current to-be-applied term in the at least one fault record text, and based on the occurrence frequency and the number of terms in the at least one fault record text, determining the to-be-processed term frequency of the current to-be-applied term under each fault record text; based on the total number of texts of the at least one fault record text and the sub-number of texts containing the current to-be-applied term, determining the inverse document frequency; and based on the inverse document frequency and the to-be-processed term frequency, determining the inverse document term frequency corresponding to the current to-be-applied term.
[0047] It should be noted that the manner of determining the inverse document term frequency of each to-be-applied term is the same. To determine the inverse document term frequency of any to-be-applied term, the inverse document term frequency of the current to-be-applied term is determined. The inverse document frequency refers to the inverse document frequency (IDF), which can be used to represent the general importance of a term. The to-be-processed term frequency (TF) refers to the frequency of a term appearing in a document.
[0048] Specifically, the occurrence frequency of the current to-be-applied term in each fault record text can be calculated, and the total number of terms in each fault record text, i.e., the number of terms, can be calculated. The occurrence frequency of the current to-be-applied term in the current fault record text can be divided by the number of terms in the current fault record text, and the quotient can be taken as the to-be-processed term frequency of the current to-be-applied term under the current fault record text. Correspondingly, the to-be-processed term frequency of the current to-be-applied term under each fault record text can be obtained. The total number of texts of all fault record texts can be calculated, and the sub-number of texts containing the current to-be-applied term can be counted. The total number of texts and the sub-number of texts can be multiplied, and the logarithm of the product can be taken as the inverse document frequency. Each to-be-processed term frequency and the inverse document frequency can be multiplied to obtain a plurality of product values, and each product value can be taken as an inverse document term frequency. The inverse document term frequencies can be averaged to obtain an average value, which can be taken as an inverse document term frequency. The inverse document term frequency corresponding to the current to-be-applied term and the preset term frequency threshold can be compared, and the to-be-applied term whose inverse document term frequency is greater than the preset term frequency threshold is determined as the to-be-used term.
[0049] For example, the to-be-processed term frequency of the ith to-be-applied term under the jth fault record text can be calculated as follows: ij
[0050]
[0051] wherein, nij Let ∑ be the number of times the i-th word to be applied appears in the j-th fault record text (i.e., the frequency of occurrence). k n kj Let be the total number of words in the j-th fault record text. Then calculate the inverse document frequency (IDF). i As shown in formula (2) below:
[0052]
[0053] Where |D| represents the total number of fault record texts, i.e., the total number of texts, and |{j:t i ∈d j}| represents the total number of fault record texts containing the i-th term to be applied, i.e., the number of text sub-sub-terms. Next, the term frequency-inverse document frequency (i.e., inverse document term frequency, tf-idf) is calculated. i As shown in formula (3):
[0054] tf-idf i =tf ij ·idf i (3)
[0055] By calculating the term frequency-inverse document frequency (IF-IVF) of each word to be applied, high-frequency words with an IF-IVF greater than a preset term frequency threshold (e.g., 0.08) can be extracted from the words to be applied. The remaining words can be filtered out, and the extracted high-frequency words can be used as the words to be used to build a retrieval model.
[0056] S130. Based on the relational attributes of relational terms, determine the association relationship between at least two content terms associated with the relational terms.
[0057] Among them, relational attributes can be used to characterize the features of relational words, such as parallel, successive, progressive, selective, adversative, hypothetical, conditional, causal, etc.
[0058] In actual application, after obtaining the to-be-used words, the to-be-used words can be compared with each relationship marker in the relationship marker table, and the to-be-used words consistent with the relationship markers are extracted as the relationship words. The remaining to-be-used words except the relationship words can be regarded as the content words. Further, for each relationship word, the content words located before and after the current relationship word can be regarded as the content words associated with the current relationship word, and the association relationship between the content words associated with the current relationship word can be determined based on the association attribute of the current relationship word, and accordingly, the content words having the association relationship can be determined. For example, for the three words of “substation, and, equipment ownership”, “and” is the relationship word, representing a parallel relationship, and at this time, it is considered that “substation” and “equipment ownership” associated with “and” have a parallel relationship, i.e., the association relationship is a parallel relationship.
[0059] For example, the relationship marker table can include relationship markers of various association types, such as parallel relationship, transition relationship, progressive relationship, selection relationship, transition relationship, assumption relationship, condition relationship, cause-effect relationship, etc. For the relationship markers of the parallel relationship, it can include one side, and, also, etc. For the relationship markers of the transition relationship, it can include then, next, only, etc. For the relationship markers of the progressive relationship, it can include moreover, not only, even, but also, etc. For the relationship markers of the selection relationship, it can include or, or, etc. For the relationship markers of the transition relationship, it can include although, but, however, etc. For the relationship markers of the assumption relationship, it can include if, if, if, if, etc. For the relationship markers of the condition relationship, it can include only, only, etc. For the relationship markers of the cause-effect relationship, it can include because, because, because, because, etc. Of course, it is not limited to this. The relationship marker table is shown in Table 2 as follows:
[0060] Table 2
[0061] Association type Relationship marker word Parallel relationship On the one hand, and, also, again Continuation relationship Then, then, just Progressive relationship And, not only, also, even, but also Selection relationship Rather, or, or Transition relationship Although, however, although, however Hypothetical relationship If, if, if, if Conditional relationship Only, only if Cause and effect relationship Because, because, because, because
[0062] S140, determining a target retrieval model based on the at least two content words and the corresponding association relationship, to determine a fault association word corresponding to the to-be-retrieved information based on the target retrieval model.
[0063] The association relationship can be represented in the form of a function, for example, the parallel relationship can be represented by a function f and , the transition relationship can be represented by a function f kind-of , the progressive relationship can be represented by a function f further , the selection relationship can be represented by a function f either , the transition relationship can be represented by a function f but , the assumption relationship can be represented by a function f kwhileThe conditional relationship can be represented by a function f if The causal relationship can be represented by a function f cause The information to be retrieved can be a search term representing the user's specific information requirement, and the fault-related term refers to information related to the information to be retrieved, and the fault-related term can have an association relationship with the information to be retrieved.
[0064] In actual application, the content words containing the association relationship can be represented by a function corresponding to the association relationship, for example, for content words A and B, the association relationship is a parallel relationship, and the function representation of the association relationship can be A f and B. The power distribution Internet of Things device fault record ontology model, i.e., the target retrieval model, can be obtained based on the constructed function relationships. In actual application scenarios, the information to be retrieved can be input into the target retrieval model, the model can analyze the semantics of the information to be retrieved and search for words associated with the information to be retrieved, the words having an association relationship with the information to be retrieved can be used as fault-related terms, and the fault-related terms can be fed back to the user end.
[0065] The technical scheme of the embodiment, by acquiring at least one fault record text, and determining at least one to-be-processed word corresponding to the fault record text; based on the standard term vocabulary, each to-be-processed word is matched and processed to determine at least one to-be-used word; based on the relationship attribute of the relationship word, the association relationship between at least two content words associated with the relationship word is determined; based on the at least two content words and the corresponding association relationship, a target retrieval model is determined to determine the fault-related term corresponding to the information to be retrieved based on the target retrieval model, solves the problem that the existing technology analyzes fault data by language recognition to construct a retrieval model, resulting in poor retrieval effect of the retrieval model, realizes the extraction of the to-be-used word related to the professional field from the fault record text based on the standard term vocabulary, and then determines the association relationship between at least two content words associated with the relationship word based on the relationship attribute of the relationship word in the to-be-used word, and then constructs a target retrieval model based on the association relationship between the content words, so that the fault-related term retrieved based on the target retrieval model is related to the professional fault record term, reducing the search of redundant and invalid words, improving the professionalism, accuracy and speed of the retrieval model, and achieving the technical effect of improving the user's retrieval requirement.
[0066] Embodiment two
[0067] Figure 2 The flowchart of the construction method of the retrieval model provided in the second embodiment of the application is based on the foregoing embodiment, and S110 is further refined. The specific implementation can be referred to the technical scheme of the embodiment. Among them, the same or corresponding technical terms as the above embodiments are not described again.
[0068] like Figure 2 As shown, the method specifically includes the following steps:
[0069] S210. Obtain at least one fault record text.
[0070] S220. Segment the fault record text to obtain at least one statement to be used.
[0071] Specifically, the fault record text can be segmented into sentences using punctuation marks as separators, resulting in multiple sentences to be used within each fault record text. Punctuation marks can include periods (.), question marks (?), exclamation marks (!), commas (,), pause marks (、), semicolons (;), colons (:), quotation marks (“,”), parentheses [(),{}], dashes (——), ellipses (……), etc.
[0072] S230. For each statement to be used, the current statement to be used and the Chinese word database are matched based on the forward maximum matching algorithm to obtain at least one first word to be compared, and the current statement to be used and the Chinese word database are matched based on the reverse maximum matching algorithm to obtain at least one second word to be compared.
[0073] The Chinese lexicon can be a Chinese dictionary containing commonly used words. The forward maximum matching algorithm can be (FMM), and the reverse maximum matching algorithm can be (RMM).
[0074] It should be noted that the word segmentation method is the same for each statement to be used. Any one of the statements to be used can be used as the current statement to explain the word segmentation of the current statement.
[0075] In this embodiment, a bidirectional matching method based on a Chinese lexicon can be used to segment the current sentence to be used. A forward maximum matching algorithm can be used to match the current sentence to be used with the Chinese lexicon to obtain at least one first word to be compared. Specifically, this can be achieved by: inputting each character in the current sentence to be used into a first moving window in ascending order to obtain a first word to be matched corresponding to each first moving window; matching each first word to be matched with the Chinese lexicon to determine the matching result for each first word to be matched; and determining at least one first word to be compared based on each first word to be matched and its corresponding matching result.
[0076] The first moving window has a preset initial length, such as 6 or 7. It can contain text of the same length; for example, a window with a length of 6 can contain 6 characters. Matching results include successful and unsuccessful matches.
[0077] It should be noted that the first moving window is a movable window. When text is entered into the window one after another, each window will contain text of the same length as the window. The text within a window can be considered as a word, i.e., the first word to be matched.
[0078] Specifically, starting with the first character in the current sentence to be used, characters are entered into the first moving window in ascending order, first entering and then removing, until the last character in the current sentence is entered. Assuming the current sentence contains ten characters A, B, C, D, E, F, G, D, R, and T, and the preset initial length of the moving window is 6, then the first moving window contains ABCDEF. After A moves forward one position, G enters the window. The second moving window contains BCDEFG. After B moves forward one position, G enters the window. The third moving window contains CDEFGD. The fourth moving window contains DEFGDR. The fifth moving window contains EFGDRT. Correspondingly, five first words to be matched are obtained. Furthermore, each first word to be matched can be matched against words in a Chinese dictionary. If a word exists in the Chinese dictionary that matches a first word, the match is considered successful; otherwise, the match is considered unsuccessful. Correspondingly, the matching result of each first word to be matched can be determined. All first words to be matched that are successfully matched can be used as the first words to be compared. When the matching fails, the window length can be adjusted, the first words to be matched can be re-determined, and the new first words to be matched can be matched with the Chinese dictionary to determine at least one first word to be compared based on the matching result.
[0079] In this embodiment, the method for determining at least one first word to be compared based on each first word to be matched and the corresponding matching result can be as follows: if the matching result contains a failed match, then based on the first word to be matched that was successfully matched, determine at least one unmatched character in the current statement to be used; recombine each unmatched character to obtain a combined statement to be used; use the combined statement to be used as the current statement to be used again, and adjust the window length of the first moving window so that each character in the current statement to be used can be input into the first moving window in ascending order based on the adjusted first moving window, and determine the first word to be matched, so that the first word to be compared can be determined based on the first word to be matched and the corresponding matching result.
[0080] In this embodiment, after determining the matching result of each first word to be matched, all successfully matched first words to be matched can be selected as the first words to be compared. It is then determined whether there are any first words to be matched that have failed to match. If so, the text containing the first words to be compared in the failed first words to be matched can be removed, resulting in the remaining text, i.e., unmatched text. These unmatched texts can be combined to form a sentence to be used. Furthermore, the combined sentence to be used can be reused as the current sentence to be used, and the window length of the first moving window can be adjusted, for example, by reducing the window length by 1. The process involves re-executing the steps of inputting each character in the current statement to be used into the first moving window in ascending order, obtaining the first word to be matched corresponding to each first moving window, and matching each first word to be matched with a Chinese dictionary to determine the matching result corresponding to each first word to be matched. Based on each first word to be matched and the corresponding matching result, at least one first word to be compared is determined. All successfully matched first words to be matched can be used as the first word to be compared. The process can end when all matching results for the second words to be matched are successful, or when the window length of the first moving window is adjusted to a preset threshold. For example, the statement to be used can be segmented using a forward maximum matching algorithm. Starting from the first character of each statement to be used and ending at the last character, the statement is segmented into words by moving the "character window". The length of the "character window" is successively reduced from l=6 to l=1, and compared with the loaded Chinese dictionary to extract the first word to be compared. For example, Dir can be set as the Chinese dictionary and Len as the window length. Step 1: Set Len = 6. For each statement to be used, take a substring str of length 6 in the forward direction. Step 2: Match the substring str with the words in Dir. Step 3: If the match is successful, the substring str is considered successfully segmented. Move the pointer to the substring str forward by Len characters and return to Step 1. Step 4: If the match fails and Len > 1, decrement Len by 1, take a substring str of length Len from the statements to be used, and return to Step 2. Otherwise, obtain a word of length 1, move the pointer to the statements to be used forward by 1 character, and return to Step 1.
[0081] Alternatively, the reverse maximum matching algorithm can be used to match the current sentence to be used with the Chinese dictionary to obtain at least one second word to be compared. The specific implementation method is as follows: input each character in the current sentence to be used into the second moving window in reverse order to obtain the second word to be matched corresponding to each second moving window; match each second word to be matched with the Chinese dictionary to determine the matching result corresponding to each second word to be matched; and determine at least one second word to be compared based on each second word to be matched and the corresponding matching result.
[0082] Specifically, any character in the current sentence to be used can be randomly selected as the initial character. Starting from this initial character, the characters in the current sentence to be used can be input into the second moving window in reverse order. Correspondingly, the second words to be matched in each second moving window can be determined. The second words to be matched can be matched with the Chinese dictionary. All successfully matched second words to be matched are taken as the second words to be compared. If the match is unsuccessful, the unmatched characters in the current sentence to be used can be selected and recombine. The window length of the second moving window is reduced. The recombined sentence can be re-input into the adjusted second moving window. The process can end when the matching results of all second words to be matched are successful, or when the window length of the second moving window is adjusted to a preset threshold. For example, a reverse random matching algorithm is used to segment the sentence to be used. Starting from any character in the sentence and ending with the first character of the sentence, the sentence is divided into words by moving the "character window". The length of the "character window" is reduced from l=6 to l=1 in turn. The words are compared with the loaded Chinese dictionary to extract the second words to be compared.
[0083] S240. Based on at least one first word to be compared and at least one second word to be compared corresponding to the current statement to be used, determine at least one word to be processed.
[0084] In this embodiment, after determining the first and second words to be compared corresponding to the current statement to be used, each first word to be compared can be compared with the second word to be compared. Words that overlap can be used as words to be processed. If there are words that do not overlap, the words that do not overlap can be combined and the word segmentation can continue.
[0085] Specifically, the method for determining at least one word to be processed based on at least one first word to be compared and at least one second word to be compared corresponding to the current statement to be used can be as follows: compare at least one first word to be compared with at least one second word to be compared to determine at least one word to be filtered; if the word to be filtered contains non-overlapping words, combine all non-overlapping words to obtain a combined statement to be processed, and use the combined statement to be processed as the new current statement to be used, and repeatedly execute the operation of determining the first word to be compared and the second word to be compared based on the new current statement to be used, so that the word to be filtered is determined based on the first word to be compared and the second word to be compared, and the word to be processed is determined based on the overlapping words in the word to be filtered.
[0086] The words to be screened include overlapping words and / or non-overlapping words.
[0087] In this embodiment, at least one first word to be compared can be compared with at least one second word to be compared, and the overlapping words and non-overlapping words between them can be selected as words to be filtered. If the words to be filtered contain non-overlapping words, then all non-overlapping words can be combined, and the combined statement can be used as the current statement to be used. The current statement to be used is then sent to step S230, where analysis is performed based on the forward maximum matching algorithm and the reverse maximum matching algorithm to determine the first word to be compared and the second word to be compared. The process can end when there are no non-overlapping words in the first word to be compared and the second word to be compared, or when all non-overlapping words have the minimum character value. Alternatively, if the words to be filtered do not contain non-overlapping words, then all overlapping words can be used as words to be processed.
[0088] For example, the first word to be compared and the second word to be compared can be compared to extract overlapping words. The non-overlapping words can be further segmented using the forward maximum matching algorithm and the backward maximum matching algorithm until there are no words or the remaining words do not overlap. Then the segmentation can be stopped and the segmentation results can be output. All overlapping words are taken as words to be processed.
[0089] S250. Based on a pre-established standard terminology vocabulary, perform matching processing on each word to be processed to determine at least one word to be used.
[0090] S260. Based on the relational attributes of relational terms, determine the association relationship between at least two content terms associated with the relational terms.
[0091] S270. Based on at least two content terms and their corresponding associations, determine the target retrieval model, and then determine the fault-related terms corresponding to the information to be retrieved based on the target retrieval model.
[0092] The technical solution of this embodiment involves segmenting the fault record text into sentences to obtain at least one sentence to be used. For each sentence to be used, a forward maximum matching algorithm is used to match the current sentence to be used with a Chinese lexicon to obtain at least one first word to be compared. A reverse maximum matching algorithm is used to match the current sentence to be used with a Chinese lexicon to obtain at least one second word to be compared. Based on the at least one first word to be compared and at least one second word to be compared corresponding to the current sentence to be used, at least one word to be processed is determined. This achieves the goal of segmenting the sentences to be used by using forward and reverse bidirectional matching algorithms, thereby improving the accuracy of word segmentation, preventing word omissions, and ultimately improving the accuracy of the retrieval model construction.
[0093] Example 3
[0094] As an optional embodiment of the above embodiments, Figure 3This is a schematic diagram of the method for constructing a retrieval model according to Embodiment 3 of the present invention. For details, please refer to the following specific content.
[0095] See Figure 3The technical solution provided by this invention can establish a standard terminology glossary by referring to relevant documents of the power distribution network. For example, the terminology of fault records of power distribution IoT devices listed in the relevant documents of the power distribution network can be exported and a standard terminology glossary can be established, as shown in Table 1. At least one fault record text can also be segmented into sentences. For example, firstly, a regular expression tokenizer (RegexpTokenizer) can be used to segment the fault record text into sentences with punctuation marks as intervals, dividing it into sentences. Multiple sentences to be used are obtained as units of sentences. Then, the sentences to be used are segmented using a bidirectional matching method based on a Chinese dictionary (i.e., a Chinese word database). For example, the sentences to be used can be segmented using a forward maximum matching algorithm. Starting from the first character of each sentence to be used and ending at the last character, the sentences are divided into words by moving the "character window". The length of the "character window" is successively reduced from l=6 to l=1. The words are compared with the loaded Chinese dictionary to extract the first word to be compared. The steps can be: Let Dir be the loaded Chinese dictionary. Step 1: Set Len = 6, and take a string str of length 6 from each sentence. Step 2: Match str with the words in Dir. Step 3: If the match is successful, the string is considered successfully segmented. Move the pointer to the sentence to be segmented forward by Len characters and return to Step 1. Step 4: If unsuccessful, if Len > 1, decrement Len by 1, take a string str of length Len from the sentence to be segmented, and return to Step 2. Otherwise, obtain a word of length 1, move the pointer to the sentence to be segmented forward by 1 character, and return to Step 1. Alternatively, a reverse random matching algorithm (i.e., reverse maximum matching algorithm) can be used for word segmentation. Starting from any character in the sentence and ending at the first character, the sentence is divided into words by moving the "character window". The length of the "character window" is successively reduced from l = 6 to l = 1, and compared with the loaded Chinese dictionary to extract the second word to be compared. Furthermore, the forward maximum matching result and the reverse random matching result of the bidirectional matching method can be compared to extract overlapping and non-overlapping words from the first and second words to be compared. It is then determined whether the non-overlapping words are empty or consist entirely of single characters. If so, the non-overlapping words are segmented using both the forward maximum matching algorithm and the reverse random matching algorithm until they are empty or consist entirely of single characters. The segmentation is then stopped, and all overlapping words are output as the segmentation result, thus obtaining the words to be processed. Further, the words to be processed are compared with a standard terminology vocabulary to extract relevant information about the faulty equipment, including terms such as line, station, type, equipment ID, and equipment ownership, as well as fault handling information, including terms such as time, defect phenomenon, defect level, defect cause, and fault mitigation plan. The extracted words are then used as the words to be applied. The frequency (tf) of the i-th word to be applied in the j-th fault record text can be calculated.ij The method can be shown in the following formula (1):
[0096]
[0097] Where, n ij Let ∑ be the number of times the i-th word to be applied appears in the j-th fault record text (i.e., the frequency of occurrence). k n kj Let be the total number of words in the j-th fault record text. Then calculate the inverse document frequency (IDF). i As shown in formula (2) below:
[0098]
[0099] Where |D| represents the total number of fault record texts, i.e., the total number of texts, and |{j:t i ∈d j}| represents the total number of fault record texts containing the i-th term to be applied, i.e., the number of text sub-sub-terms. Next, the term frequency-inverse document frequency (i.e., inverse document term frequency, tf-idf) is calculated. i As shown in formula (3):
[0100] tf-idf i =tf ij ·idf i (3)
[0101] By calculating the term frequency-inverse document frequency (IF-IVF) of each term to be applied, high-frequency words with an IF-IVF greater than a preset term frequency threshold (e.g., 0.08) can be extracted, and the remaining words can be filtered out. These extracted high-frequency words can then be used as the terms to be applied. Relationship markers (i.e., relational terms) are extracted from the terms to be applied, and the remaining words are used as ontology units (i.e., content terms) in the power distribution IoT device fault records. Based on the relational terms, the functional relationships (i.e., association relationships) between different content terms are determined, and an ontology model for the power distribution IoT device fault records is established, i.e., the target retrieval model. For an example, see [link to relevant documentation]. Figure 4 , Figure 4 This can be represented as a diagram illustrating the relationships between words representing content. The relationship between "January 1, 2021, 10kV TangXX, DKXX13, T23XXXXXX001202005060025" is a progressive relationship, which can be represented by the function f. further This indicates that the relationship between "software version mismatch," "terminal disconnection," and "critical defect" is a causal relationship, which can be expressed using the function f. cause The relationship between "reconnecting the device, upgrading the software version, and replacing the matching device" can be represented by the function f. eitherThis technical solution uses a bidirectional matching method based on a Chinese dictionary to segment the fault record text of power distribution IoT equipment. After comparing it with a standard terminology vocabulary, candidate terms for faulty equipment are extracted (including words such as line, station, type, equipment ID, equipment ownership, and fault handling information). The inverse document frequency of the candidate terms is then calculated to extract the terms in the power distribution IoT equipment fault records. Based on relational markers, the functional relationships between different ontology units are determined, and an ontology model of power distribution IoT equipment fault records is established. This enables unified management and utilization of fault record text and rapid fault diagnosis.
[0102] The technical solution of this embodiment obtains at least one fault record text and determines at least one word to be processed corresponding to the fault record text; it performs matching processing on each word to be processed based on a standard terminology vocabulary to determine at least one word to be used; it determines the association relationship of at least two content words associated with the related words based on the relational attributes of the related words; and it determines a target retrieval model based on the at least two content words and their corresponding association relationships. Based on the target retrieval model, it determines fault-related words corresponding to the information to be retrieved. This solves the problem in existing technologies where analyzing fault data through language recognition to construct a retrieval model results in poor retrieval performance. It achieves the extraction of words to be used related to the professional field from fault record text based on a standard terminology vocabulary, and then determines the association relationship of at least two content words associated with the related words based on the relational attributes of the related words. Based on the association relationships between the content words, it constructs a target retrieval model, ensuring that the fault-related words retrieved based on the target retrieval model are related to professional fault record terminology. This reduces the search for redundant and invalid words, improves the professionalism, accuracy, and speed of the retrieval model, and ultimately enhances the technical effect of meeting user retrieval needs.
[0103] Example 4
[0104] Figure 5 This is a schematic diagram of a retrieval model construction device according to Embodiment 4 of the present invention. Figure 5 As shown, the device includes: a word determination module 610, a word determination module 620, a relation determination module 630, and a retrieval model determination module 640.
[0105] The module 610 is used to acquire at least one fault record text and determine at least one word to be processed corresponding to the fault record text; the module 620 is used to match each word to be processed based on a pre-established standard terminology vocabulary to determine at least one word to be used; wherein the word to be used includes relational words and content words; the module 630 is used to determine the association relationship of at least two content words associated with the relational words based on the relational attributes of the relational words; and the module 640 is used to determine a target retrieval model based on the at least two content words and the corresponding association relationship, so as to determine fault-related words corresponding to the information to be retrieved based on the target retrieval model.
[0106] The technical solution of this embodiment obtains at least one fault record text and determines at least one word to be processed corresponding to the fault record text; it performs matching processing on each word to be processed based on a standard terminology vocabulary to determine at least one word to be used; it determines the association relationship of at least two content words associated with the related words based on the relational attributes of the related words; and it determines a target retrieval model based on the at least two content words and their corresponding association relationships. Based on the target retrieval model, it determines fault-related words corresponding to the information to be retrieved. This solves the problem in existing technologies where analyzing fault data through language recognition to construct a retrieval model results in poor retrieval performance. It achieves the extraction of words to be used related to the professional field from fault record text based on a standard terminology vocabulary, and then determines the association relationship of at least two content words associated with the related words based on the relational attributes of the related words. Based on the association relationships between the content words, it constructs a target retrieval model, ensuring that the fault-related words retrieved based on the target retrieval model are related to professional fault record terminology. This reduces the search for redundant and invalid words, improves the professionalism, accuracy, and speed of the retrieval model, and ultimately enhances the technical effect of meeting user retrieval needs.
[0107] Based on the above-mentioned device, optionally, the word determination module 610 to be processed includes a sentence determination unit to be used, a word determination unit to be compared, and a word determination unit to be processed.
[0108] The statement to be used determination unit is used to segment the fault record text into sentences to obtain at least one statement to be used.
[0109] The word to be compared unit is used to match the current statement to be used with the Chinese dictionary based on the forward maximum matching algorithm to obtain at least one first word to be compared, and to match the current statement to be used with the Chinese dictionary based on the reverse maximum matching algorithm to obtain at least one second word to be compared.
[0110] The word to be processed determination unit is used to determine at least one word to be processed based on at least one first word to be compared and at least one second word to be compared corresponding to the currently used statement.
[0111] Based on the above-mentioned device, optionally, the word to be compared determination unit includes a first word to be matched determination subunit, a matching result determination subunit, and a first word to be compared determination subunit.
[0112] The first word to be matched determination subunit is used to input each character in the current sentence to be used into the first moving window in ascending order to obtain the first word to be matched corresponding to each first moving window; wherein, the window length of the first moving window is a preset initial length;
[0113] The matching result determination subunit is used to match each first word to be matched with the Chinese word database and determine the matching result corresponding to each first word to be matched.
[0114] The first matching word determination subunit is used to determine at least one first matching word based on each first matching word and the corresponding matching result; wherein the matching result includes successful matching and failed matching.
[0115] Based on the above-mentioned device, optionally, the first sub-unit for determining words to be compared includes: a sub-unit for determining unmatched text, a sub-unit for determining combined sentences to be used, and a sub-unit for adjusting window length.
[0116] The unmatched text determination unit is used to determine at least one unmatched text in the currently used statement based on the first successfully matched word if the matching result contains a failed match.
[0117] The small units to be determined by the combination statement are used to recombine the unmatched characters to obtain the combination statement to be used.
[0118] The window length adjustment unit is used to re-use the combined statement to be used as the current statement to be used, and adjust the window length of the first moving window, so that the characters in the current statement to be used are re-input into the first moving window in ascending order based on the adjusted first moving window, and determine the first word to be matched, so that the first word to be compared is determined based on the first word to be matched and the corresponding matching results.
[0119] Optionally, based on the above-mentioned device, the word-to-be-matched determination unit may further include a second word-to-be-matched determination subunit, a result determination subunit, and a second word-to-be-matched determination subunit.
[0120] The second word to be matched determination subunit is used to input each character in the current sentence to be used into the second moving window in reverse order to obtain the second word to be matched corresponding to each second moving window;
[0121] The result determination subunit is used to match each second word to be matched with the Chinese word database and determine the matching result corresponding to each second word to be matched.
[0122] The second matching word determination subunit is used to determine at least one second matching word based on each second matching word and the corresponding matching result.
[0123] Based on the above-mentioned device, optionally, the word determination unit includes: a word selection subunit and a word processing subunit.
[0124] The subunit for determining words to be screened is used to compare the at least one first word to be compared with the at least one second word to be compared, and determine at least one word to be screened; wherein, the words to be screened include overlapping words and / or non-overlapping words;
[0125] The subunit for determining words to be processed is used to combine all non-overlapping words if the words to be filtered contain non-overlapping words to obtain a combined statement to be processed, and to use the combined statement to be processed as a new current statement to be used. Based on the new current statement to be used, the operation of determining the first and second words to be compared is repeatedly executed, so that the words to be filtered are determined based on the first and second words to be compared, and the words to be processed are determined based on the overlapping words in the words to be filtered.
[0126] Based on the above-mentioned device, optionally, the word determination module 620 includes: a word determination unit to be applied and a word determination unit to be processed.
[0127] The word to be applied determination unit is used to determine at least one word to be applied from each of the words to be processed based on the similarity between each standard term vocabulary in the standard term vocabulary list and each word to be processed;
[0128] The word to be used determination unit is used to determine the inverse document word frequency corresponding to the at least one word to be applied, and to determine at least one word to be used from the at least one word to be applied based on the inverse document word frequency and a preset word frequency threshold.
[0129] Based on the above-mentioned device, optionally, the word determination unit to be used includes a word frequency determination subunit, an inverse document frequency determination subunit, and an inverse document word frequency determination subunit.
[0130] The word frequency determination subunit is used to determine the frequency of occurrence of the current word to be applied in the at least one fault record text for each word to be applied, and to determine the word frequency to be applied in each fault record text based on the frequency of occurrence and the number of words in the at least one fault record text.
[0131] The inverse document frequency determination subunit is used to determine the inverse document frequency based on the total number of texts in the at least one fault record text and the number of text sub-texts in the fault record text containing the currently applied word.
[0132] The inverse document frequency determination subunit is used to determine the inverse document frequency corresponding to the current word to be applied based on the inverse document frequency and the frequency of each word to be processed.
[0133] The retrieval model construction apparatus provided in the embodiments of the present invention can execute the retrieval model construction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0134] Example 5
[0135] Figure 6 This is a schematic diagram of the structure of an electronic device that implements the method for constructing the retrieval model according to embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0136] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0137] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0138] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for retrieving model construction methods.
[0139] In some embodiments, the method for constructing the retrieval model may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for constructing the retrieval model described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the method for constructing the retrieval model by any other suitable means (e.g., by means of firmware).
[0140] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0141] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0142] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0143] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0144] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0145] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0146] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0147] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for constructing a retrieval model, characterized in that, include: Obtain at least one fault record text and determine at least one word to be processed corresponding to the fault record text, including: segmenting the fault record text into sentences to obtain at least one sentence to be used; for each sentence to be used, matching the current sentence to be used with a Chinese dictionary based on a forward maximum matching algorithm to obtain at least one first word to be compared, and matching the current sentence to be used with the Chinese dictionary based on a backward maximum matching algorithm to obtain at least one second word to be compared; based on the at least one first word to be compared and the at least one second word to be compared corresponding to the current sentence to be used, determining at least one word to be processed, including: The at least one first word to be compared is compared with the at least one second word to be compared to determine at least one word to be filtered; wherein, the words to be filtered include overlapping words and non-overlapping words; if the words to be filtered contain non-overlapping words, all non-overlapping words are combined to obtain a combined statement to be processed, and the combined statement to be processed is used as a new current statement to be used, and the operation of determining the first word to be compared and the second word to be compared is repeated based on the new current statement to be used, so that the words to be filtered are determined based on the first word to be compared and the second word to be compared, and the words to be processed are determined based on the overlapping words in the words to be filtered; Based on a pre-established standard terminology vocabulary, each of the words to be processed is matched to determine at least one word to be used; wherein, the word to be used includes relational words and content words; Based on the relational attributes of the relational terms, determine the association relationship between at least two content terms associated with the relational terms; Based on the at least two content terms and their corresponding relationships, a target retrieval model is determined, and fault-related terms corresponding to the information to be retrieved are determined based on the target retrieval model.
2. The method according to claim 1, characterized in that, The forward maximum matching algorithm is used to match the current sentence to be used with the Chinese lexicon to obtain at least one first word to be compared, including: Each character in the currently used sentence is input into the first moving window in ascending order to obtain the first word to be matched corresponding to each first moving window; wherein, the window length of the first moving window is a preset initial length; Each first word to be matched is matched with the Chinese dictionary to determine the matching result corresponding to each first word to be matched. Based on each first word to be matched and the corresponding matching result, at least one first word to be compared is determined; wherein, the matching result includes successful matching and failed matching.
3. The method according to claim 2, characterized in that, Based on each first word to be matched and the corresponding matching result, at least one first word to be compared is determined, including: If the matching result contains a failed match, then based on the first successfully matched word to be matched, at least one unmatched character in the current statement to be used is determined; Recombine the unmatched characters to obtain the combined sentence to be used; The combined statement to be used is re-designated as the current statement to be used, and the window length of the first moving window is adjusted. Based on the adjusted first moving window, the characters in the current statement to be used are re-input into the first moving window in ascending order to determine the first word to be matched, so that the first word to be compared is determined based on the first word to be matched and the corresponding matching results.
4. The method according to claim 1, characterized in that, The reverse maximum matching algorithm is used to match the current sentence to be used with the Chinese lexicon to obtain at least one second word to be compared, including: Input each character in the current sentence to be used into the second moving window in reverse order to obtain the second word to be matched corresponding to each second moving window; Each second word to be matched is matched with the Chinese dictionary to determine the matching result corresponding to each second word to be matched. Based on each second word to be matched and the corresponding matching results, at least one second word to be compared is determined.
5. The method according to claim 1, characterized in that, The matching process based on a pre-established standard terminology vocabulary for each of the words to be processed determines at least one word to be used, including: Based on the similarity between each standard term in the standard term vocabulary and each of the words to be processed, at least one word to be applied is determined from each of the words to be processed; Determine the inverse document frequency corresponding to the at least one word to be applied, and based on the inverse document frequency and a preset word frequency threshold, determine at least one word to be used from the at least one word to be applied.
6. The method according to claim 5, characterized in that, Determining the inverse document term frequency corresponding to the at least one word to be applied includes: For each word to be applied, determine the frequency of occurrence of the current word to be applied in the at least one fault record text, and based on the frequency of occurrence and the number of words in the at least one fault record text, determine the word frequency to be processed of the current word to be applied in each fault record text; The inverse document frequency is determined based on the total number of texts in the at least one fault record text and the number of text sub-texts in the fault record text containing the currently applied word; Based on the inverse document frequency and the frequency of each word to be processed, the inverse document frequency corresponding to the current word to be applied is determined.
7. An apparatus for constructing a retrieval model, characterized in that, include: The pending word determination module is used to acquire at least one fault record text and determine at least one pending word corresponding to the fault record text; The pending word determination module includes: a pending sentence determination unit, used to segment the fault record text into sentences to obtain at least one pending sentence; a pending comparison word determination unit, used to match the current pending sentence with a Chinese dictionary based on a forward maximum matching algorithm for each pending sentence to obtain at least one first pending comparison word, and to match the current pending sentence with the Chinese dictionary based on a reverse maximum matching algorithm to obtain at least one second pending comparison word; and a pending word determination unit, used to determine at least one pending word based on at least one first pending comparison word and at least one second pending comparison word corresponding to the current pending sentence. The word-to-be-processed determination unit includes: a word-to-be-filtered determination subunit, configured to compare the at least one first word to be compared with the at least one second word to be compared, and determine at least one word to be-filtered; wherein the words to be-filtered include overlapping words and non-overlapping words; the word-to-be-processed determination subunit is configured to, if the words to be-filtered contain non-overlapping words, combine all non-overlapping words to obtain a combined statement to be-processed, and use the combined statement to be-processed as a new current statement to be used, and repeatedly execute the operation of determining the first and second words to be compared based on the new current statement to be used, so as to determine the words to be-filtered based on the first and second words to be compared, and to determine the words to be-processed based on the overlapping words in the words to be-filtered; The word to be used determination module is used to perform matching processing on each of the words to be processed based on a pre-established standard term vocabulary to determine at least one word to be used; wherein, the words to be used include relational words and content words; The association relationship determination module is used to determine the association relationship between at least two content words associated with the relation words based on the relation attributes of the relation words; The retrieval model determination module is used to determine a target retrieval model based on the at least two content terms and their corresponding associations, so as to determine fault-related terms corresponding to the information to be retrieved based on the target retrieval model.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for constructing the retrieval model according to any one of claims 1-6.
Citation Information
Patent Citations
Power distribution Internet of Things low-voltage equipment inspection record ontology model construction method and system, and storage medium
CN115374963A
Hash index
US20180011893A1