A method, device, equipment and medium for constructing a knowledge base for aircraft manufacturing

By performing clustering and inverse clustering analysis on the initial process case set, screening and extracting key information, the problem of imperfect existing knowledge base was solved, and a more accurate and complete aircraft manufacturing knowledge base was constructed.

CN118069835BActive Publication Date: 2025-10-24CHENGDU AIRCRAFT INDUSTRY GROUP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410081137.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-10-24
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

The knowledge base for aircraft manufacturing constructed using existing methods is not perfect, resulting in a decline in R&D and manufacturing efficiency and quality.

Method used

By performing the first cluster analysis and the first inverse cluster analysis on the initial process case set, the first clusters with large similarity and the first inverse clusters with small similarity are screened out, the inverse clusters with small similarity are eliminated, and cluster analysis is performed based on the re-determined cluster center information to form a target knowledge sub-base, and finally the key information is extracted to form the target knowledge base.

Benefits of technology

It retains more process cases related to aircraft manufacturing, provides a more complete knowledge base, and improves the accuracy and completeness of the knowledge base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118069835B_ABST
    Figure CN118069835B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for constructing a knowledge base for aircraft manufacturing, equipment and a medium, relates to the technical field of databases, and solves the technical problem that the knowledge base for aircraft manufacturing obtained by the existing method is not perfect. The method comprises the following steps: performing first clustering analysis processing on an initial process case set based on initial clustering center information to obtain a first cluster; performing first inverse clustering analysis processing on the initial process case set based on initial inverse clustering center information to obtain a first inverse cluster; obtaining the initial clustering center information based on the first cluster; obtaining a target knowledge sub-library based on the initial clustering center information; extracting key information from the target knowledge sub-library to obtain target key word information; and obtaining a target knowledge base based on the target knowledge sub-library and the target key word information. Therefore, the application can retain more process cases related to composite materials for aircraft manufacturing and provide a more perfect knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of database construction, and particularly relates to a construction method, device and equipment of a knowledge base for aircraft manufacturing and a medium. BACKGROUND

[0002] At present, composite materials are the most widely used and most promising raw materials in aircraft production and manufacturing. Composite materials refer to new materials composed of two or more than two metals or non-metals as raw materials, according to the material properties, using physical or chemical methods to form new materials that can be applied in different scenarios, which can fully exert the advantages of various materials and improve the overall performance of the materials.

[0003] When the composite material knowledge base for aircraft manufacturing is constructed by using the current method, the related process cases are usually directly obtained from the composite material process cases according to the keyword information, and the knowledge base is constructed by integrating the related process cases.

[0004] However, the above method is prone to cause the loss of process cases, so that the knowledge base is not perfect enough, and the efficiency and quality of aircraft research and development and manufacturing are affected. SUMMARY

[0005] The main purpose of the present application is to provide a construction method, device, equipment and medium of a knowledge base for aircraft manufacturing, which aims to solve the technical problem that the knowledge base for aircraft manufacturing constructed by the existing method is not perfect enough.

[0006] To achieve the above purpose, the present application provides a construction method of a knowledge base for aircraft manufacturing, comprising the following steps:

[0007] obtaining initial clustering center information; wherein the initial clustering center information is obtained based on an initial process case set;

[0008] based on the initial clustering center information, performing first clustering analysis processing on the initial process case set to obtain a first cluster;

[0009] obtaining initial inverse clustering center information; based on the initial inverse clustering center information, performing first inverse clustering analysis processing on the initial process case set to obtain a first inverse cluster; deleting the first inverse cluster from the initial process case set to obtain a first process case set; wherein the initial inverse clustering center information is obtained based on the initial clustering center information;

[0010] based on the first cluster, obtaining first clustering center information; based on the first clustering center information, obtaining a target knowledge sub-base; wherein the target knowledge sub-base is obtained by performing second clustering analysis processing on the first process case set based on the first clustering center information;

[0011] performing key information extraction on the target knowledge sub-library to obtain target keyword information;

[0012] obtaining a target knowledge library based on the target knowledge sub-library and the target keyword information.

[0013] Optionally, the first clustering analysis processing on the initial process case set based on the initial clustering center information to obtain a first cluster includes:

[0014] obtaining a first cosine similarity value based on the initial clustering center information;

[0015] comparing the first cosine similarity value with a first preset threshold value, and performing the first clustering analysis processing on the initial process case set to obtain the first cluster.

[0016] Optionally, the first inverse clustering analysis processing on the initial process case set based on the initial inverse clustering center information to obtain a first inverse cluster includes:

[0017] obtaining a second cosine similarity value based on the initial inverse clustering center information;

[0018] comparing the second cosine similarity value with a second preset threshold value, and performing the first inverse clustering analysis processing on the initial process case set to obtain the first inverse cluster.

[0019] Optionally, before the obtaining of the first clustering center information based on the first cluster and the obtaining of the target knowledge sub-library based on the first clustering center information, the method further includes:

[0020] obtaining an initial target threshold value based on the second preset threshold value, wherein the initial target threshold value is obtained based on the second preset threshold value and a preset threshold value growth step;

[0021] determining whether the initial target threshold value exceeds a preset range threshold value;

[0022] if yes, performing second clustering analysis processing on the first process case set based on the first clustering center information to obtain the target knowledge sub-library;

[0023] if no, performing second clustering analysis processing and second inverse clustering analysis processing on the first process case set based on the first clustering center information to obtain an iterated target threshold value, wherein the iterated target threshold value is greater than the preset range threshold value;

[0024] obtaining the target knowledge sub-library based on the iterated target threshold value.

[0025] Optionally, the obtaining the first clustering center information based on the first type of cluster comprises:

[0026] The first text vector average value is calculated based on the first type of cluster.

[0027] The first clustering center information is obtained based on the first text vector average value.

[0028] Optionally, before the target key term information is obtained by performing key information extraction on the target knowledge base, the method further comprises:

[0029] The target knowledge base is preprocessed, wherein the preprocessing comprises removing stop words and removing punctuation processing.

[0030] Optionally, the obtaining the target key term information by performing key information extraction on the target knowledge base comprises:

[0031] The process case in the target knowledge base is subjected to length division processing to obtain sentence vector information.

[0032] The process case in the target knowledge base is subjected to candidate word extraction processing to obtain word vector information.

[0033] The target key term information is obtained based on the sentence vector information and the word vector information.

[0034] To solve the above technical problems, the embodiments of the present application further provide a construction device of a knowledge base for aircraft manufacturing, comprising:

[0035] A data acquisition module is configured to acquire initial clustering center information, wherein the initial clustering center information is obtained based on an initial process case set.

[0036] A first data processing module is configured to perform first clustering analysis processing on the initial process case set based on the initial clustering center information, and obtain a first type of cluster.

[0037] A second data processing module is configured to acquire initial inverse clustering center information, perform first inverse clustering analysis processing on the initial process case set based on the initial inverse clustering center information, and obtain a first inverse type of cluster; and delete the first inverse type of cluster from the initial process case set to obtain a first process case set, wherein the initial inverse clustering center information is obtained based on the initial clustering center information.

[0038] a first target obtaining module configured to obtain first clustering center information based on the first type of cluster, and obtain a target knowledge sub-library based on the first clustering center information, wherein the target knowledge sub-library is obtained by performing second clustering analysis on the first process case set based on the first clustering center information;

[0039] a second target obtaining module configured to perform key information extraction on the target knowledge sub-library to obtain target key term information;

[0040] a target generating module configured to obtain a target knowledge library based on the target knowledge sub-library and the target key term information.

[0041] To solve the above technical problems, the embodiments of the present application further provide an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0042] To solve the above technical problems, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the processor executes the computer program to implement the method described above.

[0043] The beneficial effects that can be achieved by the present application are as follows:

[0044] The embodiments of the present application obtain a first type of cluster with high similarity and a first inverse cluster with low similarity to the initial clustering center by performing first clustering analysis and first inverse clustering analysis on the initial process case set, and eliminate the first inverse cluster with low similarity from the initial case set, which is equivalent to preliminary screening. Then, the first clustering center information is determined again based on the first type of cluster, and clustering analysis is performed again, which is equivalent to fine screening. Through preliminary screening and fine screening, the target knowledge sub-library obtained is more accurate. Finally, key information extraction is performed from the target knowledge sub-library, and the target knowledge sub-library and the target key term information are combined to form a target knowledge library. In the traditional method of directly screening relevant process cases from process cases according to key term information, due to the different properties of composite materials used in aircraft manufacturing, the difference is large, and when the key term information is not used properly, many relevant process cases will be lost. However, the present application can retain more process cases related to aircraft manufacturing and provide a more complete knowledge library. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 A flowchart of a method for constructing a knowledge library for aircraft manufacturing according to an embodiment of the present application;

[0046] Figure 2A structural schematic diagram of a device for constructing a knowledge base for aircraft manufacturing according to an embodiment of the present application;

[0047] Figure 3 An electronic device structural schematic diagram of a hardware running environment according to an embodiment of the present application.

[0048] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0049] It should be understood that the specific embodiments described herein merely serve to explain the present application and are not intended to limit the present application.

[0050] It should be noted that if the present application has a description of "first", "second", etc., the description of "first", "second", etc. is only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features with "first", "second" can explicitly or implicitly include at least one of the features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel solutions. For example, "A and / or B" includes A solution, or B solution, or A and B solutions. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the realization of ordinary skilled in the art. When the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor is it within the scope of protection claimed by the present application.

[0051] Due to the large number of components contained in the whole aircraft, the components used in different parts are made of different composite materials through a series of processing techniques. The current knowledge base construction method is usually to obtain relevant process cases from process cases according to keyword information, integrate relevant process cases, and eliminate the rest of the process cases to construct a knowledge base. However, due to the large difference in the properties of different composite materials, the relevance of keyword information between multiple process cases is weak. When keyword information is not used properly, it will cause the loss of many relevant process cases, so that the constructed knowledge base is not perfect enough, affecting the efficiency and quality of aircraft research and development.

[0052] Based on this, the application provides a solution, that is, a construction method of an aircraft manufacturing knowledge base, which obtains a first cluster with high similarity and a first inverse cluster with low similarity to an initial cluster center by performing first clustering analysis processing and first inverse clustering analysis processing on an initial process case set, removes the first inverse cluster with low similarity from the initial case set, which is equivalent to preliminary screening; re-determines first cluster center information based on the first cluster, and performs clustering analysis processing again, which is equivalent to fine screening; through preliminary screening and fine screening, the target knowledge sub-base obtained is more accurate, and finally key information is extracted from the target knowledge sub-base, and the target knowledge sub-base and the target key term information are combined to form a target knowledge base. Through the above steps, more process cases related to aircraft manufacturing can be retained, and a more perfect knowledge base can be provided.

[0053] Based on this, with reference to Figure 1 The embodiment of the application provides a construction method of an aircraft manufacturing knowledge base, which comprises the following steps:

[0054] Step S10, obtaining initial cluster center information; wherein the initial cluster center information is obtained based on an initial process case set.

[0055] It should be noted that before the initial cluster center information is obtained, an initial process case set needs to be constructed. The initial process case set comprises a plurality of process cases and text vectors corresponding to the process cases. The initial cluster center information refers to a process case selected based on the initial process case set, which serves as an initial cluster center of a cluster. Based on the initial cluster center information, the initial process case set can be subjected to first clustering analysis processing subsequently.

[0056] In actual application, two or more initial cluster center information can be obtained based on the properties, preparation processes and formed parts of the composite materials used for aircraft manufacturing, and two or more target knowledge sub-bases can be further obtained based on the initial cluster center information, so as to improve the hierarchical nature of knowledge classification.

[0057] In actual application, a Word2vec algorithm model can be used to perform vectorization processing on the process case texts, so as to obtain the text vectors corresponding to the process cases. For example, the process case text vocabulary is input into the trained Word2vec algorithm model, and the Word2vec algorithm model outputs the corresponding text vectors.

[0058] Step S20, performing first clustering analysis processing on the initial process case set based on the initial cluster center information, and obtaining a first cluster.

[0059] It should be noted that the first clustering analysis processing refers to taking the center case as the initial clustering center information as the center point, calculating the distance or similarity measure of all process cases in the initial process case set except the center case and the center case to measure whether they belong to the same cluster. The first cluster refers to a group of process cases with high similarity to the center case.

[0060] When comparing the similarity between the process case and the center case, the process case can be subjected to clustering analysis processing by calculating the cosine similarity and comparing the cosine similarity with a preset threshold. That is, the above steps include:

[0061] Based on the initial clustering center information, a first cosine similarity value is obtained.

[0062] The first cosine similarity value is compared with a first preset threshold, the initial process case set is subjected to first clustering analysis processing, and a first cluster is obtained.

[0063] It should be noted that when the initial process case set is constructed, all the process cases have been subjected to text processing, that is, the initial process case set already includes the text vectors corresponding to each process case. The first cosine similarity value refers to the cosine similarity value calculated from the text vector of each process case and the text vector of the center case, that is, the first cosine similarity value belongs to a set which includes a plurality of first similarity values. The first preset threshold refers to a specific value for comparing with the first similarity value to determine the similarity between the corresponding process case and the center case; it should be noted that the first preset threshold can be set based on actual application scenarios, which will not be described in detail here.

[0064] In the specific implementation process, each first similarity value obtained by calculation is compared with the first preset threshold, the process cases corresponding to the first similarity values greater than the first preset threshold are clustered into one category to obtain the first cluster. Among them, the greater the first similarity value, the higher the similarity between the corresponding process case and the center case. Conversely, the smaller the first similarity value, the lower the similarity between the corresponding process case and the center case.

[0065] Specifically, the calculation formula of the first similarity value is as follows:

[0066]

[0067] Among them, x k is the text vector of the first cluster case processed by the Word2vec algorithm model; y kThe text vector to be judged whether the clustering can be completed; cos(x, y) is the size of the similarity value between vectors; n is the length of the vector. Among them, if the value of cos(x, y) reaches the slow point threshold, it means that the clustering can be completed, which can be used as related data in the knowledge base.

[0068] Step S30, obtaining initial inverse clustering center information; based on the initial inverse clustering center information, performing first inverse clustering analysis processing on the initial process case set to obtain a first inverse cluster; deleting the first inverse cluster from the initial process case set to obtain a first process case set; wherein the initial inverse clustering center information is obtained based on the initial clustering center information.

[0069] It should be noted that through the above steps, the first similarity value of all process cases in the initial process case set has been calculated and obtained. The initial inverse clustering center information refers to the process case with the smallest first similarity value in the initial process case set (hereinafter referred to as the inverse center case). The inverse center case is taken as the clustering center point of a class cluster, and the distance or similarity measure between the inverse center case and all the remaining process cases in the initial process case set except the inverse center case and the first class cluster is calculated to measure whether they belong to the same class cluster. The first inverse cluster refers to a group of process cases with high similarity to the inverse center case.

[0070] It can be understood that the process cases included in the first inverse cluster have low similarity to the center case, and therefore need to be deleted from the initial case set to obtain the first process case set. That is, the first process case set refers to the process case set after the initial process case set is removed from the first inverse cluster. Through the first inverse clustering analysis processing, process cases that do not belong to the target to be obtained can be quickly removed to construct an effective and accurate target knowledge base.

[0071] As described above, when comparing the similarity between the process case and the inverse center case, the method of calculating the cosine similarity and using the preset threshold for clustering can be used to obtain the similarity, and therefore this step includes:

[0072] Based on the initial inverse clustering center information, a second cosine similarity value is obtained;

[0073] The second cosine similarity value is compared with a second preset threshold, and the initial process case set is subjected to first inverse clustering analysis processing to obtain a first inverse cluster.

[0074] It should be noted that the second cosine similarity value refers to the cosine similarity value calculated by each process case text vector and the inverse center case text vector, that is, the second cosine similarity value belongs to a set, which includes several second similarity values. The second preset threshold refers to a specific value for comparing with the second similarity value to determine the similarity of the corresponding process case and the inverse center case.

[0075] In the specific implementation process, the second similarity value obtained by calculation is compared with the second preset threshold, and the process cases corresponding to the second similarity values greater than the second preset threshold are clustered into one class to obtain a first inverse cluster. Wherein, the greater the second similarity value, the higher the similarity between the corresponding process case and the inverse center case. Conversely, the smaller the second similarity value, the lower the similarity between the corresponding process case and the inverse center case.

[0076] In order to avoid the loss of too much knowledge data caused by the first inverse clustering analysis processing, the present application adopts the iterative way for inverse clustering analysis processing.

[0077] Specifically, before the first clustering center information is obtained based on the first cluster, and the target knowledge sub-library is obtained based on the first clustering center information, it further includes:

[0078] Based on the second preset threshold, an initial target threshold is obtained; wherein the initial target threshold is obtained based on the second preset threshold and a preset threshold growth step;

[0079] It is judged whether the initial target threshold exceeds a preset range threshold;

[0080] If yes, the first process case set is subjected to second clustering analysis processing based on the first clustering center information to obtain a target knowledge sub-library;

[0081] If not, the first process case set is subjected to second clustering analysis processing and second inverse clustering analysis processing based on the first clustering center information to obtain an iterated target threshold; the iterated target threshold is greater than the preset range threshold;

[0082] Based on the iterated target threshold, a target knowledge sub-library is obtained.

[0083] It should be noted that the preset range threshold is set according to actual conditions, which is not limited here. The target threshold refers to the sum of the second preset threshold and the preset threshold growth step.

[0084] In a specific implementation, the second preset threshold used in the first inverse clustering analysis process can take the minimum value of the preset range threshold. After completing the first inverse clustering analysis process, the target threshold is first calculated based on the second preset threshold and the preset threshold growth step, and it is determined whether the target threshold exceeds the preset range threshold. If so, it indicates that the first process case set obtained does not need to be further case eliminated, and the subsequent steps can be performed. If not, it indicates that the first process case set obtained still needs to continue case elimination, and iteration begins at this time. That is, with the first cluster center information obtained by the first cluster calculation as the center point, the second clustering analysis process and the second inverse clustering analysis process are performed until the target threshold exceeds the preset range threshold, and then the subsequent steps are performed.

[0085] The second cluster analysis process and the second inverse cluster analysis process are performed in the same manner as the first cluster analysis process and the first inverse cluster analysis process described above. It should be noted that the terms "first" and "second" are used only to indicate the order before and after iterations, and are not intended to limit the first or second iterations. The specific number of iterations can be adjusted based on actual conditions. That is, before the target threshold exceeds the preset range threshold, steps S20 and S30 will be executed in a loop until the target threshold exceeds the preset range threshold.

[0086] However, it should be noted that when step S20 is repeated, the initial cluster center information is not used as the cluster center. Instead, as described in step S40, the cluster center is re-determined based on the most recently obtained first cluster. Similarly, when step S30 is repeated, the initial inverse cluster center information is obtained based on the re-determined cluster center. During the iteration process, the second preset threshold is determined by a gradient increase within a preset range (i.e., the preset threshold is incremented by a predetermined step size each time).

[0087] For example, assuming the preset threshold range is [0.6, 0.8] and the preset threshold increment is 0.05, in actual application, the second preset threshold may start at 0.6 to perform the first inverse clustering analysis. In subsequent iterations, the second preset threshold may be incremented by 0.05 each time.

[0088] Step S40: obtaining first cluster center information based on the first cluster; obtaining a target knowledge sub-base based on the first cluster center information; wherein the target knowledge sub-base is obtained by performing a second cluster analysis on the first process case set based on the first cluster center information.

[0089] It should be noted that the first clustering center information refers to the process case determined as a clustering center based on the first cluster. The target knowledge sub-library refers to the knowledge sub-library in the target knowledge base to be obtained. In the specific implementation process, specifically, according to the first cluster obtained newly, the average value of all process case text vectors contained in the first cluster is calculated, and the average value obtained by the calculation is taken as the re-determined clustering center. The second clustering analysis processing is performed with the re-determined first clustering center information as the clustering center, the second cluster is obtained, and the second cluster is taken as the target knowledge sub-library, so as to improve the knowledge recommendation accuracy of the target knowledge sub-library.

[0090] Specifically, the first clustering center information is obtained based on the first cluster, including:

[0091] The first text vector average value is calculated based on the first cluster;

[0092] The first clustering center information is obtained based on the first text vector average value.

[0093] It should be noted that the first text vector average value refers to the average value of all process case text vectors contained in the first cluster.

[0094] Specifically, the first clustering center information satisfies the following relationship:

[0095]

[0096] Wherein, C represents the re-determined clustering center; S[m] represents the mth process case text vector in the cluster, m=1, 2, …, M, M represents the number of all text vectors in the cluster.

[0097] In addition, it should be noted that when two or more initial clustering center information is obtained, that is, the above all steps of the application are executed multiple times, for each initial clustering center information, a corresponding knowledge sub-library is constructed. In the process of constructing each knowledge sub-library, the values of the first preset threshold, the second preset threshold and the preset range threshold can be the same or different, which is specifically set according to the actual situation.

[0098] Step S50, extracting key information from the target knowledge sub-library to obtain target key word information.

[0099] It should be noted that the key information includes material name, process, parts and components, etc. The target key word information refers to the word information with high relevance to the target knowledge base obtained by the target.

[0100] Since the data content of the process case in the target knowledge base is messy and contains a large number of meaningless words, in order to improve the accuracy and efficiency of key information extraction, specifically, before the key information extraction of the target knowledge base is performed and the target key word information is obtained, the method further comprises the following steps:

[0101] Preprocessing the target knowledge base; wherein the preprocessing comprises removing stop words and removing punctuation processing.

[0102] It should be noted that the stop word removal processing refers to a method for removing stop words in the process case text of the target knowledge base, which can improve the accuracy of subsequent extraction of process knowledge. The stop word list can be provided by the commonly used stop word list in the art, or the stop word list can be generated according to the words with higher probability but no content meaning in the text. The punctuation removal processing refers to a method for processing the word segmentation of the process case text of the target knowledge base, so as to facilitate the subsequent extraction of key words.

[0103] Specifically, the key information extraction of the target knowledge base to obtain target key word information comprises the following steps:

[0104] Length division processing is performed on the process case in the target knowledge base to obtain sentence vector information;

[0105] The candidate word extraction processing is performed on the process case in the target knowledge base to obtain word vector information.

[0106] Based on the sentence vector information and the word vector information, the target key word information is obtained.

[0107] It should be noted that the sentence vector information includes a plurality of sentences. The word vector information includes a plurality of candidate words. By calculating the similarity of the sentences and the candidate words, the target key word information is determined.

[0108] In order to facilitate understanding, the above steps are described in detail taking the SBERT algorithm as an example. After preprocessing the process case of the target knowledge base, the steps of obtaining the target key word information are as follows:

[0109] According to the length of the process case text, the text is divided into a plurality of sentences; specifically, the preprocessed process case text is divided, and in this paper, the length of the text is divided, the text A is composed of a plurality of sentences a, and the same number of words is set, then the segmentation formula can be obtained:

[0110]

[0111] Wherein, A represents the number of words in the text part, a represents the number of divided sentences, and z represents the number of words contained in a sentence.

[0112] A number of candidate words are extracted from the process case text; specifically, all the names appearing in the process case text are taken as candidate words. The similarity between the sentence and the candidate words is calculated based on the SBERT algorithm model. First, the divided sentence and the candidate words are input into the two trained BERT models to obtain the sentence vector and the candidate word vector with the same dimension; then, the cosine similarity between the sentence vector and the candidate word vector is calculated according to the cosine similarity calculation formula described above, when the cosine similarity is greater than the third preset threshold, it means that the semantic of the sentence vector and the candidate word vector are close; the candidate word corresponding to the word vector can be taken as the target keyword.

[0113] For each divided sentence, the similarity between it and each candidate word extracted is calculated to obtain the target keyword.

[0114] The above steps are repeated to determine the target keywords of each process case, and all the target keywords are combined to form the target keyword information.

[0115] For example, the D1 sentence vector is: There is a park near my home, There are a lot of beautiful trees, flowers and birds in the park;

[0116] The D2 word vector is: banana;

[0117] The D3 word vector is: apple;

[0118] The D4 word vector is: park;

[0119] The D5 word vector is: tree;

[0120] The cosine similarity between D1 and D2 is calculated by the cosine similarity calculation formula to obtain:

[0121] The cosine similarity between D1 and D3 is: 0.6228;

[0122] The cosine similarity between D1 and D4 is: 0.7990;

[0123] The cosine similarity between D1 and D5 is: 0.7187.

[0124] The cosine similarity between D1 and D5 is: 0.7187.

[0125] Suppose the third preset threshold is 0.7, it can be seen that the cosine similarity between D4 and D5 is greater than the third preset threshold, which means that D4 and D5 are close in semantic to D1. Therefore, D4 and D5 are taken as the target keywords.

[0126] Step S60, obtaining a target knowledge base based on the target knowledge sub-base and the target keyword information.

[0127] It should be noted that the target knowledge base is constructed by merging the target knowledge sub-base and the target keyword information.

[0128] Referring to Figure 2 Based on the same inventive concept, the embodiment of the present application also proposes a construction device of a knowledge base for aircraft manufacturing, comprising:

[0129] The data acquisition module is configured to acquire initial clustering center information, wherein the initial clustering center information is obtained based on an initial process case set.

[0130] The first data processing module is configured to perform first clustering analysis processing on the initial process case set based on the initial clustering center information, and obtain a first cluster.

[0131] The second data processing module is configured to acquire initial inverse clustering center information, perform first inverse clustering analysis processing on the initial process case set based on the initial inverse clustering center information, and obtain a first inverse cluster. The first inverse cluster is deleted from the initial process case set to obtain a first process case set. The initial inverse clustering center information is obtained based on the initial clustering center information.

[0132] The first target obtaining module is configured to obtain first clustering center information based on the first cluster, and obtain a target knowledge sub-base based on the first clustering center information. The target knowledge sub-base is obtained by performing second clustering analysis processing on the first process case set based on the first clustering center information.

[0133] The second target obtaining module is configured to extract key information from the target knowledge sub-base to obtain target keyword information.

[0134] The target generation module is configured to obtain a target knowledge base based on the target knowledge sub-base and the target keyword information.

[0135] It should be noted that the modules in the construction device of the knowledge base for aircraft manufacturing in the embodiment correspond one by one to the steps in the construction method of the knowledge base for aircraft manufacturing in the foregoing embodiment. Therefore, the specific embodiments of the present embodiment can refer to the embodiments of the construction method of the knowledge base for aircraft manufacturing, which will not be described here.

[0136] Referring to Figure 3 , Figure 3 The electronic device structure schematic diagram of the hardware running environment involved in the embodiment of the present application.

[0137] AsFigure 3 As shown in the figure, the electronic device can include a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to realize the connection and communication between the components. The user interface 1003 can include a display, an input unit such as a keyboard, and can also include a standard wired interface, a wireless interface. The network interface 1004 can optionally include a standard wired interface, a wireless interface (such as a wireless fidelity (WIreless-FIdelity, WI-FI) interface). The memory 1005 can be a high-speed random access memory (RAM) memory, or a stable non-volatile memory (Non-Volatile Memory, NVM) such as a disk memory. The memory 1005 can also be a storage device independent of the aforementioned processor 1001.

[0138] Those skilled in the art can understand that the structure shown in the figure does not constitute a limitation on the electronic device, and can include more or fewer components than the figure, or combine certain components, or different component arrangements. Figure 3

[0139] As shown in the figure, the memory 1005 as a storage medium can include an operating system, a network communication module, a user interface module, and an electronic program. Figure 3

[0140] In the electronic device shown in the figure, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the electronic device of the present application can be arranged in the electronic device, and the electronic device calls the construction device of the knowledge base for aircraft manufacturing stored in the memory 1005 through the processor 1001, and executes the construction method of the knowledge base for aircraft manufacturing provided by the embodiment of the present application. Figure 3

[0141] In addition, in an embodiment, the embodiment of the present application also provides a computer program product, which, when executed by a processor, implements the method described above.

[0142] In addition, in an embodiment, the embodiment of the present application also provides a computer storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the method in the foregoing embodiment.

[0143] ​​​In some embodiments, the computer readable storage medium can be a memory such as a FRAM, ROM, PROM, EPROM, EEPROM, flash memory, a magnetic surface memory, an optical disk, or a CD-ROM, etc.; or can be various devices including one or any combination of the above memories. The computer can be various computing devices including a smart terminal and a server.

[0144] In some embodiments, the executable instructions can be in the form of a program, software, software modules, scripts, or code, written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0145] By way of example, the executable instructions can, but need not, reside in a file system's files, can be stored in a part of a file that holds other programs or data, for example, in one or more scripts stored in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files, for example, files that store one or more modules, sub programs, or portions of code.

[0146] By way of example, the executable instructions can be deployed to be executed on one computer, or on multiple computers of a location that are interconnected through a communication network.

[0147] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article, or system that includes a list of elements not only includes those elements, but also includes other elements not expressly listed or inherent to such process, method, article, or system. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or system including the element.

[0148] The above-mentioned sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0149] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, can also be through hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application essentially or say the part of the prior art contribution can be embodied in the form of software products, the computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk), including a plurality of instructions to make a multimedia terminal device (may be a mobile phone, computer, television receiver, or network equipment, etc.) executes the method described in various embodiments of the present application.

[0150] The above is only the preferred embodiment of the present application, not therefore limit the patent scope of the present application, any use of the present application specification and drawings content of the equivalent structure or equivalent process transformation, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method of constructing a knowledge base for aircraft manufacturing, characterized by, The method comprises the following steps: obtaining initial clustering center information; wherein the initial clustering center information is obtained based on an initial process case set; performing first clustering analysis processing on the initial process case set based on the initial clustering center information to obtain a first cluster; obtaining initial reverse clustering center information; performing first reverse clustering analysis processing on the initial process case set based on the initial reverse clustering center information to obtain a first reverse cluster; and deleting the first reverse cluster from the initial process case set to obtain a first process case set; wherein the initial reverse clustering center information is obtained based on the initial clustering center information; obtaining first clustering center information based on the first cluster; and obtaining a target knowledge sub-library based on the first clustering center information; wherein the target knowledge sub-library is obtained by performing second clustering analysis processing on the first process case set based on the first clustering center information; extracting key information from the target knowledge sub-library to obtain target key term information; obtaining a target knowledge library based on the target knowledge sub-library and the target key term information; the first clustering analysis processing on the initial process case set based on the initial clustering center information to obtain a first cluster comprises: obtaining a first cosine similarity value based on the initial clustering center information; comparing the first cosine similarity value with a first preset threshold value to perform first clustering analysis processing on the initial process case set, and clustering process cases corresponding to the first cosine similarity value greater than the first preset threshold value into a cluster to obtain a first cluster; the first reverse clustering analysis processing on the initial process case set based on the initial reverse clustering center information to obtain a first reverse cluster comprises: obtaining a second cosine similarity value based on the initial reverse clustering center information; comparing the second cosine similarity value with a second preset threshold value to perform first reverse clustering analysis processing on the initial process case set, and clustering process cases corresponding to the second cosine similarity value greater than the second preset threshold value into a cluster to obtain a first reverse cluster.

2. The method of constructing a knowledge base for aircraft manufacturing according to Claim 1, characterized by, the first clustering center information obtained based on the first cluster comprises: before the target knowledge sub-library obtained based on the first clustering center information, further comprising: obtaining an initial target threshold value based on the second preset threshold value; wherein the initial target threshold value is obtained based on the second preset threshold value and a preset threshold value growth step; determining whether the initial target threshold value exceeds a preset range threshold value; if yes, performing second clustering analysis processing on the first process case set based on the first clustering center information to obtain a target knowledge sub-library; if no, performing second clustering analysis processing and second reverse clustering analysis processing on the first process case set based on the first clustering center information to obtain an iterated target threshold value; the iterated target threshold value is greater than the preset range threshold value; obtaining a target knowledge sub-library based on the iterated target threshold value.

3. The method of constructing a knowledge base for aircraft manufacturing according to Claim 1, wherein, the first clustering center information obtained based on the first cluster comprises: Based on the first type of cluster, a first text vector average is calculated and obtained; Based on the first text vector average, first clustering center information is obtained.

4. The method of constructing a knowledge base for aircraft manufacturing according to Claim 1, wherein, Before the target knowledge sub-library is subjected to key information extraction to obtain target key term information, the method further includes: The target knowledge sub-library is preprocessed, and the preprocessing includes removing stop words and removing punctuation processing.

5. The method of constructing a knowledge base for aircraft manufacturing according to Claim 1, wherein, The target knowledge sub-library is subjected to key information extraction to obtain target key term information, including: The process cases in the target knowledge sub-library are subjected to length division processing to obtain sentence vector information; The process cases in the target knowledge sub-library are subjected to candidate word extraction processing to obtain word vector information; Based on the sentence vector information and the word vector information, target key term information is obtained.

6. An apparatus for constructing a knowledge base for aircraft manufacturing, for implementing the method for constructing a knowledge base for aircraft manufacturing according to any one of claims 1 to 5, characterized in that, The method includes: A data acquisition module is configured to acquire initial clustering center information, wherein the initial clustering center information is obtained based on an initial process case set; A first data processing module is configured to perform first clustering analysis processing on the initial process case set based on the initial clustering center information to obtain a first type of cluster; A second data processing module is configured to acquire initial inverse clustering center information, perform first inverse clustering analysis processing on the initial process case set based on the initial inverse clustering center information to obtain a first inverse type of cluster, and delete the first inverse type of cluster from the initial process case set to obtain a first process case set, wherein the initial inverse clustering center information is obtained based on the initial clustering center information; A first target acquisition module is configured to obtain first clustering center information based on the first type of cluster, and obtain a target knowledge sub-library based on the first clustering center information, wherein the target knowledge sub-library is obtained by performing second clustering analysis processing on the first process case set based on the first clustering center information; A second target acquisition module is configured to extract key information from the target knowledge sub-library to obtain target key term information; A target generation module is configured to obtain a target knowledge library based on the target knowledge sub-library and the target key term information.

7. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the method of any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the processor executes the computer program to implement the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Context enhanced Dirichlet model supporting online clustering of short text streams

    CN115827861A

  • Feature clustering dimension reduction method for commercial text classification

    CN116010603A