A method, system, medium, and apparatus for automatically recovering girth weld missing data items
By constructing a sample vector matrix and a failure data vector, and using a preset correlation function to calculate missing data items, the problem of unrecoverable missing data items in circumferential welds was solved, and risk prediction of circumferential welds was realized.
Patent Information
- Application Number
- CN202310217643.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-03
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2043-03-03
AI Technical Summary
Existing technologies cannot effectively recover missing data items of girth welds, resulting in the inability to perform risk prediction.
Construct a sample vector matrix and a missing data vector, calculate missing data items using a preset correlation function, and recover missing data using fitting coefficients.
Quickly recover missing data items to enable normal risk prediction for circumferential welds.
Smart Images

Figure CN116205362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of girth weld failure risk data processing, and in particular to a method, system, medium and equipment for automatically recovering missing data items of a girth weld. Background Art
[0002] Girth weld failure is the most significant factor affecting pipeline failure. Girth weld failure prediction is achieved by using various machine learning algorithms to learn from representative training samples. While the selection of representative training samples ensures data completeness, in actual predictions, the basic data for girth welds in production often suffers from missing data items. Missing data items can render the trained model incapable of prediction. Currently, there is no method to recover missing data items for girth welds based on representative training samples, making risk prediction impossible for girth welds with missing data items. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to address the deficiencies of the prior art and provide a method, system, medium and equipment for automatically recovering missing data items of a girth weld.
[0004] The technical solution of the present invention to solve the above technical problems is as follows:
[0005] A method for automatically recovering missing data items of a girth weld, comprising:
[0006] Construct a sample vector matrix of sample data including complete data items;
[0007] Constructing a first failure data vector based on the missing data of the girth weld including the missing data item; calculating a preset correlation function based on the first failure data vector in combination with the sample vector matrix; and calculating the missing data item based on the preset correlation function;
[0008] Recover missing data for girth welds based on the calculated missing data terms.
[0009] The beneficial effect of the present invention is that this solution can quickly recover missing data items by fitting coefficients of typical samples of girth welds for missing data items, using the fitting coefficients and typical samples, making it possible to achieve normal prediction of girth welds that were previously impossible to risk predict.
[0010] Furthermore, the constructing of a sample vector matrix of sample data including complete data items specifically includes:
[0011] Construct multiple sample vectors based on the sample data of the complete data item;
[0012] Perform a preset normalization process on each sample vector, and construct a sample vector matrix based on all normalized sample vectors.
[0013] Furthermore, the preset normalization process specifically includes:
[0014] Normalization is performed using the first formula, which is:
[0015]
[0016] in, is the normalized value of the jth data in the i-th sample vector of the sample data; x ij is the actual value of the jth data in the i-th sample vector of the sample data; max j is the actual maximum value of the jth data in all sample vectors.
[0017] Furthermore, the constructing of the first failure data vector including the missing data items specifically includes:
[0018] Construct a failure data vector including missing data items;
[0019] The preset normalization process is performed on the failure data vector to obtain a first failure data vector after the normalization process.
[0020] Another technical solution of the present invention to solve the above technical problems is as follows:
[0021] A system for automatically recovering missing data items of girth welds, comprising: a full data vector construction module, a missing data vector construction module, a correlation function calculation module, a missing item calculation module, and a recovery module;
[0022] The full data vector construction module is used to construct a first failure data vector based on the girth weld missing data including the missing data item;
[0023] The missing data vector construction module is used to construct a first failure data vector including missing data items;
[0024] The correlation function calculation module is used to calculate a preset correlation function based on the first failure data vector and the sample vector matrix;
[0025] The missing item calculation module is used to calculate the missing data items according to the preset correlation function; the recovery module is used to restore the missing data of the girth weld according to the calculated missing data items.
[0026] The beneficial effect of the present invention is that this solution can quickly recover missing data items by fitting coefficients of typical samples of girth welds for missing data items, using the fitting coefficients and typical samples, making it possible to achieve normal prediction of girth welds that were previously impossible to risk predict.
[0027] Furthermore, the full data vector construction module is specifically used to construct a plurality of sample vectors based on the sample data of the complete data item;
[0028] Perform a preset normalization process on each sample vector, and construct a sample vector matrix based on all normalized sample vectors.
[0029] Furthermore, the full data vector construction module is specifically configured to perform normalization processing using a first formula, where the first formula is:
[0030]
[0031] in, is the normalized value of the jth data in the i-th sample vector of the sample data; x ij is the actual value of the jth data in the i-th sample vector of the sample data; max j is the actual maximum value of the jth data in all sample vectors.
[0032] Furthermore, the missing data vector construction module is specifically used to construct a failure data vector including missing data items;
[0033] The preset normalization process is performed on the failure data vector to obtain a first failure data vector after the normalization process.
[0034] Another technical solution of the present invention to solve the above technical problems is as follows:
[0035] A storage medium, characterized in that instructions are stored in the storage medium, and when a computer reads the instructions, the computer is caused to execute a method for automatically recovering missing data items of a girth weld as described in any of the above schemes.
[0036] Another technical solution of the present invention to solve the above technical problems is as follows:
[0037] An electronic device, characterized in that it includes a processor and the storage medium described in the above solution, and the processor executes instructions in the storage medium.
[0038] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic flow chart of a method for automatically recovering missing data items of a girth weld provided by an embodiment of the present invention;
[0040] Figure 2A structural block diagram of a system for automatically recovering missing data items of girth welds provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0041] The principles and features of the present invention are described below with reference to the accompanying drawings. The embodiments given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0042] like Figure 1 As shown, a method for automatically recovering missing data items of a girth weld provided by an embodiment of the present invention includes:
[0043] S1, construct a sample vector matrix of sample data including complete data items;
[0044] It should be noted that the sample vector matrix construction process may include:
[0045] Define V i is the sample vector consisting of the data contained in the i-th sample. Each data in the sample vector is normalized according to formula (1).
[0046]
[0047] in, is the normalized value of the jth data in the i-th sample vector. The jth data item can be any data item, namely, any one of diameter, toughness, defect type, defect clockwise direction, defect length, defect height, and pressure. ij is the actual value of the jth data in the i-th sample vector. maxj is the actual maximum value of the j-th data in all sample vectors.
[0048] Among them, x ij It's V i Specific items, V i =[x i1 , x i2 , x i3 ,...x ij ...]. It should be noted that the sample data contains all data items, with no missing items. Use the complete sample data to complete a data item with missing items.
[0049] The sample vector matrix is constructed according to the normalized sample vector, as shown in formula (2).
[0050] V=「V1,V2,...,V N ] (2)
[0051] S2, constructing a first failure data vector based on the missing data of the girth weld including the missing data item; it should be noted that, in a certain embodiment, constructing the first failure data vector may include:
[0052] The failure data vector with missing data items is represented by V x Represent it, normalize it according to formula (1), and set the missing data items to 0. Use p to represent V x The missing data item number. Obviously, v xp =0.
[0053] Construct the function shown in formula (3).
[0054]
[0055] Among them, v i,j V i The jth data item of h i is the i-th data item of h. xi V x The i-th data item in . N is the number of sample vectors. M is the dimension of a single sample vector. H is the coefficient vector for fitting the missing data item vector to the sample vector.
[0056] It should be noted that formula (3) is the association method between the failure data vector with missing data items and the sample vector matrix, |f(h) is a function of the coefficient h, and the sample vector can best match the failure data vector with missing data items through h.
[0057] S3, calculating a preset correlation function based on the first failure data vector and the sample vector matrix; wherein the preset correlation function may be a coefficient vector for fitting the missing data item vector to the sample vector, and minimizing the fitting error through S3.
[0058] It should be noted that calculating the preset correlation function may include: using formula (4) to solve h1~h N , that is, h vector, h=[h1,h2,...h i ,...h N ].
[0059]
[0060] Among them, pinv() is the function for finding the pseudo-inverse matrix. i ~h N is the coefficient vector, V x1 ~V xM Is a data vector with missing data items, the missing data items are vxp. V11~v1M are the data items of the first sample vector, V 21 ~V 2M For each data item of the second sample vector, V N1 ~V NM is the data item of the Nth sample vector.
[0061] S4, calculating the missing data items according to the preset correlation function; it should be noted that, according to the preset correlation function, that is, h=[h1, h2, ...h i ,...h N ]The coefficient vector of the sample vector fitting the missing data term vector is calculated according to formula (5) xp .
[0062]
[0063] Where p represents V x The missing data item number. xp is the missing term in Vx to be solved, which is explained in S2. N is the number of sample vectors, which has the same meaning as in formula (2). hi is the value of each vector h in the solved vector, i.e. h = [h1, h2, ..h i ,..h N ]. ip It is the pth data item of the i-th sample vector. The sample vector is provided by the site and all data items are not missing.
[0064] S5, recovering the missing data of the girth weld according to the calculated missing data items.
[0065] It should be noted that V x The data vector that represents the missing data item needs to be completed using the sample vector. The missing data item number is represented by p. Therefore, V xp Refers to the specific missing data items that need to be solved. v1, v2, v i , v N All refer to sample data vectors, 1, 2, i, N are numbers. ij It refers to the jth data item of the i-th sample vector. The sample values are all complete data provided on site. The data vectors with missing data items are also provided on site, but only the missing items are missing. Automatic completion of missing items is the purpose of this invention.
[0066] This solution uses fitting coefficients for typical samples of girth welds with missing data items. Utilizing the fitting coefficients and typical samples, the missing data items can be quickly restored, enabling normal prediction of girth welds that were previously impossible to predict.
[0067] Optionally, in some embodiments, constructing a sample vector matrix of sample data including complete data items specifically includes:
[0068] Construct multiple sample vectors based on the sample data of the complete data item;
[0069] Perform a preset normalization process on each sample vector, and construct a sample vector matrix based on all normalized sample vectors.
[0070] Optionally, in some embodiments, the preset normalization process specifically includes:
[0071] Normalization is performed using the first formula, which is:
[0072]
[0073] in, is the normalized value of the jth data in the i-th sample vector of the sample data; x ij is the actual value of the jth data in the i-th sample vector of the sample data; max j is the actual maximum value of the jth data in all sample vectors.
[0074] Optionally, in some embodiments, constructing the first failure data vector including missing data items specifically includes:
[0075] Construct a failure data vector including missing data items;
[0076] The preset normalization process is performed on the failure data vector to obtain a first failure data vector after the normalization process.
[0077] In one embodiment, typical samples and data for girth weld failure prediction are shown in Table 1. Sample V with missing data items x As shown in Table 2, the defect type is missing.
[0078]
[0079] Table 1
[0080]
[0081] Table 2
[0082] S11: Define V i is the sample vector consisting of the data contained in the i-th sample. The data in the sample vector in Table 1 are normalized according to formula (1), and Table 2 is obtained.
[0083]
[0084] in, is the normalized value of the jth data in the i-th sample vector. ij is the actual value of the jth data in the i-th sample vector. max jis the actual maximum value of the jth data in all sample vectors. The normalized typical sample data table is shown in Table 3:
[0085]
[0086] Table 3
[0087] S12: Construct a sample vector matrix based on the normalized sample vector, as shown in formula (2).
[0088]
[0089] S13: The failure data vector with missing data items is V x Represent it, normalize it according to formula (1), and set the missing data items to 0. Use p to represent V x The missing data item number. Obviously, V xp =0, where p=3.
[0090] V x =[O.750123052, 0.9, 0.0, 0.380388842, 0.065454545, 0.13740458, 0.75],
[0091] S14: Construct a function as shown in formula (3).
[0092]
[0093] S15: Use equation (4) to solve h1~h 10 .
[0094]
[0095] Among them, pinv() is a function for finding the pseudo-inverse matrix.
[0096] S16: Calculate the missing item v according to formula (5) xp , where p=3.
[0097]
[0098] Among them, V i,3 =[0.9, 0.3, 0.9, 0.3, 0.9, 0.3, 0.6, 0.9, 0.6, 0.6], i=1,...,10.
[0099] h=[-0.01363, 0.9086, -0.00624, -0.00315, 0.075487, 0.168597, -0.05056, 0.056874, 0.0647, -0.196754286238276].
[0100] The actual number of missing data items is 0.3, and the auto-completion result is 0.31388, with a relative error of only 4.6%.
[0101] In one embodiment, if Figure 2 As shown, a system for automatically recovering missing data items of girth welds includes: a full data vector construction module 1101, a missing data vector construction module 1102, a correlation function calculation module 1103, a missing item calculation module 1104 and a recovery module 1105;
[0102] The full data vector construction module 1101 is used to construct a sample vector matrix of sample data including complete data items;
[0103] The missing data vector construction module 1102 is used to construct a first failure data vector based on the girth weld missing data including the missing data item;
[0104] The correlation function calculation module 1103 is used to calculate a preset correlation function based on the first failure data vector and the sample vector matrix;
[0105] The missing item calculation module 1104 is used to calculate the missing data items according to the preset correlation function;
[0106] The recovery module 1105 is used to recover the missing data of the girth weld according to the calculated missing data items.
[0107] This solution uses fitting coefficients for typical samples of girth welds with missing data items. Utilizing the fitting coefficients and typical samples, the missing data items can be quickly restored, enabling normal prediction of girth welds that were previously impossible to predict.
[0108] Optionally, in some embodiments, the full data vector construction module 1101 is specifically configured to construct a plurality of sample vectors based on sample data of a complete data item;
[0109] Perform a preset normalization process on each sample vector, and construct a sample vector matrix based on all normalized sample vectors.
[0110] Optionally, in some embodiments, the full data vector construction module 1101 is specifically configured to perform normalization processing using a first formula, where the first formula is:
[0111]
[0112] in, is the normalized value of the jth data in the i-th sample vector of the sample data; x ij is the actual value of the jth data in the i-th sample vector of the sample data; maxj is the actual maximum value of the jth data in all sample vectors.
[0113] Optionally, in some embodiments, the missing data vector construction module 1102 is specifically configured to construct a failure data vector including missing data items;
[0114] The preset normalization process is performed on the failure data vector to obtain a first failure data vector after the normalization process.
[0115] Another technical solution of the present invention to solve the above technical problems is as follows:
[0116] It can be understood that in some embodiments, some or all of the optional implementation methods in the above embodiments may be included.
[0117] It should be noted that the above embodiments are product embodiments corresponding to the previous method embodiments. For the description of the optional implementation methods in the product embodiments, please refer to the corresponding description in the above method embodiments, which will not be repeated here.
[0118] In a certain embodiment, a storage medium stores instructions, and when a computer reads the instructions, the computer is caused to execute a method for automatically recovering missing data items of a girth weld as described in any of the above embodiments.
[0119] In one embodiment, an electronic device includes a processor and the storage medium described in the above embodiment, wherein the processor executes instructions in the storage medium.
[0120] The reader should understand that, in the description of this specification, reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0121] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the method embodiments described above are merely illustrative. For example, the division of steps is merely a logical function division. In actual implementation, other division methods may be used. For example, multiple steps may be combined or integrated into another step, or some features may be ignored or not performed.
[0122] If the above method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0123] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for automatically recovering missing data items of girth welds, characterized in that: include: Construct a sample vector matrix of sample data including complete data items; constructing a first failure data vector based on the missing data of the girth weld including the missing data item; The first failure data vector is: Among them, v i,j V i The jth data item, V i is the sample vector consisting of the data contained in the i-th sample, h i is the i-th data item of h, v xi V x The i-th data item, V x is the failure data vector with missing data items, N is the number of sample vectors, M is the dimension of a single sample vector, and h is the coefficient vector of the sample vector fitting the missing data item vector; the above formula is the association method between the failure data vector with missing data items and the sample vector matrix; Calculate the preset correlation function based on the first failure data vector combined with the sample vector matrix; Calculating the preset correlation function includes: solving h1~h using the following formula N , that is, h vector, h=[h1,h2,…h i ,…h N ]: Among them, pinv() is the pseudo-inverse matrix function, h i ~h N is the coefficient vector, v x1 ~v xM is a data vector with missing data items, where the missing data items are v xp , v 1,1 ~v 1,M For each data item of the first sample vector, v 2,1 ~v 2,M For each data item of the second sample vector, v N,1 ~v N,M is the data item of the Nth sample vector; The missing data items are calculated according to the preset correlation function; wherein, according to the preset correlation function, that is, h=[h1,h2,…h i ,…h N ]The coefficient vector of the sample vector fitting the missing data term vector, press to calculate the missing data term v xp : Where p represents V x The missing data item number, v xp It is V x The missing data item to be found in v i,p is the pth data item of the i-th sample vector; Recover missing data for girth welds based on the calculated missing data terms.
2. The method for automatically recovering missing data items of a girth weld according to claim 1, characterized in that: The constructing of the sample vector matrix of the sample data including the complete data items specifically includes: Construct multiple sample vectors based on the sample data of the complete data item; Perform a preset normalization process on each sample vector, and construct a sample vector matrix based on all normalized sample vectors.
3. The method for automatically recovering missing data items of a girth weld according to claim 2, characterized in that: The preset normalization process specifically includes: Normalization is performed using the first formula, which is: in, is the normalized value of the jth data in the i-th sample vector of the sample data; x ij is the actual value of the jth data in the i-th sample vector of the sample data; max j is the actual maximum value of the jth data in all sample vectors.
4. A method for automatically recovering missing data items of girth welds according to claim 2 or 3, characterized in that: The constructing of the first failure data vector including the missing data items specifically includes: Construct a failure data vector including missing data items; The preset normalization process is performed on the failure data vector to obtain a first failure data vector after the normalization process.
5. A system for automatically recovering missing data items of girth welds, characterized in that: include: Full data vector construction module, missing data vector construction module, correlation function calculation module, missing item calculation module and recovery module; The full data vector construction module is used to construct a sample vector matrix of sample data including complete data items; The missing data vector construction module is used to construct a first failure data vector according to the missing data of the girth weld including the missing data item; the first failure data vector is: Among them, v i,j V i The jth data item, V i is the sample vector consisting of the data contained in the i-th sample, h i is the i-th data item of h, v xi V x The i-th data item, V x is the failure data vector with missing data items, N is the number of sample vectors, M is the dimension of a single sample vector, and h is the coefficient vector of the sample vector fitting the missing data item vector; the above formula is the association method between the failure data vector with missing data items and the sample vector matrix; The correlation function calculation module is used to calculate the preset correlation function based on the first failure data vector combined with the sample vector matrix; calculating the preset correlation function includes: solving h1~h using the following formula N , that is, h vector, h=[h1,h2,…h i ,…h N ]: Among them, pinv() is the pseudo-inverse matrix function, h i ~h N is the coefficient vector, v x1 ~v xM is a data vector with missing data items, where the missing data items are v xp , v 1,1 ~v 1,M For each data item of the first sample vector, v 2,1 ~v 2,M For each data item of the second sample vector, v N,1 ~v N,M is the data item of the Nth sample vector; The missing item calculation module is used to calculate the missing data items according to the preset correlation function; wherein, according to the preset correlation function, that is, h=[h1,h2,…h i ,…h N ]The coefficient vector of the sample vector fitting the missing data term vector, press to calculate the missing data term v xp : Where p represents V x The missing data item number, v xp It is V x The missing data item to be found in v i,p is the pth data item of the i-th sample vector; The recovery module is used to recover the missing data of the girth weld according to the calculated missing data items.
6. The system for automatically recovering missing girth weld data items according to claim 5, characterized in that: The full data vector construction module is specifically used to construct multiple sample vectors based on the sample data of the complete data item; Perform a preset normalization process on each sample vector, and construct a sample vector matrix based on all normalized sample vectors.
7. The system for automatically recovering missing girth weld data items according to claim 6, characterized in that: The full data vector construction module is specifically used to perform normalization processing using a first formula, where the first formula is: in, is the normalized value of the jth data in the i-th sample vector of the sample data; x ij is the actual value of the jth data in the i-th sample vector of the sample data; max j is the actual maximum value of the jth data in all sample vectors.
8. A system for automatically recovering missing girth weld data items according to claim 6 or 7, characterized in that: The missing data vector construction module is specifically used to construct a failure data vector including missing data items; The preset normalization process is performed on the failure data vector to obtain a first failure data vector after the normalization process.
9. A storage medium, characterized in that: The storage medium stores instructions, and when a computer reads the instructions, the computer is caused to execute a method for automatically recovering missing data items of a girth weld according to any one of claims 1 to 4.
10. An electronic device, characterized in that: The device comprises a processor and the storage medium according to claim 9, wherein the processor executes instructions in the storage medium.
Citation Information
Patent Citations
Pipeline magnetic flux leakage internal detection missing data interpolation method based on an LS-KNN
CN109492708A
GAN-based reconstruction method for pipeline magnetic flux leakage detection data loss
CN110929376A