Method and device for completing missing data for a hydrometallurgical plant
By using a missing data completion method for hydrometallurgical plants and employing octal vector reconstruction technology, the problem of low data processing efficiency caused by the loss of raw material data in hydrometallurgical plants was solved, achieving efficient data completion and saving computing resources.
Patent Information
- Application Number
- CN202511180600.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-22
AI Technical Summary
In hydrometallurgical plants, raw material data is easily lost due to various reasons. Existing technologies require a large amount of hardware resources to process large amounts of data, resulting in low data processing efficiency.
By acquiring missing triplet data, the target data is determined to fill in the missing data using an octet vector reconstruction method. This includes acquiring missing triplet data, determining the first and second octet vectors, reconstructing them to obtain the third and fourth octet vectors, and determining the target data from the preset candidate data based on these vectors.
It reduces computational overhead, decreases the use of hardware computing resources, improves data processing efficiency, and optimizes data computation paths.
Smart Images

Figure CN120723759B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of big data, and more particularly to a missing data completion method and device for a hydrometallurgical plant. BACKGROUND
[0002] With the continuous change and improvement of metal extraction processes, the use amount and supply chain of raw materials of the hydrometallurgical plant are also changing. In order to reduce the cost of raw material use while promoting process improvement, it is necessary to analyze various use data of raw materials. However, in actual production application, various use data of raw materials are prone to data loss due to various reasons. In the related art, data related to missing data is matched in historical raw material data to determine the missing data.
[0003] In the process of implementing the present application concept, at least the following problems exist in the related art: In the case of a large amount of data, the matching method in the related art needs to occupy a large amount of hardware resources for data query and matching, resulting in low data processing efficiency. SUMMARY
[0004] In view of the above problems, the present application provides a missing data completion method and device for a hydrometallurgical plant to improve data processing efficiency.
[0005] According to a first aspect of the present application, a missing data completion method for a hydrometallurgical plant is provided, comprising: obtaining missing triple data, wherein the missing triple data includes a raw material name, a raw material attribute related to the raw material name, and relationship data representing the association between the raw material name and the raw material attribute, and the raw material name or the raw material attribute in the missing triple data is missing; determining a first eight-dimensional vector and a second eight-dimensional vector from a plurality of eight-dimensional vectors according to non-missing data in the missing triple data; reconstructing the first eight-dimensional vector and the second eight-dimensional vector to obtain a third eight-dimensional vector and a fourth eight-dimensional vector; and determining target data from a plurality of candidate data according to the third eight-dimensional vector and the fourth eight-dimensional vector to complete the missing triple data.
[0006] According to an embodiment of the present application, the first eight-dimensional vector and the second eight-dimensional vector are reconstructed to obtain the third eight-dimensional vector and the fourth eight-dimensional vector, comprising: interchanging a plurality of imaginary part data of the same position and the same number in the first eight-dimensional vector and the second eight-dimensional vector to obtain the third eight-dimensional vector and the fourth eight-dimensional vector.
[0007] According to an embodiment of the present application, the determining the target data from the plurality of candidate data according to the third octuple vector and the fourth octuple vector comprises: evaluating the plurality of candidate data according to the third octuple vector and the fourth octuple vector to obtain a confidence degree of each of the plurality of candidate data; and determining the candidate data with the highest confidence degree as the target data.
[0008] According to an embodiment of the present application, the evaluating the plurality of candidate data according to the third octuple vector and the fourth octuple vector to obtain a confidence degree of each of the plurality of candidate data comprises: determining an interaction vector according to the third octuple vector and the fourth octuple vector; and determining the confidence degree of each of the plurality of candidate data according to a product of the interaction vector and an octuple vector represented by each of the plurality of candidate data.
[0009] According to an embodiment of the present application, the determining the interaction vector according to the third octuple vector and the fourth octuple vector comprises: determining a unit octuple vector according to a ratio of a norm of the fourth octuple vector to the fourth octuple vector; and determining the interaction vector according to a product of the third octuple vector and the unit octuple vector.
[0010] According to an embodiment of the present application, the method further comprises: encoding the initial data to obtain an embedding vector according to a target dimension and a value mapping range, the initial data being a raw material name, a raw material attribute or relationship data; dividing the embedding vector to obtain eight embedding sub-vectors; and obtaining an octuple vector corresponding to the initial data based on the eight embedding sub-vectors.
[0011] According to an embodiment of the present application, the encoding the initial data to obtain an embedding vector according to a target dimension and a value mapping range comprises: randomly initializing a plurality of values same as the target dimension in the value mapping range to obtain the embedding vector.
[0012] According to an embodiment of the present application, the method further comprises: repeatedly performing the following operations until the target dimension is determined from a plurality of embedding vector dimensions: determining an embedding vector dimension of a pth round from the plurality of embedding vector dimensions, where p is greater than or equal to 1 and p is a positive integer; determining a pth loss value corresponding to the embedding vector dimension of the pth round according to a preset loss function and a sample triple data of the pth round, the sample triple data of the pth round comprising a sample raw material name, sample relationship data and sample raw material attribute; and determining the embedding vector dimension of the pth round as the target dimension in a case where the pth loss value meets a preset threshold.
[0013] According to the embodiment of the present application, the determining of the pth loss value corresponding to the embedding vector dimension of the pth round according to the preset loss function and the obtained sample triple data of the pth round comprises: determining the first sample octuple vector of the pth round and the second sample octuple vector of the pth round from a plurality of octuple vectors according to the sample material name and sample relationship data included in the sample triple data of the pth round; reconstructing the first sample octuple vector of the pth round and the second sample octuple vector of the pth round to obtain the third sample octuple vector of the pth round and the fourth sample octuple vector of the pth round; evaluating the plurality of candidate data according to the third sample octuple vector of the pth round and the fourth sample octuple vector of the pth round to obtain the first confidence of each of the plurality of candidate data of the pth round; evaluating the sample material attribute included in the sample triple data of the pth round according to the third sample octuple vector of the pth round and the fourth sample octuple vector of the pth round to obtain the second confidence of the sample material attribute; and determining the pth loss value according to the loss function, the first confidence with the largest value in the pth round and the second confidence.
[0014] According to the first aspect of the present application, a data completion device for raw materials of a hydrometallurgical plant is provided, the device comprising: a data acquisition module for acquiring missing triple data, wherein the missing triple data comprises a material name, a material attribute related to the material name, and relationship data representing the association between the material name and the material attribute, and the material name or the material attribute in the missing triple data is missing; a vector determination module for determining a first octuple vector and a second octuple vector in a data encoding set according to non-missing data in the missing triple data; a vector reconstruction module for reconstructing the first octuple vector and the second octuple vector to obtain a third octuple vector and a fourth octuple vector; and a data determination module for determining target data in a plurality of candidate data according to the third octuple vector and the fourth octuple vector to complete the data of the missing triple data.
[0015] According to the embodiment of the present application, the first octuple vector and the second octuple vector are reconstructed to obtain the information lossless third octuple vector and the fourth octuple vector under lower computational overhead, and the plurality of times of interaction calculation according to the first octuple vector and the second octuple vector is replaced by the reconstruction, thereby reducing the data calculation amount in the process of determining the target data, further reducing the occupation of hardware computing resources, saving the subsequent computational overhead, improving the data processing efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0016] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:
[0017] Figure 1 An application scenario diagram of the missing data completion method for hydrometallurgical plants according to an embodiment of the present application is shown.
[0018] Figure 2 A flowchart of the missing data completion method for hydrometallurgical plants according to an embodiment of the present application is shown.
[0019] Figure 3 A schematic diagram of the missing data completion method for hydrometallurgical plants according to an embodiment of the present application is shown.
[0020] Figure 4 A flowchart of determining the target dimension according to an embodiment of the present application is shown.
[0021] Figure 5 A structural block diagram of the data completion device for hydrometallurgical plant raw materials according to an embodiment of the present application is shown.
[0022] Figure 6 A block diagram of an electronic device suitable for implementing the data completion method for hydrometallurgical plant raw materials according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0023] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, that the description is merely exemplary of the present application, and is intended to provide a thorough description for those skilled in the art to understand the present application. Therefore, the description is not intended to limit the scope of the present application. In the following detailed description of the embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present application.
[0024] The terms used herein are merely used to describe specific embodiments, and are not intended to limit the present application. The terms "include" and "have" and the like used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0025] All terms used herein, including technical and scientific terms, have the same meanings as those generally understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the present specification, and should not be interpreted in an idealized or excessively formal manner.
[0026] In the case of using expressions similar to "at least one of A, B, and C, etc.", it is generally intended to include at least one of A, B, and C, etc. (for example, "a system having at least one of A, B, and C" should include a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).
[0027] In the technical solutions of the present application, the collection, updating, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information) comply with the relevant legal regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to maintain user personal information security and network security.
[0028] With the continuous change of the metal extraction process of the hydrometallurgical plant, the data related to the usage amount and supply chain of the raw materials of the hydrometallurgical plant are also changing, and the hydrometallurgical plant staff often needs to determine the raw material cost according to the raw material data for the hydrometallurgical plant to reduce the raw material usage cost on the premise of promoting process improvement, but in actual production application, the raw material data is easy to cause data missing due to various reasons.
[0029] In the related art, in the database that stores historical raw material data, matching is performed in the database based on known information corresponding to the missing data to determine the missing data.
[0030] In the case of a large amount of raw material data, the method for determining missing data by matching in the related art needs to occupy a large amount of hardware resources for data query and matching, resulting in low data processing efficiency.
[0031] Embodiments of the present application provide a missing data completion method for a hydrometallurgical plant, characterized in that the method comprises: obtaining missing triple data, wherein the missing triple data comprises a raw material name, a raw material attribute related to the raw material name, and relationship data representing an association relationship between the raw material name and the raw material attribute, and the raw material name or the raw material attribute in the missing triple data is missing; determining a first eight-dimensional vector and a second eight-dimensional vector from a plurality of eight-dimensional vectors according to non-missing data in the missing triple data; reconstructing the first eight-dimensional vector and the second eight-dimensional vector to obtain a third eight-dimensional vector and a fourth eight-dimensional vector; and determining target data from a plurality of candidate data according to the third eight-dimensional vector and the fourth eight-dimensional vector to complete the data of the missing triple data.
[0032] Figure 1 An application scenario diagram of the missing data completion method for a hydrometallurgical plant according to an embodiment of the present application is shown.
[0033] As Figure 1 shown, the application scenario 100 according to this embodiment can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or fiber optic cables, and the like.
[0034] A user can use the first terminal device 101, the second terminal device 102, the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, and the like. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, and the like (only as examples).
[0035] The first terminal device 101, the second terminal device 102, the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, and the like.
[0036] The server 105 can be a server providing various services, such as a background management server providing support for a website browsed by a user using the first terminal device 101, the second terminal device 102, the third terminal device 103 (only as an example). The background management server can analyze and process received user requests and the like, and feed back the processing results (such as web pages, information, or data, and the like obtained or generated according to user requests) to the terminal device.
[0037] It should be noted that the missing data completion method for a hydrometallurgical plant provided by the embodiments of the present application can generally be executed by the server 105. Accordingly, the missing data completion device for a hydrometallurgical plant provided by the embodiments of the present application can generally be arranged in the server 105. The missing data completion method for a hydrometallurgical plant provided by the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Accordingly, the missing data completion device for a hydrometallurgical plant provided by the embodiments of the present application can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0038] It should be understood that Figure 1The number of the first terminal device, the second terminal device, the third terminal device, the network and the server in the figure is only an example. According to the implementation needs, there can be any number of the first terminal device, the second terminal device, the third terminal device, the network and the server.
[0039] The following will be based on Figure 1 The scenario described above, by Figures 2-4 The missing data completion method for hydrometallurgical plants of the embodiment of the application is described in detail.
[0040] Figure 2 A flowchart of the missing data completion method for hydrometallurgical plants according to the embodiment of the application is shown.
[0041] As Figure 2 shown, the missing data completion method for hydrometallurgical plants of the embodiment includes operation S210 to operation S240.
[0042] In operation S210, missing triple data is obtained, wherein the missing triple data includes a raw material name, a raw material attribute related to the raw material name, and relationship data representing an association relationship between the raw material name and the raw material attribute, and the raw material name or the raw material attribute in the missing triple data is missing.
[0043] In operation S220, according to the non-missing data in the missing triple data, a first eight-dimensional vector and a second eight-dimensional vector are determined from a plurality of eight-dimensional vectors.
[0044] In operation S230, the first eight-dimensional vector and the second eight-dimensional vector are reconstructed to obtain a third eight-dimensional vector and a fourth eight-dimensional vector.
[0045] In operation S240, according to the third eight-dimensional vector and the fourth eight-dimensional vector, target data is determined from a plurality of candidate data to complete the missing triple data.
[0046] According to the embodiment of the application, the missing triple data can be a raw material name and relationship data; the missing triple data can also be a raw material attribute and relationship data. The raw material name can be the name of the raw material used by the hydrometallurgical plant, and the raw material name can include electricity, sulfuric acid, lime, steam, etc. The relationship data can include material type, unit consumption per kiloton of solution, raw material origin, raw material unit price, etc. The raw material attribute can include chemical reagent, auxiliary material, energy medium, A province, 120 kg, etc.
[0047] For example, the raw material name is electricity, the relationship data of electricity is material type, and the raw material attribute corresponding to electricity and material type is energy medium. For another example, the raw material name is sulfuric acid, the relationship data of sulfuric acid is raw material origin, and the raw material attribute corresponding to sulfuric acid and raw material origin is B province.
[0048] The non-missing data represents data saved in the missing triple data, and the non-missing data can be the raw material name and the relationship data, or the raw material attribute and the relationship data.
[0049] The form of the eight-dimensional vector can be, for example:
[0050] ;
[0051] wherein, The eight-dimensional vector can represent any one of the raw material name, the raw material attribute, and the relationship data. The eight-dimensional vector is composed of eight vectors composed of k real numbers, that is, Any one of the eight-dimensional vectors is k-dimensional. The imaginary unit is represented.
[0052] The missing triple data is obtained, and according to the non-missing data in the missing triple data, matching is performed in the storage space in which a plurality of eight-dimensional vectors are stored, to determine the eight-dimensional vector corresponding to the non-missing data.
[0053] For example, in the case where the non-missing data can be the raw material name and the relationship data, the raw material name and the relationship data are matched in the storage space in which a plurality of eight-dimensional vectors corresponding to the non-missing data are stored, to determine the eight-dimensional vector corresponding to the raw material name and the eight-dimensional vector corresponding to the relationship data, that is, the first eight-dimensional vector and the second eight-dimensional vector.
[0054] The first eight-dimensional vector and the second eight-dimensional vector are reconstructed to mix the first eight-dimensional vector elements and the data of the second eight-dimensional vector in each eight-dimensional vector, to obtain a third eight-dimensional vector and a fourth eight-dimensional vector.
[0055] The plurality of candidate data can be the raw material name and the raw material attribute that have appeared in the history, and according to the matching degree between the third eight-dimensional vector and the fourth eight-dimensional vector and the plurality of candidate data, the candidate data with the highest matching degree with the second eight-dimensional vector and the fourth eight-dimensional vector is determined as the target data from the plurality of candidate data, and the target data is used to complete the data of the missing triple data.
[0056] Suppose the number of raw material names, relationship data, and raw material attributes is n, the same raw material name corresponds to n relationship data, and even the same raw material name and the same relationship data correspond to multiple raw material attributes, then the number of raw material name, relationship data, and raw material attribute triplets is n 3 In the process of determining the target data using the matching method, the time complexity of the algorithm is O (n 3 ).
[0057] In the method implementation process, the operation S210 and the operation S230 are operations that do not need to be executed in a loop, and the time complexity of the code for implementing the operation S210 and the operation S230 is O(1); in the operation S220, at most 2n rounds of loop execution are required to determine the first octuple vector and the second octuple vector, and therefore the time complexity of the code corresponding to the operation S220 is O(n); the operation S240 needs to calculate at most 2n candidate data, and the execution process can involve various vector operations of the third octuple vector and the fourth octuple vector, wherein the time complexity of the vector operation is related to the dimension of the vector, and the dimension of the vector is determined, and therefore the time complexity of the code corresponding to the operation S240 is at most O(n).
[0058] Therefore, the time complexity of the operation S210 to the operation S240 is at most O(n), and the operation S210 to the operation S240 reduces the time complexity required for determining the target data, thereby reducing the occupation time of the processor and other computing resources in the operation process and improving the data processing efficiency.
[0059] According to the embodiments of the present application, the first octuple vector and the second octuple vector are reconstructed to obtain the information-lossless third octuple vector and the fourth octuple vector at a lower calculation overhead, and the multiple times of interaction calculation based on the first octuple vector and the second octuple vector is replaced by the reconstruction, thereby reducing the data calculation amount in the process of determining the target data, further reducing the occupation of the hardware computing resources, saving the subsequent calculation overhead, and improving the data processing efficiency.
[0060] According to the embodiments of the present application, the missing data completion method for the hydrometallurgical plant further includes: encoding the initial data according to the target dimension and the value mapping range to obtain an embedding vector, the initial data being a raw material name, a raw material attribute, or relationship data; dividing the embedding vector to obtain eight embedding sub-vectors; and determining an octuple vector corresponding to the initial data based on the eight embedding sub-vectors.
[0061] According to the embodiments of the present application, the initial data is any one of the raw material data, that is, the initial data represents any one of the raw material name, the raw material attribute, or the relationship data. The initial data can be obtained from a database or other storage space for storing historical raw material data, and the initial data is encoded according to the target dimension and the data mapping range, so that the vector dimension of the embedding vector is the target dimension, and the value range of each element in the embedding vector is the data mapping range, wherein the target dimension can be an integer multiple of eight; and the value mapping range can be an interval (0, 1).
[0062] The embedding vector obtained by encoding the initial data is divided, and in the case that the dimension of the embedding vector is 8k, the embedding vector is divided into eight embedding sub-vectors of k dimensions. Based on the eight embedding sub-vectors and the preset seven imaginary units, an octonion vector is constructed.
[0063] For example, the dimension of 8k is divided to obtain eight vectors composed of k real numbers , that is, eight k-dimensional embedding sub-vectors are obtained, based on the preset seven data units , an octonion vector is constructed , wherein the octonion vector includes a real part data and seven imaginary part data .
[0064] The plurality of initial data in the database or other storage space is traversed to determine a plurality of octonion vectors, and the plurality of initial data and the plurality of octonion vectors corresponding to the initial data are stored in the storage space in the form of key-value pairs, so as to determine a first octonion vector and a second octonion vector from the plurality of octonion vectors according to the non-missing data in the missing triple data.
[0065] Each element in the octonion vector is an octonion, which has a richer structure than a real number, carries a larger amount of information, and increases the interpretability of the initial data.
[0066] According to the embodiment of the application, the embedding vector obtained by encoding is divided to determine the octonion vector corresponding to the initial data based on the eight embedding sub-vectors obtained by division, and the initial data is represented by the octonion vector, thereby reducing the dimension of the vector and saving the subsequent calculation overhead and improving the data processing efficiency.
[0067] According to the embodiment of the application, the initial data is encoded to obtain an embedding vector according to a target dimension and a value mapping range, including: initializing a plurality of values same as the target dimension randomly in the value mapping range to obtain the embedding vector.
[0068] According to the embodiment of the application, the initial data is encoded by randomly selecting data in the value mapping range to determine the value of each element in the embedding vector. This random initialization method no longer requires a high vector dimension, i.e., does not require elements with multiple values of 0 to construct the embedding vector, and can also distinguish between multiple different initial data, thereby reducing the dimension and sparsity of the embedding vector and reducing the calculation and storage overhead in the early stage.
[0069] Figure 3 A schematic diagram of a missing data completion method for a hydrometallurgical plant according to an embodiment of the application is shown.
[0070] As Figure 3As shown, the non-missing data includes first data 301 and second data 302, in the case of the first data 301 being relational data, the second data 302 is a raw material name or a raw material attribute; in the case of the first data 301 being a raw material name or a raw material attribute, the second data 302 is relational data.
[0071] For the first data 301 and the second data 302, a plurality of numerical values of the same dimension as the target dimension are randomly initialized in the numerical mapping range, to obtain a first embedding vector 303 corresponding to the first data 301 and a second embedding vector 304 corresponding to the second data 302.
[0072] The first embedding vector 303 and the second embedding vector 304 are respectively divided to obtain eight first embedding sub-vectors corresponding to the embedding vector 303, i.e., the first embedding sub-vector 305_1, the first embedding sub-vector 305_2, the first embedding sub-vector 305_3, the first embedding sub-vector 305_4, the first embedding sub-vector 305_5, the first embedding sub-vector 305_6, the first embedding sub-vector 305_7, and the first embedding sub-vector 305_8, and eight second embedding sub-vectors corresponding to the embedding vector 304, i.e., the second embedding sub-vector 306_1, the second embedding sub-vector 306_2, the second embedding sub-vector 306_3, the second embedding sub-vector 306_4, the second embedding sub-vector 306_5, the second embedding sub-vector 306_6, the second embedding sub-vector 306_7, and the second embedding sub-vector 306_8.
[0073] Based on the first embedding sub-vector 305_1 to the first embedding sub-vector 305_8, a first eight-dimensional vector 307 corresponding to the first data 301 is constructed and determined; based on the second embedding sub-vector 306_1 to the first embedding sub-vector 306_8, a second eight-dimensional vector 308 corresponding to the second data 302 is constructed and determined.
[0074] The first eight-dimensional vector 307 and the second eight-dimensional vector 308 are reconstructed to obtain a third eight-dimensional vector 309 and a fourth eight-dimensional vector 310. According to the third eight-dimensional vector 309 and the fourth eight-dimensional vector 310, a target data 312 is determined from the preset candidate data 311_1, candidate data 311_2, …, candidate data 311_m, to perform data completion on the missing triple data, wherein m is a positive integer.
[0075] According to the embodiment of the application, by means of random initialization, a higher vector dimension is no longer required, i.e., a plurality of numerical values of 0 are not required to construct the embedding vector, and a plurality of different initial data can also be distinguished, thereby reducing the dimension and sparsity of the embedding vector.
[0076] Further, the first eight-dimensional vector and the second eight-dimensional vector are constructed based on the first embedding vector and the second embedding vector with smaller dimensions, compared to vector operations based on the first embedding vector and the second embedding vector, the first eight-dimensional vector and the second eight-dimensional vector have smaller dimensions, and operations on the first eight-dimensional vector and the second eight-dimensional vector reduce the amount of data operations.
[0077] The first eight-dimensional vector and the second eight-dimensional vector are reconstructed to obtain the third eight-dimensional vector and the fourth eight-dimensional vector for determining the target data from the plurality of candidate data, which optimizes the data calculation path for determining the target data and further improves the data processing efficiency.
[0078] The target dimension, i.e., the dimension of the embedding vector, determines the dimensions of the eight embedding sub-vectors in the eight-dimensional vector. The higher the dimension of the embedding sub-vector, the lower the data processing efficiency, but the more detailed and complex the representation information of the eight-dimensional vector. It also needs to be considered that the higher the dimension of the embedding sub-vector, the more redundant and noisy the eight-dimensional vector may be. Therefore, the target dimension needs to be determined to meet the accuracy requirement of the target data and have a higher data processing efficiency.
[0079] According to the embodiments of the present application, the missing data completion method for a hydrometallurgical plant further comprises: repeatedly performing the following operations until the target dimension is determined from the plurality of embedding vector dimensions: determining the embedding vector dimension of the p th round from the plurality of embedding vector dimensions, where p ≥ 1 and p is a positive integer; determining the p th loss value corresponding to the embedding vector dimension of the p th round according to the preset loss function and the obtained sample triple data of the p th round, the sample triple data of the p th round comprising sample raw material name, sample relationship data and sample raw material attribute; and in the case that the p th loss value meets the preset threshold, determining the embedding vector dimension of the p th round as the target dimension.
[0080] According to the embodiments of the present application, the plurality of embedding vector dimensions can be preset based on the experience of determining the target dimension in the historical process, and the embedding vector dimension is a multiple of eight. In the plurality of embedding vector dimensions, the embedding vector dimensions can be selected in ascending order to determine the embedding vector dimension of the p th round.
[0081] The sample triple data comprises sample raw material name, sample relationship data and sample raw material attribute. The description of the sample raw material name, sample relationship data and sample raw material attribute can be referred to the above-mentioned related content of the raw material name, relationship data and raw material attribute, which will not be repeated here.
[0082] The sample triple data of the p th round is obtained, the sample raw material name or the sample raw material attribute is taken as the missing data, and the p th loss value corresponding to the embedding vector dimension of the p th round is determined according to the preset loss function.
[0083] The preset threshold value can be 0.05, and in a case where the pth loss value meets the preset threshold value, the embedding vector dimension of the pth round is determined as the target dimension.
[0084] According to the embodiment of the present application, the loss value corresponding to different embedding vector dimensions is determined by using sample triple data, and the target dimension meeting the preset threshold value is determined, so that the target dimension with smaller embedding vector dimension is determined in a case where the accuracy requirement of the target data can be met, and the time for occupying the processor and other computing resources in the process of determining the target data is reduced, and the data processing efficiency is improved.
[0085] According to the embodiment of the present application, the pth loss value corresponding to the embedding vector dimension of the pth round is determined according to the preset loss function and the sample triple data of the pth round, comprising: determining the first sample eight vector of the pth round and the second sample eight vector of the pth round from the plurality of eight vectors according to the sample material name and sample relationship data included in the sample triple data of the pth round; reconstructing the first sample eight vector of the pth round and the second sample eight vector of the pth round to obtain the third sample eight vector of the pth round and the fourth sample eight vector of the pth round; evaluating the plurality of candidate data according to the third sample eight vector of the pth round and the fourth sample eight vector of the pth round to obtain the first confidence of each of the plurality of candidate data; evaluating the sample material attribute included in the sample triple data of the pth round according to the third sample eight vector of the pth round and the fourth sample eight vector of the pth round to obtain the second confidence of the sample material attribute; and determining the pth loss value according to the loss function, the first confidence and the second confidence with the largest value in the pth round.
[0086] According to the embodiment of the present application, the description about determining the first sample eight vector and the second sample eight vector, and reconstructing the first sample eight vector and the second sample eight vector can refer to the related content about determining the first eight vector and the second eight vector, and reconstructing the first sample eight vector and the second sample eight vector, which will not be repeated here.
[0087] The first confidence is used to evaluate the possibility that the candidate data is the missing data in the sample triple data. The eight vector corresponding to each of the plurality of candidate data is determined from the plurality of eight vectors, and the third sample eight vector, the fourth sample eight vector and the eight vector of each of the candidate data are calculated respectively to evaluate the matching degree of the candidate data and the third eight vector and the fourth eight vector, and the first confidence of the candidate data is obtained. In the plurality of candidate data, the candidate data with the highest first confidence is determined as the target candidate data.
[0088] The second confidence degree is used as a reference value of the first confidence degree. The eight-dimensional vector corresponding to the sample raw material attribute is determined from the plurality of eight-dimensional vectors, and the second confidence degree of the sample raw material attribute is calculated according to the third sample eight-dimensional vector, the fourth sample eight-dimensional vector and the eight-dimensional vector corresponding to the sample raw material attribute.
[0089] The pth loss value is determined according to the loss function, the first confidence degree with the maximum value in the pth round and the second confidence degree of the pth round.
[0090] According to an embodiment of the present application, the loss value corresponding to different embedding vector dimensions is determined by using sample triple data, and a target dimension satisfying a preset threshold is determined, so that the target dimension with smaller embedding vector dimension is determined in a case where the accuracy requirement of the target data can be met, and the time for occupying the processor and other computing resources in the process of determining the target data is reduced, and the data processing efficiency is improved. According to an embodiment of the present application, a plurality of different sample triple data can be used in each round to calculate the loss value of one round. Specifically, the pth loss value can be determined by formula (1):
[0091] (1);
[0092] Wherein, L represents the loss value, represents the sample raw material name, represents the sample relationship data, represents the sample raw material attribute, represents the candidate data with the maximum first confidence degree in the case where the preset missing data is the sample raw material name, represents the candidate data with the maximum first confidence degree in the case where the preset missing data is the sample raw material attribute.
[0093] represents a data set composed of a plurality of sample triple data. In the case where each round includes a plurality of sample triple data, the loss value corresponding to the plurality of sample triple data can be obtained, and the sum of the loss values corresponding to the plurality of sample triple data is taken as the pth loss value corresponding to the embedding vector dimension of the pth round. The number of sample triple data in the data set affects the speed and accuracy of determining the target dimension.
[0094] is a confidence degree calculation formula, represents an exponential function with a natural constant e (approximately equal to 2.71828) as the base, the function is a logarithmic function with 10 as the base. represents the first confidence degree, represents the second confidence degree.
[0095] According to an embodiment of the present application, a negative sample is introduced in each round to determine the target dimension, the negative sample being a triple data with data error, for example, the sample triple data being (sulfuric acid, material type, chemical reagent), and the negative sample being (sulfuric acid, material type, energy medium). The loss value after introduction of the negative sample is determined by formula (2) to improve the accuracy of the target dimension, thereby improving the accuracy of the target data.
[0096] (2);
[0097] wherein, represents a set of triple data composed of the negative sample. represents the loss value after introduction of the negative sample.
[0098] Figure 4 A flowchart for determining the target dimension according to an embodiment of the present application is shown.
[0099] As shown in Figure 4 , the process for determining the target dimension of this embodiment includes operations S410-S460.
[0100] In operation S410, the embedding vector dimension of the 1st round is determined from a plurality of embedding vector dimensions.
[0101] In operation S420, the 1st loss value corresponding to the embedding vector dimension of the 1st round is determined according to the preset loss function and the obtained sample triple data of the 1st round.
[0102] In operation S430, the embedding vector dimension of the qth round is determined from a plurality of embedding vector dimensions, wherein q≥2, and q is a positive integer.
[0103] In operation S440, the qth loss value corresponding to the embedding vector dimension of the qth round is determined according to the preset loss function and the obtained sample triple data of the qth round.
[0104] In operation S450, it is determined whether the qth loss value and the q-1th loss value satisfy a preset condition.
[0105] The preset condition can be that the qth loss value and the q-1th loss value are less than a preset gradient change threshold, wherein the preset gradient change threshold can be 0.001, 0.002, etc.
[0106] If the qth loss value and the q-1th loss value satisfy the preset condition, operation S460 is performed, otherwise, operation S430 is performed.
[0107] In operation S460, the embedding vector dimension of the qth round is determined as the target dimension.
[0108] According to the embodiment of the present application, the target dimension is determined when the change trend of the loss values of the two continuous rounds tends to be stable, the calculation of redundant rounds is reduced, and the data processing efficiency is improved.
[0109] According to the embodiment of the present application, the first eight-element vector and the second eight-element vector are reconstructed to obtain a third eight-element vector and a fourth eight-element vector, including: exchanging the same position and the same number of imaginary part data in the first eight-element vector and the second eight-element vector to obtain the third eight-element vector and the fourth eight-element vector.
[0110] According to the embodiment of the present application, the imaginary part data represents data with imaginary units. The same position and the same number of imaginary part data in the first eight-element vector and the second eight-element vector are exchanged, so that the third eight-element vector includes both the data of the first eight-element vector and the data of the second eight-element vector, and the fourth eight-element vector includes both the data of the first eight-element vector and the data of the second eight-element vector.
[0111] Specifically, the number of exchanged imaginary part data can be four.
[0112] For example, the first eight-element vector is , and the second eight-element vector is:
[0113] .
[0114] The vector with a dimension of k is represented by real numbers. The last four imaginary part data of the first eight-element vector and the second eight-element vector are exchanged to obtain a third eight-element vector and a fourth eight-element vector .
[0115] According to the embodiment of the present application, the first eight-element vector and the second eight-element vector are reconstructed by exchanging the same position and the same number of imaginary part data in the first eight-element vector and the second eight-element vector, which saves the calculation overhead of multiple vector interaction calculations on the first eight-element vector and the second eight-element vector, and improves the data processing efficiency.
[0116] According to the embodiment of the present application, the target data is determined from the third eight-element vector and the fourth eight-element vector among the plurality of candidate data, including: evaluating the plurality of candidate data according to the third eight-element vector and the fourth eight-element vector to obtain the confidence of each of the plurality of candidate data; and determining the candidate data with the highest confidence as the target data.
[0117] According to an embodiment of the present application, the confidence is used to represent the possibility of the candidate data belonging to the missing triple data. The octuple vector corresponding to each of the candidate data is determined from the plurality of octuple vectors, and the confidence of each of the candidate data is calculated according to the third octuple vector, the fourth octuple vector and the octuple vector corresponding to each of the candidate data, respectively, to evaluate the matching degree of each of the candidate data with the third octuple vector and the fourth octuple vector. The candidate data with the highest confidence is determined as the target candidate data from the plurality of candidate data.
[0118] According to an embodiment of the present application, the target data is determined from the plurality of candidate data, which reduces the selectable space of selecting the target data, thereby reducing the number of data operations and improving the data processing efficiency.
[0119] According to an embodiment of the present application, the confidence of each of the candidate data is determined according to the third octuple vector and the fourth octuple vector, including: determining the interaction vector according to the third octuple vector and the fourth octuple vector; and determining the confidence of each of the candidate data according to the product of the interaction vector and the octuple vector corresponding to each of the candidate data.
[0120] According to an embodiment of the present application, the fourth octuple vector is fitted according to the vector direction of the third octuple vector to determine the interaction vector, or the third octuple vector is fitted according to the vector direction of the fourth octuple vector to determine the interaction vector.
[0121] The confidence of the candidate data can be determined by formula (3):
[0122] (3);
[0123] wherein, represents the confidence, represents the interaction vector, represents the octuple vector corresponding to the candidate data, represents the inner product.
[0124] The confidence of each of the candidate data is determined according to the product of the interaction vector and the octuple vector corresponding to each of the candidate data. Since the interaction vector is determined according to the third octuple vector and the fourth octuple vector, the greater the product of the octuple vector corresponding to the candidate data and the interaction vector, the smaller the angle between the interaction vector and the octuple vector corresponding to the candidate data, the higher the similarity between the interaction vector and the octuple vector corresponding to the candidate data, and the higher the possibility of the candidate data belonging to the missing triple data.
[0125] According to the embodiment of the present application, the confidence of each candidate data is determined by the product of the interaction vector and the eight-dimensional vector represented by each candidate data, the high-dimensional, multi-step and redundant confidence calculation is compressed into a low-dimensional, single-step vector multiplication operation, and the data processing efficiency is improved.
[0126] According to the embodiment of the present application, the interaction vector is determined according to the third eight-dimensional vector and the fourth eight-dimensional vector, comprising: determining a unit eight-dimensional vector according to the ratio of the fourth eight-dimensional vector and the norm of the fourth eight-dimensional vector; and determining the interaction vector according to the product of the third eight-dimensional vector and the unit eight-dimensional vector.
[0127] According to the embodiment of the present application, the fourth eight-dimensional vector is converted into a unit eight-dimensional vector to retain the direction of the fourth eight-dimensional vector and discard the norm of the fourth eight-dimensional vector, so as to determine the unit eight-dimensional vector for fitting the third eight-dimensional vector.
[0128] The unit eight-dimensional vector can be determined by formula (4):
[0129] (4);
[0130] wherein, represents the fourth eight-dimensional vector, represents the norm of the fourth eight-dimensional vector, represents the unit eight-dimensional vector.
[0131] The interaction vector can be determined by formula (5):
[0132] (5);
[0133] wherein, represents the third eight-dimensional vector, represents the product.
[0134] The interaction vector is determined according to the product of the third eight-dimensional vector and the unit eight-dimensional vector, and the unit eight-dimensional vector is determined according to the ratio of the fourth eight-dimensional vector and the norm of the fourth eight-dimensional vector, so the unit eight-dimensional vector discards the norm of the fourth eight-dimensional vector, thereby reducing the probability of unstable deviation of the interaction vector caused by the introduction of the norm of the fourth eight-dimensional vector in the calculation process.
[0135] According to the embodiment of the present application, the interaction vector is determined according to the product of the third eight-dimensional vector and the unit eight-dimensional vector, and the unit eight-dimensional vector discards the norm of the fourth eight-dimensional vector, so the stability of the interaction vector is improved, the process of processing the interaction vector with unstable deviation is omitted, and the data processing efficiency is improved.
[0136] According to an embodiment of the present application, the interaction vector is determined according to the third octonion vector and the fourth octonion vector, and the method further comprises: determining a unit octonion vector corresponding to the third octonion vector according to a ratio of a norm of the third octonion vector and the third octonion vector; and determining the interaction vector according to a product of the fourth octonion vector and the unit octonion vector.
[0137] According to an embodiment of the present application, the method of the present application for determining target data (Our) and the method of using other network models for determining target data are compared in three evaluation indexes, including Mean Absolute Error (MAE), Root Mean Squared Error (RMSE) and R-Squared or Coefficient of Determination (R 2 ) in Table 1, wherein the other network models include Long Short-Term Memory (LSTM) model, Gated Recurrent Unit (GRU) model, Bidirectional Temporal Convolutional Network (BTDC) model, and a hybrid network model combining Convolutional Neural Network (CNN) and gated neural network; and the K-Nearest Neighbors Imputation (KNNI) is also included in Table 1.
[0138] Table 1
[0139]
[0140] Based on the above data completion method for raw materials of hydrometallurgical plants, the present application further provides a data completion device for raw materials of hydrometallurgical plants. Figure 5 The device will be described in detail below.
[0141] Figure 5 A structural block diagram of the data completion device for raw materials of hydrometallurgical plants according to an embodiment of the present application is shown.
[0142] As Figure 5 shown, the data completion device for raw materials of hydrometallurgical plants of this embodiment comprises a data acquisition module 510, a vector determination module 520, a vector reconstruction module 530 and a data determination module 540.
[0143] The data acquisition module 510 is configured to acquire missing triple data, wherein the missing triple data comprises a raw material name, a raw material attribute related to the raw material name, and relationship data representing an association relationship between the raw material name and the raw material attribute, and the raw material name or the raw material attribute in the missing triple data is missing.
[0144] The vector determination module 520 is configured to determine a first eight-element vector and a second eight-element vector in the data encoding set according to non-missing data in the missing triple data.
[0145] The vector reconstruction module 530 is configured to reconstruct the first eight-element vector and the second eight-element vector to obtain a third eight-element vector and a fourth eight-element vector.
[0146] The data determination module 540 is configured to determine target data in a plurality of candidate data according to the third eight-element vector and the fourth eight-element vector, so as to complete data of the missing triple data.
[0147] According to the embodiment of the present application, the vector reconstruction module 530 comprises a data interchanging subunit configured to interchange a plurality of imaginary part data of the same position and the same quantity in the first eight-element vector and the second eight-element vector to obtain the third eight-element vector and the fourth eight-element vector.
[0148] According to the embodiment of the present application, the data determination module 540 comprises a confidence degree determination sub-module and a target determination sub-module.
[0149] The confidence degree determination sub-module is configured to evaluate the plurality of candidate data according to the third eight-element vector and the fourth eight-element vector to obtain a confidence degree of each of the plurality of candidate data.
[0150] The target determination sub-module is configured to determine a candidate data with the highest confidence degree in the candidate data as the target data.
[0151] According to the embodiment of the present application, the confidence degree determination sub-module comprises a vector determination unit and a confidence degree determination unit.
[0152] The vector determination unit is configured to determine an interaction vector according to the third eight-element vector and the fourth eight-element vector.
[0153] The confidence degree determination unit is configured to determine the confidence degree of each of the plurality of candidate data according to a product of the interaction vector and an eight-element vector representing each of the plurality of candidate data.
[0154] According to the embodiment of the present application, the vector determination unit comprises a unit determination subunit and a vector determination subunit.
[0155] The unit determination subunit is configured to determine a unit eight-element vector according to a ratio of a norm of the fourth eight-element vector to the fourth eight-element vector.
[0156] A vector determination subunit is configured to determine the interaction vector according to a product of the third eight-element vector and the unit eight-vector.
[0157] According to an embodiment of the present application, the data completion device for raw materials of a hydrometallurgical plant further comprises an encoding module, a vector division module, and a vector determination module.
[0158] The encoding module is configured to encode the initial data according to the target dimension and the value mapping range, to obtain an embedding vector, the initial data being a raw material name, or a raw material attribute, or relationship data.
[0159] The vector division module is configured to divide the embedding vector to obtain eight embedding sub-vectors.
[0160] The vector determination module is configured to determine, based on the eight embedding sub-vectors, an eight-element vector corresponding to the initial data.
[0161] According to an embodiment of the present application, the encoding module comprises a vector determination submodule.
[0162] The vector determination submodule is configured to randomly initialize a plurality of values identical to the target dimension within the value mapping range, to obtain the embedding vector.
[0163] According to an embodiment of the present application, the data completion device for raw materials of a hydrometallurgical plant further comprises a dimension determination module.
[0164] The dimension determination module comprises a dimension determination submodule, a loss determination submodule, and a target determination submodule, and is configured to repeatedly perform the following operations until the target dimension is determined from a plurality of embedding vector dimensions: the dimension determination submodule is configured to determine, in the plurality of embedding vector dimensions, an embedding vector dimension of the pth round, where p≥1 and p is a positive integer. The loss determination submodule is configured to determine, according to a preset loss function and a pth round of sample triple data, a pth loss value corresponding to the embedding vector dimension of the pth round, the pth round of sample triple data comprising a sample raw material name, sample relationship data, and sample raw material attributes. The target determination submodule is configured to determine, in a case where the pth loss value satisfies a preset threshold, the embedding vector dimension of the pth round as the target dimension.
[0165] According to an embodiment of the present application, the loss determination submodule comprises a sample determination unit, a vector reconstruction unit, a first evaluation unit, a second evaluation unit, and a loss determination unit.
[0166] The sample determination unit is configured to determine, according to the sample raw material name and the sample relationship data included in the pth round of sample triple data, a first sample eight-element vector of the pth round and a second sample eight-element vector of the pth round from the plurality of eight-element vectors.
[0167] a vector reconstruction unit configured to reconstruct the first sample octonion vector of the pth round and the second sample octonion vector of the pth round to obtain a third sample octonion vector of the pth round and a fourth sample octonion vector of the pth round.
[0168] a first evaluation unit configured to evaluate the plurality of candidate data according to the third sample octonion vector of the pth round and the fourth sample octonion vector of the pth round to obtain a first confidence degree of each of the plurality of candidate data of the pth round.
[0169] a second evaluation unit configured to evaluate a sample raw material attribute included in the sample triple data of the pth round according to the third sample octonion vector of the pth round and the fourth sample octonion vector of the pth round to obtain a second confidence degree of the sample raw material attribute.
[0170] a loss determination unit configured to determine a pth loss value according to a loss function and according to the first confidence degree with the largest value and the second confidence degree of the pth round.
[0171] According to an embodiment of the present application, any of the data acquisition module 510, the vector determination module 520, the vector reconstruction module 530 and the data determination module 540 can be combined in one module, or any of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of other modules, and implemented in one module. According to an embodiment of the present application, at least one of the data acquisition module 510, the vector determination module 520, the vector reconstruction module 530 and the data determination module 540 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging a circuit, etc. hardware or firmware, or in any one of software, hardware and firmware or in any appropriate combination of any of them. Alternatively, at least one of the data acquisition module 510, the vector determination module 520, the vector reconstruction module 530 and the data determination module 540 can be at least partially implemented as a computer program module which can perform the corresponding function when executed.
[0172] Figure 6 A block diagram of an electronic device suitable for implementing the data completion method for raw materials of hydrometallurgical plants according to an embodiment of the present application is shown.
[0173] As Figure 6As shown, the electronic device 600 according to an embodiment of the present application includes a processor 601 which can perform various appropriate actions and processes in accordance with a program stored in a read only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chip set, and / or a special purpose microprocessor (e.g., an application specific integrated circuit (ASIC)), and so on. The processor 601 can also include an on-board memory for cache use. The processor 601 can include a single processing unit or multiple processing units to perform the various actions of the method processes according to embodiments of the present application.
[0174] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method processes according to embodiments of the present application by executing the programs in the ROM 602 and / or the RAM 603. Note that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method processes according to embodiments of the present application by executing the programs stored in the one or more memories.
[0175] According to an embodiment of the present application, the electronic device 600 can further include an input / output (I / O) interface 605 which is also connected to the bus 604. The electronic device 600 can further include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as necessary. A removable recording medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 610 as necessary, so that a computer program read out therefrom is installed in the storage section 608 as necessary.
[0176] The application further provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments, or can exist independently without being assembled into the device / apparatus / system. The computer readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the application.
[0177] According to the embodiments of the application, the computer readable storage medium can be a non-volatile computer readable storage medium, which can include, but is not limited to, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In the application, the computer readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in connection with an instruction execution system, apparatus or device. For example, according to the embodiments of the application, the computer readable storage medium can include one or more memories of the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603.
[0178] The embodiments of the application also include a computer program product, which includes a computer program containing program codes for executing the method shown in the flow chart. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the data completion method for raw materials of hydrometallurgical plants provided by the embodiments of the application.
[0179] The above functions defined in the system / apparatus of the embodiments of the application are performed when the computer program is executed by the processor 601. According to the embodiments of the application, the system, apparatus, module, unit, etc. described above can be implemented by computer program modules.
[0180] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, downloaded and installed in the form of signals on a network medium, and be downloaded and installed through the communication part 609 and / or installed from the detachable medium 611. The program codes contained in the computer program can be transmitted by any appropriate network medium, including but not limited to wireless, wired, etc., or any appropriate combination thereof.
[0181] In such embodiments, the computer program can be downloaded and installed from the network via the communication section 609, and / or installed from the removable media 611. When the computer program is executed by the processor 601, the above-described functions defined in the system of the embodiments of the present application are executed. The system, device, apparatus, module, unit, etc. described above can be realized by the computer program modules according to the embodiments of the present application.
[0182] According to the embodiments of the present application, the program code for executing the computer program provided by the embodiments of the present application can be written in any combination of one or more programming languages, and specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming language, and / or assembly / machine language. The programming language includes, but is not limited to, such as Java, C++, python, "C" language or similar programming language. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, connected to the Internet through an Internet service provider).
[0183] The flowcharts and block diagrams in the drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the figures. For example, two blocks noted in succession can actually be executed substantially concurrently or in the opposite order, depending on the functionality involved. It should also be noted that each block in the flowcharts or block diagrams, and combinations of blocks in the flowcharts or block diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0184] Those skilled in the art can understand that the features described in various embodiments of the present application can be combined and / or integrated in various combinations, even if such combinations are not explicitly described in the present application. In particular, the features described in various embodiments of the present application can be combined and / or integrated in various combinations without departing from the spirit and teachings of the present application. All such combinations and / or integrations are within the scope of the present application.
[0185] The embodiments of the application have been described. However, these embodiments are merely for illustration and are not intended to limit the scope of the application. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Various alternatives and modifications to the embodiments described herein will be apparent to those skilled in the art in view of the foregoing without departing from the scope of the application.
Claims
1. A method for missing data imputation for hydrometallurgical plants, characterized by, The method comprises: acquiring missing triple data, wherein the missing triple data comprises a raw material name, a raw material attribute related to the raw material name, and relationship data representing an association between the raw material name and the raw material attribute, and the raw material name or the raw material attribute in the missing triple data is missing; determining a first eight-dimensional vector and a second eight-dimensional vector from a plurality of eight-dimensional vectors according to non-missing data in the missing triple data; interchanging a plurality of imaginary part data of the same position and the same number in the first eight-dimensional vector and the second eight-dimensional vector to obtain a third eight-dimensional vector and a fourth eight-dimensional vector; determining an interaction vector according to the third eight-dimensional vector and the fourth eight-dimensional vector; determining a confidence degree of each of a plurality of candidate data according to a product of the interaction vector and an eight-dimensional vector representing each of the plurality of candidate data; determining a candidate data with the highest confidence degree as target data to complete the missing triple data.
2. The method of claim 1, wherein, The method further comprises: encoding initial data according to a target dimension and a numerical value mapping range to obtain an embedding vector, wherein the initial data is the raw material name, the raw material attribute, or the relationship data; dividing the embedding vector to obtain eight embedding sub-vectors; 3. The method of claim 1, wherein, determining an eight-dimensional vector corresponding to the initial data based on the eight embedding sub-vectors. The method further comprises: randomly initializing a plurality of numerical values same as the target dimension within the numerical value mapping range to obtain the embedding vector. The method further comprises:
4. The method of claim 3, wherein, repeating the following operations until the target dimension is determined from a plurality of embedding vector dimensions: determining an embedding vector dimension of the p th round from the plurality of embedding vector dimensions, wherein p≥1 and p is a positive integer; 5. The method of claim 4, wherein, determining a p th loss value corresponding to the embedding vector dimension of the p th round according to a preset loss function and a sample triple data of the p th round, wherein the sample triple data of the p th round comprises a sample raw material name, a sample relationship data, and a sample raw material attribute; in a case where the p th loss value meets a preset threshold, determining the embedding vector dimension of the p th round as the target dimension. The method further comprises: determining a p th sample eight-dimensional vector of the first and a p th sample eight-dimensional vector of the second from the plurality of eight-dimensional vectors according to the sample raw material name and the sample relationship data comprised in the sample triple data of the p th round. 6. The method of claim 5, wherein, reconstructing the first sample octuple vector of the p th round and the second sample octuple vector of the p th round to obtain a third sample octuple vector of the p th round and a fourth sample octuple vector of the p th round; evaluating the plurality of candidate data according to the third sample octuple vector of the p th round and the fourth sample octuple vector of the p th round to obtain a first confidence degree of each of the plurality of candidate data of the p th round; evaluating the sample raw material attribute included in the sample triple data of the p th round according to the third sample octuple vector of the p th round and the fourth sample octuple vector of the p th round to obtain a second confidence degree of the sample raw material attribute; determining the p th loss value according to the loss function, the first confidence degree and the second confidence degree with the largest numerical value in the p th round.
7. A data completion device for raw material of a hydrometallurgical plant, characterized by, The device comprises: a data acquisition module configured to acquire missing triple data, wherein the missing triple data comprises a raw material name, a raw material attribute related to the raw material name, and relationship data representing an association between the raw material name and the raw material attribute, and the raw material name or the raw material attribute in the missing triple data is missing; a vector determination module configured to determine a first octuple vector and a second octuple vector in a data encoding set according to non-missing data in the missing triple data; a vector reconstruction module configured to reconstruct the first octuple vector and the second octuple vector to obtain a third octuple vector and a fourth octuple vector; a data determination module configured to determine target data in a plurality of candidate data according to the third octuple vector and the fourth octuple vector to perform data completion on the missing triple data; The vector reconstruction module comprises: a data exchange subunit configured to exchange a plurality of imaginary part data of the same position and the same number in the first octuple vector and the second octuple vector to obtain the third octuple vector and the fourth octuple vector; The data determination module comprises: a confidence degree determination submodule configured to evaluate the plurality of candidate data according to the third octuple vector and the fourth octuple vector to obtain a confidence degree of each of the plurality of candidate data; a target determination submodule configured to determine a candidate data with the highest confidence degree in the candidate data as the target data; The confidence degree determination submodule comprises: a vector determination unit configured to determine an interaction vector according to the third octuple vector and the fourth octuple vector; a confidence degree determination unit configured to determine a confidence degree of each of the plurality of candidate data according to a product of the interaction vector and an octuple vector representing each of the plurality of candidate data.
Citation Information
Patent Citations
Complementation method and device of power distribution network knowledge graph, electronic equipment, storage medium and program product
CN118410182A
Internet software development management system and method
CN120215889A