Training Method, Device, Electronic Device and Storage Medium of Global Preprocessing Model
Through federal training technology, each logistics platform trains local preprocessing models based on its own logistics address data sources, and performs parameter exchange and aggregation, solving the problem of poor training effect of preprocessing models in the existing technology and achieving higher recognition accuracy and universality.
Patent Information
- Application Number
- CN202210090185.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-01-25
AI Technical Summary
During the training process, the existing preprocessing model cannot be collected on a large scale due to the privacy of address data, resulting in poor training results and poor accuracy and universality.
The model training method of federated training is adopted, and the local preprocessing model is obtained through the training of each logistics platform based on its own logistics address data source, and then the model parameter exchange and aggregation are performed to obtain the global preprocessing model.
The recognition accuracy and universality of the preprocessing model are improved, so that the model can better identify a larger range of address data, especially address data in remote areas.
Smart Images

Figure CN114492406B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for training a global preprocessing model. Background Art
[0002] With the rapid development of e-commerce business and the increasing popularity of online shopping, the volume of logistics business has also increased sharply. In the process of creating a logistics order for a logistics business, the identification of logistics address information including the recipient address and the sender address is indispensable.
[0003] In the process of identifying logistics address information, it is necessary to use a preprocessing model to identify and process the address encoding of a handwritten logistics address or a logistics address input by a user, so as to obtain the address encoding of the logistics address, so as to obtain the logistics address information of the logistics order by using the address encoding. In the existing training process of the preprocessing model, the recognition accuracy of the preprocessing model for the address encoding is related to the coverage and accuracy of the address data used to train the preprocessing model.
[0004] However, since address data belongs to private data, such data cannot be collected and included on a large scale. This also results in poor training effects for the existing training of the preprocessing model, and both the accuracy and universality of the preprocessing model are relatively poor. Summary of the Invention
[0005] Embodiments of this application provide a method, apparatus, electronic device, and storage medium for training a global preprocessing model, which are used to train the preprocessing model to obtain a preprocessing model with high recognition universality and recognition accuracy.
[0006] In a first aspect, this application provides a method for training a global preprocessing model, including:
[0007] A first terminal obtains a first training data set, and uses the corpus training data in the first training data set to train an original preprocessing model to obtain a first local preprocessing model; wherein, the first training data set includes corpus training data from a first logistics address data source;
[0008] Obtain the model parameters of a second local preprocessing model sent by a second terminal; wherein, the second local preprocessing model is a model obtained by the second terminal using a second training data set to train the original preprocessing model, and the second training data set includes corpus training data from a second logistics address data source;
[0009] The first terminal performs parameter aggregation and weighting processing on the first local preprocessing model by using the model parameters of the second local preprocessing model to obtain a global preprocessing model.
[0010] In an alternative embodiment, the first terminal obtains a first training data set, including:
[0011] The first terminal obtains text corpus data of each logistics address from a first logistics address data source;
[0012] Perform data cleaning and feature engineering processing on the text corpus data to obtain corpus training data in the first training data set.
[0013] In an alternative embodiment, training the original preprocessing model using the corpus training data in the first training data set to obtain a first local preprocessing model includes:
[0014] Perform vectorization processing on the corpus training data in the first training data set to obtain character vectors;
[0015] Use the character vectors to train the original preprocessing model until the parameters in the original preprocessing model converge to obtain the first local preprocessing model.
[0016] In an alternative embodiment, performing parameter aggregation and weighting processing on the first local preprocessing model using the model parameters of the second local preprocessing model to obtain a global preprocessing model includes:
[0017] Perform decompression processing on the model parameters of the second local preprocessing model to obtain the neuron node distribution of the second local preprocessing model and the parameter values on each neuron node of the second local preprocessing model;
[0018] Determine the mapping relationship between neuron nodes between the neuron node distribution of the first local preprocessing model and the neuron node distribution of the second local preprocessing model;
[0019] Using the mapping relationship, perform aggregation and weighting processing on the parameter values on the corresponding neuron nodes in the first local preprocessing model according to the parameter values on each neuron node of the second local preprocessing model to obtain the global preprocessing model.
[0020] In an alternative embodiment, after obtaining the first local preprocessing model, the training method further includes:
[0021] The first terminal determines the neuron node distribution of the first local preprocessing model and the parameter values on each neuron node, and performs compression processing on the neuron node distribution and the parameter values on each neuron node to obtain the model parameters of the first local preprocessing model;
[0022] Send the model parameters of the first local preprocessing model to the second terminal, so that the second terminal performs parameter aggregation and weighting on the second local preprocessing model in the second terminal according to the model parameters of the first local preprocessing model to obtain the global preprocessing model.
[0023] In an alternative embodiment, the training method further includes:
[0024] The first terminal obtains a first incremental training dataset, and uses the first incremental training dataset to perform incremental training on the current global preprocessing model to obtain an incremented first global preprocessing model;
[0025] Wherein, the first incremental training dataset includes first corpus incremental training data, and the first corpus incremental training data is data obtained by performing data increment processing on the corpus training data in the first training dataset.
[0026] In an alternative embodiment, after obtaining the incremented first global preprocessing model, the training method further includes:
[0027] Determine the difference between the parameter values of the global preprocessing model before increment at each neuron node and the parameter values of the incremented first global preprocessing model at each neuron node;
[0028] Perform compression processing according to the difference between the parameter values at each neuron node and the neuron node distribution of the global preprocessing model after incremental training to obtain the model parameters of the incremented first global preprocessing model;
[0029] Send the model parameters of the first global preprocessing model after incremental training to the second terminal, so that the second terminal performs parameter aggregation and weighting on the current global preprocessing model in the second terminal according to the model parameters of the incremented first global preprocessing model to obtain the incremented first global preprocessing model.
[0030] In an alternative embodiment, after the first terminal performs parameter aggregation and weighting on the first local preprocessing model using the model parameters of the second local preprocessing model to obtain the global preprocessing model, the training method further includes:
[0031] Obtain the model parameters of the second globally pre - processed model after increment; wherein, the second globally pre - processed model after increment is a model obtained by the second terminal performing incremental training on the current globally pre - processed model in the second terminal using a second incremental training dataset, and the second incremental training dataset includes second corpus incremental training data, and the second corpus incremental training data is data obtained by performing data increment processing on the corpus training data in the second training dataset;
[0032] The first terminal performs parameter aggregation and weighting processing on the current globally pre - processed model using the model parameters of the second globally pre - processed model to obtain the second globally pre - processed model after increment.
[0033] In an alternative embodiment, when the first terminal performs parameter aggregation and weighting processing on the current globally pre - processed model using the model parameters of the second globally pre - processed model to obtain the second globally pre - processed model after increment, it further includes:
[0034] Perform decompression processing on the model parameters of the second globally pre - processed model to obtain the neuron node distribution of the second globally pre - processed model, the difference between the parameter values on each neuron node, and the neuron node distribution of the globally pre - processed model after incremental training;
[0035] Determine the mapping relationship between the neuron node distribution of the current globally pre - processed model and the neuron node distribution of the second globally pre - processed model;
[0036] Use the parameter values on each neuron node in the model parameters of the second globally pre - processed model to perform aggregation and weighting processing on the parameter values on the corresponding neuron nodes in the current globally pre - processed model to obtain the second globally pre - processed model after increment.
[0037] In an alternative embodiment, the training method further includes:
[0038] Store the globally pre - processed model obtained from this training in a model repository; wherein, multiple globally pre - processed models obtained from multiple trainings are stored in the model repository;
[0039] Call each model in the model repository to train a preset NLP task model to obtain a trained NLP task model.
[0040] In a second aspect, the present application provides a training device for a globally pre - processed model. The training device for the globally pre - processed model is applied to a first terminal, and the training device includes:
[0041] An acquisition module, configured to obtain a first training data set; wherein, the first training data set includes corpus training data from a first logistics address data source;
[0042] A training module, configured to train an original preprocessing model by using the corpus training data in the first training data set to obtain a first local preprocessing model;
[0043] The acquisition module is further configured to obtain model parameters of a second local preprocessing model sent by a second terminal; wherein, the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model by using a second training data set, and the second training data set includes corpus training data from a second logistics address data source;
[0044] An aggregation and weighting processing module, configured to perform parameter aggregation and weighting processing on the first local preprocessing model by using the model parameters of the second local preprocessing model to obtain a global preprocessing model.
[0045] In a third aspect, the present application provides an electronic device, including: at least one processor and a memory;
[0046] The memory stores computer-executable instructions;
[0047] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the method described in the first aspect.
[0048] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the method described in the first aspect is implemented.
[0049] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0050] The embodiments of the present application provide a training method, apparatus, electronic device, and storage medium for a global preprocessing model. By obtaining a first training data set and using the corpus training data in the first training data set to train an original preprocessing model, a first local preprocessing model is obtained; wherein, the first training data set includes corpus training data from a first logistics address data source; obtaining model parameters of a second local preprocessing model sent by a second terminal; wherein, the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model using a second training data set, and the second training data set includes corpus training data from a second logistics address data source; finally, using the model parameters of the second local preprocessing model to perform parameter aggregation and weighting processing on the first local preprocessing model to obtain a global preprocessing model. In this solution, a model training method of federated training is used to obtain a global training model with better accuracy and stronger universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0052] Figure 1 It is a schematic diagram of a network architecture on which the present application is based;
[0053] Figure 2 It is a schematic flowchart of a training method for a global preprocessing model provided by the present application;
[0054] Figure 3 It is a data flow schematic diagram of a training method for a global preprocessing model provided by an embodiment of the present application;
[0055] Figure 4 It is a schematic structural diagram of the mapping relationship between neuron nodes provided by an embodiment of the present application;
[0056] Figure 5 It is a data flow schematic diagram of another training method for a global preprocessing model provided by an embodiment of the present application;
[0057] Figure 6 It is a data flow schematic diagram of yet another training method for a global preprocessing model provided by an embodiment of the present application;
[0058] Figure 7 It is a schematic diagram of a network architecture based on a global preprocessing model provided by an embodiment of the present application;
[0059] Figure 8 It is a schematic structural diagram of a training apparatus for a global preprocessing model provided by an embodiment of the present application;
[0060] Figure 9 Schematic structural diagram of the electronic device provided by an embodiment of the present invention.
[0061] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be provided hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments
[0062] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of systems and methods consistent with some aspects of the present application as detailed in the appended claims.
[0063] With the rapid development of e-commerce business and the increasing popularity of online shopping, the business volume of the logistics business of the logistics platform has also increased sharply. In the process of creating a logistics order on the logistics platform, it is inseparable from the identification of the logistics address information of the recipient or the logistics address information of the sender.
[0064] In the process of identifying the logistics address information, first, the logistics platform will obtain the handwritten or input logistics address to be identified, and then the logistics platform will use the preprocessing model to perform encoding and recognition processing on these logistics addresses. Among them, in the process of the preprocessing model performing encoding and recognition processing on the logistics address, the preprocessing model will perform a series of processes on the logistics address to be identified, including vectorization processing, word segmentation processing, etc., to obtain the address code corresponding to the logistics address to be identified.
[0065] Subsequently, the logistics platform will use the multi-task NLP task model to perform a series of processes on the address code, including address error correction, address word segmentation, address serialization annotation, etc., and finally obtain the logistics address information that can be used to generate a logistics order.
[0066] In the above process, whether the logistics address information obtained by the multi-task NLP task model processing the address code is accurate will depend on the accuracy of the address code. And the accuracy of the address code will be related to the training quality of the preprocessing model.
[0067] In the training process of the existing preprocessing model, in order to enable the preprocessing model to recognize different address expression forms of the same logistics address as the same address code. The logistics platform needs to collect a large amount of address data with different precisions as the training data of the preprocessing model and train the preprocessing model.
[0068] However, since address data belongs to private data, a logistics platform can generally only obtain partial address data of the areas covered by its own business scope. Training the preprocessing model with such address data will lead to the problem of low accuracy when the preprocessing model encodes and recognizes address data in areas not covered by the logistics platform's business, that is, the existing preprocessing model has low universality in recognizing address codes. This will also affect the accuracy of the logistics address information output by the subsequent multi-task NLP task model.
[0069] To address the above problems, the solution based on this application uses the model training method of federated training, so that each logistics platform trains the same preprocessing model with training data based on different logistics address data sources to obtain local preprocessing models. Then, the model parameters of each local preprocessing model are exchanged and aggregated, so that each logistics platform can obtain the same global preprocessing model.
[0070] Among them, for the terminal of any logistics platform, it obtains the first training data set and trains the original preprocessing model with the corpus training data in the first training data set to obtain the first local preprocessing model; where the first training data set includes corpus training data from the first logistics address data source; obtains the model parameters of the second local preprocessing model sent by the second terminal; where the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model with the second training data set, and the second training data set includes corpus training data from the second logistics address data source; finally, uses the model parameters of the second local preprocessing model to perform parameter aggregation and weighting processing on the first local preprocessing model to obtain the global preprocessing model.
[0071] Through such a model training method of federated training, the finally obtained global preprocessing model is trained based on different data sources. Compared with the preprocessing model obtained in the prior art, the global preprocessing model obtained in this application can effectively recognize address data in a larger range and has better universality; at the same time, due to the broadening of the data source, the global preprocessing model has better encoding recognition effect and higher accuracy for special addresses, especially remote address data.
[0072] The method provided by this application will be described below in combination with different implementation manners.
[0073] Refer to Figure 1 , Figure 1 which is a schematic diagram of a network architecture based on this application. The Figure 1 shown network architecture may specifically include a first terminal 1 and a second terminal 2.
[0074] Among them, the first terminal 1 and the second terminal 2 can specifically be hardware devices that can be used to execute data acquisition, processing, and interaction, including but not limited to desktop computers, tablet computers, and cloud computing platforms.
[0075] Through the network, the first terminal 1 and the second terminal 2 can perform encrypted exchange of model parameters, and respectively perform aggregation processing of the model parameters according to the exchanged model parameters to obtain a global preprocessing model.
[0076] It can be known that in the training method of the global preprocessing model provided in this application, the first terminal and the second terminal will adopt the same training method to train and process the same original preprocessing model.
[0077] For the convenience of description, this application will take the first terminal as the execution subject and combine the accompanying drawings to describe the training method of the global preprocessing model in this case. For the second terminal, the method for implementing the training of the global preprocessing model is similar to the training method used in the first terminal, and this application will not elaborate on it.
[0078] Embodiment 1
[0079] Figure 2 is a schematic flowchart of a training method for a global preprocessing model provided by this application. As Figure 2 shown, the execution subject of this method is the aforementioned first terminal, and this method includes:
[0080] Step 201, obtain a first training data set; among them, the first training data set includes corpus training data from a first logistics address data source.
[0081] Step 202, use the corpus training data in the first training data set to train the original preprocessing model to obtain a first local preprocessing model.
[0082] Step 203, obtain the model parameters of the second local preprocessing model sent by the second terminal; among them, the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model using a second training data set, and the second training data set includes corpus training data from a second logistics address data source.
[0083] Step 204, use the model parameters of the second local preprocessing model to perform parameter aggregation and weighting processing on the first local preprocessing model to obtain a global preprocessing model.
[0084] The training method for the global preprocessing model provided by this application can specifically be applied to a training device for the global preprocessing model. In this embodiment, it will be described by taking the training device for the global preprocessing model integrated in the first terminal as an example.
[0085] It should be noted that in the implementation mode of the present application, the first terminal and the second terminal can specifically be processing terminals from different logistics platforms, and the same original preprocessing model will be preset in the first terminal and the second terminal. The original preprocessing model is specifically based on algorithms such as word2vector algorithm or bert algorithm. This model can convert the corpus data of the logistics address into a word vector representation, and obtain the address encoding of the logistics address based on the word vector representation.
[0086] Since the relevant data of the logistics address in the logistics address data source all involve personal privacy, that is, these data cannot be exchanged between different terminals of different logistics platforms. Based on this, the implementation mode of the present application will adopt the method of federated training to overcome this problem. Specifically, in the implementation mode of the present application, the first terminal and the second terminal will respectively accumulate their own logistics address data sources, and generate corpus training data for training their respective original preprocessing models according to their respective logistics address data sources. It is precisely by using the predicted training data obtained from different logistics address data sources to train the same original preprocessing model that the model parameters in the locally trained preprocessing models obtained by training will be different.
[0087] At this time, the first terminal trains to obtain a first locally preprocessed model, and the second terminal trains to obtain a second locally preprocessed model. Subsequently, the first terminal will receive the model parameters from the second terminal, that is, the model parameters of the second locally preprocessed model, and perform parameter aggregation and weighting processing on the model parameters of its own existing first locally preprocessed model according to the received model parameters of the second locally preprocessed model, so as to obtain a processed model, that is, a globally preprocessed model.
[0088] That is to say, the first terminal and the second terminal respectively train the same original preprocessing model by using their own unique predicted training data, then exchange the model parameters of the locally trained preprocessing models, and finally, use the exchanged model parameters to perform aggregation and weighting on their own existing locally preprocessed models to obtain the final globally preprocessed model.
[0089] Compared with the method of training the preprocessing model using a single logistics address data source in the prior art, the implementation of the present application can, without the first terminal and the second terminal sharing the logistics address data source, obtain a globally preprocessed model through the method of federated training that can accurately identify the address encoding of the address data in the first logistics address data source and can also accurately identify the address encoding of the address data in the second logistics address data source. This model has stronger universality and higher accuracy.
[0090] Next, the implementation mode of the present application will be further described in combination with an actual scenario. Figure 3Data flow diagram of a training method for a global preprocessing model provided by an embodiment of this application, as Figure 3 shown, in the framework shown in this figure, it includes logistics platform A and logistics platform B.
[0091] Among them, the first terminal in this application will be set in logistics platform A, and the second terminal will be set in logistics platform B.
[0092] As mentioned before, since different logistics platforms need to accumulate their respective logistics address data to obtain their respective logistics address data sources.
[0093] Suppose the business scope of logistics platform A is mainly oriented to City A, and its first terminal will obtain text corpus data of each logistics address including those from City A (the first logistics address data source). And the business scope of logistics platform B is mainly oriented to City B, and its second terminal will obtain text corpus data of each logistics address including those from City B (the second logistics address data source).
[0094] After the first terminal obtains the text corpus data of each logistics address of the first logistics address data source, it will perform data cleaning and feature engineering processing on these text corpus data to obtain corpus training data for the first training dataset that can be used to train the model.
[0095] Similarly, after the second terminal obtains the text corpus data of each logistics address of the second logistics address data source, it will perform data cleaning and feature engineering processing on these text corpus data to obtain corpus training data for the second training dataset that can be used to train the model.
[0096] And by performing data cleaning and feature engineering processing on the text corpus data, the obtained corpus training data can be made more standardized and effective to ensure the training quality of the model. In particular, since the preprocessing model based on this application is a model for identifying address codes, it can adopt an unlabeled training method during training, that is, in the process of obtaining the training dataset, the corpus training data does not need to be labeled to obtain the training dataset.
[0097] Subsequently, as Figure 3 shown, the same original preprocessing model will be pre-installed in the first terminal of logistics platform A and the second terminal of logistics platform B respectively.
[0098] At this time, the first terminal will use the aforementioned corpus training data of City A to train the original preprocessing model to obtain a first local preprocessing model that can be used to identify the logistics addresses in City A and convert them into address codes.
[0099] In an alternative implementation, the training of the original preprocessing model by the first terminal may include: vectorizing the corpus training data in the first training dataset to obtain character vectors; and training the original preprocessing model using the character vectors until the parameters in the original preprocessing model converge, thereby obtaining the first local preprocessing model.
[0100] Meanwhile, the second terminal will use the aforementioned corpus training data of City B to train the original preprocessing model, obtaining a second local preprocessing model that can be used to identify the logistics addresses in City B and convert them into address codes. The training process is similar to that of the first terminal and will not be elaborated herein.
[0101] After the first terminal and the second terminal have respectively completed the training of the original preprocessing model to obtain local preprocessing models, as Figure 3 shown, the two of them will also encrypt and exchange the model parameters to further process the local preprocessing models using the obtained model parameters.
[0102] Specifically, as shown in Figure 3 , taking the first terminal as an example, the first terminal will first determine the neuron node distribution of the first local preprocessing model and the parameter values on each neuron node, and perform compression processing on the neuron node distribution and the parameter values on each neuron node to obtain the model parameters of the first local preprocessing model; then, the first terminal will send the model parameters of the first local preprocessing model to the second terminal. At this time, the second terminal will perform parameter aggregation and weighting processing on the second local preprocessing model in the second terminal using the model parameters of the first local preprocessing model to obtain the global preprocessing model.
[0103] Among them, when the first terminal performs compression processing on the parameter values on each neuron node to obtain model parameters, the following method can be used, that is, through dimensionality reduction algorithms such as SVD, the current parameter matrix of the model is decomposed into multiple low-rank matrices, which are then transmitted to the second terminal. After the second terminal obtains these low-rank matrices, it performs matrix reconstruction to restore to the original parameter matrix, and then performs parameter weighted aggregation and weighting processing to obtain the global preprocessing model.
[0104] While the first terminal sends the model parameters of the first local preprocessing model, the first terminal also receives the model parameters of the second local preprocessing model sent from the second terminal. Since the model parameters are obtained after the second terminal compresses the neuron node distribution of the second local preprocessing model and the parameter values on each neuron node, the first terminal will first decompress the model parameters of the second local preprocessing model to obtain the neuron node distribution of the second local preprocessing model and the parameter values of the second local preprocessing model on each neuron node. Then, the first terminal will determine the mapping relationship between the neuron nodes of the neuron node distribution of the first local preprocessing model and the neuron node distribution of the second local preprocessing model. Finally, the first terminal will use the mapping relationship to perform aggregation and weighting processing on the parameter values of the corresponding neuron nodes in the first local preprocessing model according to the parameter values of the second local preprocessing model on each neuron node to obtain the global preprocessing model.
[0105] Specifically, in the step where the first terminal determines the mapping relationship between the neuron nodes of the neuron node distribution of the first local preprocessing model and the neuron node distribution of the second local preprocessing model, it is implemented by using the feature that the neuron parameter matrix has permutation invariance.
[0106] Figure 4 This is a schematic structural diagram of the mapping relationship between neuron nodes provided by the embodiments of the present application. After two preprocessing models to be trained are trained with different inputs, the order of neuron nodes in each hidden layer of the preprocessing models will change (as Figure 4 shown), which will cause the order of elements in the matrix expressions of the model parameters in the first local preprocessing model and the model parameters in the second local preprocessing model to change. However, no matter how the order of its elements changes, its final output result is not affected.
[0107] Based on this, after the first terminal obtains the neuron node distribution of the second local preprocessing model, the following method will be used to match it with the neuron node distribution of the first local preprocessing model.
[0108] The global hidden layer obtained by initialization or the previous iteration is respectively matched with the local hidden layers of each party. Assuming that the global hidden layer (the first local preprocessing model) of the central node is matched with a local hidden layer (the second local preprocessing model) of a certain platform, the objective function is as follows:
[0109]
[0110] where w jlrepresents the l-th neuron of the j-th platform hidden layer; θ i represents the i-th neuron of the global hidden layer; c(.,.) represents the similarity function of a pair of neurons; represents the similarity weight value between the l-th neuron of the local hidden layer of the j-th platform and the i-th neuron of the global hidden layer.
[0111] Subsequently, it is solved through the Hungarian matching algorithm to obtain the weight matrix of the local hidden layer (the second local preprocessing model) of this platform relative to the global hidden layer (the first local preprocessing model), until the local hidden layers of all platforms are matched, forming a neuron mapping between the two.
[0112] After completing the neuron node matching of the first local preprocessing model and the second local preprocessing model, the first local preprocessing model can be multiplied by the corresponding weight matrix and weighted aggregated to obtain a new global hidden layer, that is, the global preprocessing model.
[0113] At this point, the first terminal will obtain the global preprocessing model. Since this global preprocessing model is based on the first local preprocessing model that can be used to identify the logistics addresses in City A and convert them into address codes, and the model parameters are aggregated and weighted to aggregate and weight the model parameters of the second local preprocessing model that can be used to identify the logistics addresses in City B and convert them into address codes into the model parameters of the first local preprocessing model, the obtained global preprocessing model can also, like the second local preprocessing model, identify and convert the logistics addresses in City B into address codes.
[0114] Similar to the first terminal obtaining the global preprocessing model, the second terminal will also first train to obtain the second local preprocessing model, and then use the model parameters of the first local preprocessing model to aggregate and weight the model parameters of the second local preprocessing model to obtain the same global preprocessing model as the global preprocessing model obtained by the first terminal. Similarly, for the second terminal, using this global preprocessing model can not only, like the second local preprocessing model, identify and convert the logistics addresses in City B into address codes, but also has the ability to identify and convert the logistics addresses in City A into address codes.
[0115] Based on this, as Figure 3 shown, the first terminal of logistics platform A and the second terminal of logistics platform B will obtain the same global preprocessing model. Subsequently, using this global preprocessing model can be used to identify the to-be-processed logistics address data to obtain the address code corresponding to the logistics address, for the subsequent NLP task models of different tasks to use the address code for training.
[0116] To ensure the timeliness of the preprocessing model, the logistics platform will change the data in the logistics address data source at regular intervals to ensure the authenticity and effectiveness of the data. At this time, it is necessary to retrain the preprocessing model using the updated data to ensure that the global preprocessing model can recognize the updated data. However, if the model is retrained in the aforementioned manner, it will result in a large resource overhead.
[0117] Based on this, on the basis of the above implementation, this implementation also provides a method for incrementally training the obtained global preprocessing model, so that the incremented global preprocessing model can effectively recognize and process the updated incremental data.
[0118] Figure 5 It is a data flow diagram of another training method of the global preprocessing model provided by the embodiment of the present application. As Figure 5 shown, the first terminal of logistics platform A and the second terminal of logistics platform B have respectively completed the federated training of the original preprocessing model to obtain the same global preprocessing model.
[0119] At this time, if the first terminal of logistics platform A performs data increment processing, that is, the first terminal will perform data increment processing on the corpus training data in the first training data set to obtain the first corpus increment training data.
[0120] The first terminal will, as Figure 5 shown, after obtaining the first corpus increment training data, use the first corpus increment training data to train the existing global preprocessing model to obtain the incremented first global preprocessing model.
[0121] The incremented first global preprocessing model obtained at this time is the global preprocessing model that can be used to recognize the incremented data.
[0122] Subsequently, the first terminal will also send the model parameters of the incremented first global preprocessing model to the second terminal again.
[0123] Different from the previous sending of the model parameters of the local preprocessing model to the second terminal, since the model parameters of the incremented first global preprocessing model are sent this time, some of the model parameters may be the same as the model parameters of the previously obtained global preprocessing model. Based on this, in order to improve the sending speed of the model parameters, the following process can be adopted during the sending of the model parameters this time:
[0124] The first terminal determines the difference between the parameter values of the global preprocessing model before the increment at each neuron node and the parameter values of the first global preprocessing model after the increment at each neuron node. Then, according to the difference between the parameter values at each neuron node and the neuron node distribution of the global preprocessing model after incremental training, compression processing is performed to obtain the model parameters of the first global preprocessing model after the increment. Finally, the model parameters of the first global preprocessing model after incremental training are sent to the second terminal.
[0125] As Figure 5 shown, the second terminal will perform parameter aggregation and weighting processing on the current global preprocessing model in the second terminal according to the model parameters of the first global preprocessing model after the increment to obtain the first global preprocessing model after the increment. At this time, in the first terminal of logistics platform A and the second terminal of logistics platform B, there is the first global preprocessing model after the increment, and this first global preprocessing model after the increment can be used to identify the first corpus incremental training data included in the first incremental dataset to output the address code corresponding to the first corpus incremental training data.
[0126] Similar to the Figure 5 process is that Figure 6 is a data flow diagram of another training method of the global preprocessing model provided by the embodiment of the present application. As Figure 6 shown, the first terminal of logistics platform A and the second terminal of logistics platform B have respectively completed the federated training of the original preprocessing model to obtain the same global preprocessing model.
[0127] At this time, if the second terminal of logistics platform B performs data increment processing, that is, the second terminal will perform data increment processing on the corpus training data in the second training dataset to obtain the second corpus incremental training data. Then, the second terminal will use the second corpus incremental training data to perform incremental training on the current global preprocessing model in the second terminal to obtain the second global preprocessing model after the increment. Correspondingly, the second terminal will also encrypt and send the model parameters of the second global preprocessing model after the increment to the first terminal for the first terminal to use.
[0128] As Figure 6 shown, for the first terminal, after obtaining the model parameters of the second global preprocessing model after the increment sent by the second terminal, the first terminal performs parameter aggregation and weighting processing on the current global preprocessing model using the model parameters of the second global preprocessing model to obtain the second global preprocessing model after the increment.
[0129] Among them, when the first terminal uses the model parameters of the second global preprocessing model to perform parameter aggregation and weighting on the current global preprocessing model to obtain the incremented second global preprocessing model, the first terminal first decompresses the model parameters of the second global preprocessing model to obtain the neuron node distribution of the second global preprocessing model, the difference between the parameter values on each neuron node, and the neuron node distribution of the globally preprocessed model after incremental training; then, the first terminal determines the mapping relationship between the neuron node distribution of the current global preprocessing model and the neuron node distribution of the second global preprocessing model; finally, the parameter values on the corresponding neuron nodes of the current global preprocessing model are aggregated and weighted using the parameter values on each neuron node in the model parameters of the second global preprocessing model to obtain the incremented second global preprocessing model.
[0130] At this time, as Figure 6 shown, both the first terminal and the second terminal obtain the incremented second global preprocessing model, and the second global preprocessing model can be used to identify the second corpus incremental training data and output the corresponding address encoding.
[0131] In summary, Figure 5 shows the relevant processing flow for the first terminal to retrain the global preprocessing model to obtain the incremented first global preprocessing model when the data of the first logistics address data source at the first terminal increases. Figure 6 shows the relevant processing flow for the first terminal to retrain the global preprocessing model to obtain the incremented second global preprocessing model when the data of the second logistics address data source at the second terminal increases.
[0132] Regardless of which logistics address data source sends data increments, the first terminal will retrain its current global preprocessing model so that the obtained incremented model can be applied to the identification and processing of incremented data. Among them, in order to improve the training efficiency of the incremented model, when encrypting and transmitting the model parameters of the incremented model, only the change amount of the model parameters can be transmitted to improve the transmission efficiency of the model parameters.
[0133] In order to facilitate the subsequent use of the global preprocessing model by the NLP task model. On the basis of the above embodiments, the terminal will also store the globally preprocessed model obtained from this training in the model repository; among them, multiple globally preprocessed models obtained from multiple trainings are stored in the model repository; each model in the model repository is called to train the preset NLP task model to obtain the trained NLP task model.
[0134] Figure 7A schematic diagram of a network architecture provided by an embodiment of the present application, as follows Figure 7 As shown, this network architecture uses the aforementioned global preprocessing model as a pre-model for the NLP task model to output address encodings that can be used to train this model for the NLP task model.
[0135] Specifically, first, the first terminal of logistics platform A can use the aforementioned method to train the original preprocessing model to obtain a global preprocessing model, which will be stored in the model repository; at the same time, whenever the global preprocessing model undergoes one training based on incremental data to obtain an incremented first global preprocessing model or an incremented second global preprocessing model, these incremented models will also be stored in the model repository. That is, multiple global preprocessing models obtained through multiple trainings are stored in the model repository.
[0136] Then, the first terminal of logistics platform A will sequentially call out each model from the model repository and input the training data (such as Figure 7 the labeled address data shown) for training the NLP task model into each model (such as Figure 7 the global preprocessing model 1, global preprocessing model 2... global preprocessing model N shown) to obtain multiple address encodings (such as Figure 7 address encoding 1, address encoding 2... address encoding N shown). Then, the obtained multiple address encodings are used to train the NLP task model respectively to obtain a trained NLP task model.
[0137] Of course, during the process of using multiple address encodings to train the NLP task model respectively, the federated training method can also be adopted to further improve the training quality of the NLP task model. For its specific implementation process, this embodiment will not be elaborated further.
[0138] The embodiment of the present application provides a training method for a global preprocessing model. By obtaining a first training data set and using the corpus training data in the first training data set to train an original preprocessing model, a first local preprocessing model is obtained. Among them, the first training data set includes corpus training data from a first logistics address data source; obtaining the model parameters of a second local preprocessing model sent by a second terminal. Among them, the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model using a second training data set, and the second training data set includes corpus training data from a second logistics address data source; finally, using the model parameters of the second local preprocessing model to perform parameter aggregation and weighting processing on the first local preprocessing model to obtain a global preprocessing model. In this solution, a model training method of federated training is used to obtain a global training model with better accuracy and stronger universality.
[0139] Embodiment 2
[0140] Corresponding to the training method of the global preprocessing model in the above embodiment, Figure 8 It is a schematic structural diagram of a training device for a global preprocessing model provided by an embodiment of the present application, as Figure 8 shown. The training device for the global preprocessing model can be specifically applied to a first terminal, and it includes: an acquisition module 10, a training module 20, and an aggregation and weighting processing module 30.
[0141] Among them, the acquisition module 10 is used to obtain a first training data set. Among them, the first training data set includes corpus training data from a first logistics address data source;
[0142] The training module 20 is used to use the corpus training data in the first training data set to train an original preprocessing model to obtain a first local preprocessing model;
[0143] The acquisition module 10 is further used to obtain the model parameters of a second local preprocessing model sent by a second terminal. Among them, the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model using a second training data set, and the second training data set includes corpus training data from a second logistics address data source;
[0144] The aggregation and weighting processing module 30 is used to use the model parameters of the second local preprocessing model to perform parameter aggregation and weighting processing on the first local preprocessing model to obtain a global preprocessing model.
[0145] In an alternative embodiment, the acquisition module 10 is specifically used for:
[0146] Obtain the text corpus data of each logistics address from the first logistics address data source; perform data cleaning and feature engineering processing on the text corpus data to obtain the corpus training data in the first training dataset.
[0147] In an alternative embodiment, the training module 20 is specifically configured to:
[0148] Perform vectorization processing on the corpus training data in the first training dataset to obtain character vectors; use the character vectors to train the original preprocessing model until the parameters in the original preprocessing model converge, to obtain the first local preprocessing model.
[0149] In an alternative embodiment, the aggregation and weighting processing module 30 is specifically configured to:
[0150] Perform decompression processing on the model parameters of the second local preprocessing model to obtain the neuron node distribution of the second local preprocessing model and the parameter values on each neuron node of the second local preprocessing model; determine the mapping relationship between the neuron nodes between the neuron node distribution of the first local preprocessing model and the neuron node distribution of the second local preprocessing model; use the mapping relationship to perform aggregation and weighting processing on the parameter values on the corresponding neuron nodes in the first local preprocessing model according to the parameter values on each neuron node of the second local preprocessing model, to obtain the global preprocessing model.
[0151] In an alternative embodiment, the aggregation and weighting processing module 30 is further configured to:
[0152] Determine the neuron node distribution of the first local preprocessing model and the parameter values on each neuron node, and perform compression processing on the neuron node distribution and the parameter values on each neuron node to obtain the model parameters of the first local preprocessing model;
[0153] The obtaining module is further configured to send the model parameters of the first local preprocessing model to the second terminal, for the second terminal to perform parameter aggregation and weighting processing on the second local preprocessing model in the second terminal according to the model parameters of the first local preprocessing model, to obtain the global preprocessing model.
[0154] In an alternative embodiment, the obtaining module 10 is further configured to:
[0155] Obtain a first incremental training dataset;
[0156] The training module 20 is further configured to: perform incremental training on the current global preprocessing model by using the first incremental training dataset to obtain an incremented first global preprocessing model; wherein, the first incremental training dataset includes first corpus incremental training data, and the first corpus incremental training data is data obtained by performing data increment processing on the corpus training data in the first training dataset.
[0157] In an alternative embodiment, the aggregation and weighting processing module 30 is further configured to:
[0158] Determine the difference between the parameter values of the global preprocessing model before increment at each neuron node and the parameter values of the incremented first global preprocessing model at each neuron node; perform compression processing based on the difference between the parameter values at each neuron node and the neuron node distribution of the global preprocessing model after incremental training to obtain the model parameters of the incremented first global preprocessing model.
[0159] The obtaining module 10 is further configured to: send the model parameters of the incremented first global preprocessing model to the second terminal for the second terminal to perform parameter aggregation and weighting processing on the current global preprocessing model in the second terminal according to the model parameters of the incremented first global preprocessing model to obtain an incremented first global preprocessing model.
[0160] In an alternative embodiment, the obtaining module 10 is further configured to: obtain the model parameters of the incremented second global preprocessing model sent by the second terminal; wherein, the incremented second global preprocessing model is a model obtained by the second terminal performing incremental training on the current global preprocessing model in the second terminal by using a second incremental training dataset, and the second incremental training dataset includes second corpus incremental training data, and the second corpus incremental training data is data obtained by performing data increment processing on the corpus training data in the second training dataset;
[0161] The aggregation and weighting processing module 30 is further configured to: perform parameter aggregation and weighting processing on the current global preprocessing model by using the model parameters of the second global preprocessing model to obtain an incremented second global preprocessing model.
[0162] In an alternative embodiment, the aggregation and weighting processing module 30 is further configured to:
[0163] Perform decompression processing on the model parameters of the second global preprocessing model to obtain the neuron node distribution of the second global preprocessing model, the difference between the parameter values at each neuron node, and the neuron node distribution of the global preprocessing model after incremental training;
[0164] Determine the mapping relationship between the neuron node distribution of the current global preprocessing model and the neuron node distribution of the second global preprocessing model;
[0165] Use the parameter values on each neuron node in the model parameters of the second global preprocessing model to perform aggregation and weighting processing on the parameter values on the corresponding neuron nodes in the current global preprocessing model, and obtain the second global preprocessing model after increment.
[0166] In an alternative embodiment, the training device for the global preprocessing model further includes: a downstream training module;
[0167] The downstream training module is configured to: store the globally preprocessed model obtained in this training in a model repository; wherein, multiple globally preprocessed models obtained from multiple trainings are stored in the model repository; call each model in the model repository to train a preset NLP task model, and obtain a trained NLP task model.
[0168] The embodiment of the present application provides a training device for a global preprocessing model. By obtaining a first training data set and using the corpus training data in the first training data set to train an original preprocessing model, a first local preprocessing model is obtained; wherein, the first training data set includes corpus training data from a first logistics address data source; obtain the model parameters of a second local preprocessing model sent by a second terminal; wherein, the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model using a second training data set, and the second training data set includes corpus training data from a second logistics address data source; finally, use the model parameters of the second local preprocessing model to perform parameter aggregation and weighting processing on the first local preprocessing model to obtain a global preprocessing model. In this solution, a model training method of federated training is used to obtain a globally trained model with better accuracy and stronger universality.
[0169] Embodiment III
[0170] Figure 9 For the structural schematic diagram of the electronic device provided by the embodiment of the present invention, as Figure 9 shown, the embodiment of the present invention further provides an electronic device 1400, including: a memory 1401, a processor 1402, and a computer program.
[0171] Wherein, the computer program is stored in the memory 1401 and is configured to be executed by the processor 1402 to implement the training method of the global preprocessing model provided by any embodiment of the present invention. Relevant descriptions can be understood by referring to the relevant descriptions and effects corresponding to the steps in the accompanying drawings, and will not be elaborated here.
[0172] Among them, in this embodiment, the memory 1401 and the processor 1402 are connected through a bus.
[0173] Embodiment 4
[0174] The embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the training method of the global preprocessing model provided in any embodiment of the present invention.
[0175] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems or modules can be in electrical, mechanical or other forms.
[0176] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0177] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware, or in the form of hardware plus software functional modules.
[0178] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer or other programmable question-answering systems, so that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0179] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0180] In addition, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the training method of the global preprocessing model described above.
[0181] Moreover, although the operations are depicted in a particular order, this should be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.
[0182] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A training method for a global preprocessing model, characterized in that, it includes: The first terminal obtains a first training data set, and uses the corpus training data in the first training data set to train an original preprocessing model to obtain a first local preprocessing model; wherein, the first training data set includes corpus training data from a first logistics address data source; Obtain the model parameters of the second local preprocessing model sent by the second terminal; wherein, the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model using a second training data set, and the second training data set includes corpus training data from a second logistics address data source; The first terminal performs parameter aggregation and weighting processing on the first local preprocessing model using the model parameters of the second local preprocessing model to obtain a global preprocessing model; The training method further includes: The first terminal obtains a first incremental training data set, and uses the first incremental training data set to perform incremental training on the current global preprocessing model to obtain an incremented first global preprocessing model; wherein, the first incremental training data set includes first corpus incremental training data, and the first corpus incremental training data is data obtained by performing data increment processing on the corpus training data in the first training data set; Determine the difference between the parameter values of the global preprocessing model before the increment on each neuron node and the parameter values of the incremented first global preprocessing model on each neuron node; Perform compression processing according to the difference between the parameter values on each neuron node and the neuron node distribution of the global preprocessing model after the incremental training to obtain the model parameters of the incremented first global preprocessing model; Send the model parameters of the first global preprocessing model after the incremental training to the second terminal for the second terminal to perform parameter aggregation and weighting processing on the current global preprocessing model in the second terminal according to the model parameters of the incremented first global preprocessing model to obtain the incremented first global preprocessing model.
2. The training method according to claim 1, characterized in that, The first terminal obtains a first training data set, including: The first terminal obtains the text corpus data of each logistics address from a first logistics address data source; Perform data cleaning and feature engineering processing on the text corpus data to obtain the corpus training data in the first training data set.
3. The training method according to claim 1, characterized in that, The using the corpus training data in the first training data set to train the original preprocessing model to obtain a first local preprocessing model includes: Perform vectorization processing on the corpus training data in the first training data set to obtain character vectors; Use the character vectors to train the original preprocessing model until the parameters in the original preprocessing model converge to obtain the first local preprocessing model.
4. The training method according to claim 1, characterized in that, Performing parameter aggregation and weighting on the first local preprocessing model by using the model parameters of the second local preprocessing model to obtain a global preprocessing model, including: Performing decompression processing on the model parameters of the second local preprocessing model to obtain the neuron node distribution of the second local preprocessing model and the parameter values of the second local preprocessing model on each neuron node; Determining the mapping relationship between neuron nodes between the neuron node distribution of the first local preprocessing model and the neuron node distribution of the second local preprocessing model; Using the mapping relationship, performing aggregation and weighting on the parameter values of the corresponding neuron nodes in the first local preprocessing model according to the parameter values of the second local preprocessing model on each neuron node to obtain the global preprocessing model.
5. The training method according to claim 1, wherein, after obtaining the first local preprocessing model, the training method further includes: The first terminal determines the neuron node distribution of the first local preprocessing model and the parameter values on each neuron node, and compresses the neuron node distribution and the parameter values on each neuron node to obtain the model parameters of the first local preprocessing model; Sending the model parameters of the first local preprocessing model to the second terminal for the second terminal to perform parameter aggregation and weighting on the second local preprocessing model in the second terminal by using the model parameters of the first local preprocessing model to obtain the global preprocessing model.
6. The training method according to claim 1, wherein, after the first terminal performs parameter aggregation and weighting on the first local preprocessing model by using the model parameters of the second local preprocessing model to obtain a global preprocessing model, the training method further includes: Obtaining the model parameters of the second global preprocessing model after increment; wherein, the second global preprocessing model after increment is a model obtained by the second terminal performing incremental training on the current global preprocessing model in the second terminal by using a second incremental training data set, wherein the second incremental training data set includes second corpus incremental training data, and the second corpus incremental training data is data obtained by performing data increment processing on the corpus training data in the second training data set; The first terminal performs parameter aggregation and weighting on the current global preprocessing model by using the model parameters of the second global preprocessing model to obtain the second global preprocessing model after increment.
7. The training method according to claim 6, wherein, the first terminal performing parameter aggregation and weighting on the current global preprocessing model by using the model parameters of the second global preprocessing model to obtain the second global preprocessing model after increment further includes: Decompress the model parameters of the second global preprocessing model to obtain the neuron node distribution of the second global preprocessing model, the difference between the parameter values on each neuron node, and the neuron node distribution of the global preprocessing model after incremental training; Determine the mapping relationship between the neuron node distribution of the current global preprocessing model and the neuron node distribution of the second global preprocessing model; Use the parameter values on each neuron node in the model parameters of the second global preprocessing model to perform aggregation and weighting processing on the parameter values on the corresponding neuron nodes in the current global preprocessing model to obtain the second global preprocessing model after increment; 8. The training method according to any one of claims 1-7, characterized in that, further comprising: Store the globally preprocessed model obtained from this training in the model repository; wherein, multiple globally preprocessed models obtained from multiple trainings are stored in the model repository; Call each model in the model repository to train a preset NLP task model to obtain a trained NLP task model.
9. A training device for a global preprocessing model, characterized in that, The training device for the global preprocessing model is applied to a first terminal, and the training device includes: An acquisition module, configured to obtain a first training data set; wherein, the first training data set includes corpus training data from a first logistics address data source; A training module, configured to use the corpus training data in the first training data set to train an original preprocessing model to obtain a first local preprocessing model; The acquisition module is further configured to obtain the model parameters of the second local preprocessing model sent by a second terminal; wherein, the second local preprocessing model is a model obtained by the second terminal training the original preprocessing model using a second training data set, and the second training data set includes corpus training data from a second logistics address data source; An aggregation and weighting processing module, configured to perform parameter aggregation and weighting processing on the first local preprocessing model using the model parameters of the second local preprocessing model to obtain a global preprocessing model; The acquisition module is further configured to: obtain a first incremental training data set; The training module is further configured to: use the first incremental training data set to perform incremental training on the current global preprocessing model to obtain a first globally preprocessed model after increment; wherein, the first incremental training data set includes first corpus incremental training data, and the first corpus incremental training data is data obtained by performing data increment processing on the corpus training data in the first training data set; The aggregation and weighting processing module is further configured to: Determine the difference between the parameter values on each neuron node of the global preprocessing model before increment and the parameter values on each neuron node of the first globally preprocessed model after increment; perform compression processing according to the difference between the parameter values on each neuron node and the neuron node distribution of the global preprocessing model after incremental training to obtain the model parameters of the first globally preprocessed model after increment; The obtaining module is further configured to: send the model parameters of the first globally preprocessed model after the incremental training to the second terminal, so that the second terminal performs parameter aggregation and weighting processing on the current globally preprocessed model in the second terminal according to the model parameters of the first globally preprocessed model after the increment, to obtain the first globally preprocessed model after the increment.
10. An electronic device, characterized in that it includes: at least one processor and a memory; the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the method according to any one of claims 1-8 is implemented.
12. A computer program product, including a computer program, characterized in that when the computer program is executed by the processor, the method according to any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Joint model training method, system and device and computer readable storage medium
CN109871702A
Multi-terminal model compression method and device based on knowledge federation, task prediction method and device and electronic equipment
CN112052938A