Information recommendation method, system and equipment and storage medium
By applying reinforcement learning model and blockchain technology in the charging pile recommendation system, combining the user's real-time status and environmental parameters, the problem of failure to fully consider the user's personal needs in the existing technology is solved, and efficient and real-time charging pile recommendations are achieved.
Patent Information
- Application Number
- CN202311471539.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-06
AI Technical Summary
The existing charging pile recommendation service fails to fully consider the user's personal needs, resulting in the failure to optimize the recommendation results.
By applying a reinforcement learning model in the first information recommendation node, combining the reinforcement learning model parameter update in the blockchain node, the environmental parameters and current state parameters of the target energy equipment are obtained, and prediction processing is performed to recommend a suitable energy supply device.
The method of recommending charging piles based on the user's real-time status is realized, fully considering the user's personal needs, ensuring the real-timeness of the prediction model and the optimization of the recommended results.
Smart Images

Figure CN119939006A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computing service technology, and in particular to an information recommendation method, system, device and storage medium. Background Art
[0002] With the rapid development of Internet technology, its application is becoming more and more extensive. Internet technology is not only limited to work, but also has a significant impact on people's life and study. For example, in life, with the rapid development of new energy vehicles, the number of public charging piles in cities continues to grow, and charging facility operators are also constantly expanding. How to reliably recommend charging facilities to users has become an urgent problem to be solved. The commonly used technical solution is to recommend charging equipment through real-time data provided by the charging facility operators themselves, that is, to recommend charging facilities to users through simple rules.
[0003] However, the above recommendation scheme only recommends charging facilities to users through simple rules, and does not fully consider the individual needs of users, resulting in non-optimal recommendation results.
[0004] Application Contents
[0005] In order to solve the above technical problems, the present application hopes to provide an information recommendation method, system, device and storage medium, which solves the problem that the current charging pile recommendation service does not fully consider the user's individual needs, and proposes a method for recommending charging piles based on the user's real-time status, which fully considers the user's individual needs and ensures that the obtained prediction model has strong real-time performance.
[0006] The technical solution of this application is implemented as follows:
[0007] In a first aspect, an information recommendation method is provided, the method being applied to a first information recommendation node, the method comprising:
[0008] If it is determined that the target energy device has an energy replenishment demand, obtain the environmental parameters of the area where the target energy device is located and the current state parameters of the target energy device; wherein the environmental parameters include the distribution information of energy providing devices that provide energy in the area where the target energy device is located;
[0009] Determine a trained reinforcement learning model; wherein the reinforcement learning model is obtained by updating the parameters of the reinforcement learning model stored in the blockchain node;
[0010] The reinforcement learning model performs prediction processing according to the environmental parameters and the current state parameters to obtain a recommendation result of an energy providing device that provides energy replenishment for the target energy equipment.
[0011] In a second aspect, an information recommendation system is provided, the system comprising: a first information recommendation node and a blockchain node; wherein:
[0012] The first information recommendation node is used to implement the steps of any one of the above information recommendation methods;
[0013] The blockchain node is used to run an intelligent protocol and store a reinforcement learning model recommended by an energy provider for replenishing energy for at least one energy device.
[0014] In a third aspect, an information recommendation device is provided, the information recommendation device is used to run the first information recommendation node, the device at least comprising: a memory, a processor and a communication bus; wherein:
[0015] The memory is used to store executable instructions;
[0016] The communication bus is used to realize the communication connection between the processor and the memory;
[0017] The processor is used to execute the information recommendation program stored in the memory to implement the steps of any one of the above-mentioned information recommendation methods.
[0018] In a fourth aspect, a storage medium is provided, wherein an information recommendation program is stored on the storage medium, and when the information recommendation program is executed by a processor, the steps of any of the above-mentioned information recommendation methods are implemented.
[0019] The embodiment of the present application provides an information recommendation method, system, device and storage medium. If it is determined that the target energy device has an energy replenishment demand, the first information recommendation node obtains the environmental parameters of the area where the target energy device is located and the current state parameters of the target energy device, and determines the trained reinforcement learning model. Finally, the environmental parameters and the current state parameters are predicted and processed by the reinforcement learning model to obtain the recommended result of the energy providing device that provides energy replenishment for the target energy device. In this way, when the target energy device has an energy replenishment demand, the first information recommendation node uses the reinforcement learning model stored in the blockchain node to locally predict the environmental parameters of the area where the target energy device is located and the current state parameters of the target energy device, and obtains the recommended result of the energy providing device that provides energy replenishment for the target energy device, which solves the problem that the current charging pile recommendation service does not fully consider the individual needs of users, and proposes a method for recommending charging piles according to the real-time status of users, which fully considers the individual needs of users and ensures that the obtained prediction model has strong real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Schematic diagram of the process of the information recommendation method provided in the embodiment of the present application Figure 1 ;
[0021] Figure 2 Schematic diagram of the process of the information recommendation method provided in the embodiment of the present application Figure 2 ;
[0022] Figure 3 Schematic diagram of the process of the information recommendation method provided in the embodiment of the present application Figure 3 ;
[0023] Figure 4 Schematic diagram of the process of the information recommendation method provided in the embodiment of the present application Figure 4 ;
[0024] Figure 5 An application scenario of the information recommendation method provided in the embodiment of the present application;
[0025] Figure 6 A schematic diagram of establishing a Markov decision model provided in an embodiment of the present application;
[0026] Figure 7 A schematic diagram of a model training process provided in an embodiment of the present application;
[0027] Figure 8 A schematic diagram of the structure of an information recommendation system provided in an embodiment of the present application;
[0028] Fig. 9 A schematic diagram of the structure of an information recommendation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0030] The embodiment of the present application provides an information recommendation method, referring to Figure 1 As shown, the method is applied to the first information recommendation node, and the method includes the following steps:
[0031] Step 101: If it is determined that the target energy device has an energy replenishment demand, obtain the environmental parameters of the area where the target energy device is located and the current state parameters of the target energy device.
[0032] The environmental parameters at least include distribution information of energy providing devices that provide energy in the area where the target energy equipment is located.
[0033] In an embodiment of the present application, the first information recommendation node corresponds to the target energy device, that is, the first information recommendation node only recommends energy providing devices for the target energy device, that is, the first information recommendation node runs in a terminal device bound to the target energy device, and the terminal device bound to the target energy device can be, for example, an onboard computer of the target energy device, a central control device, or a third-party mobile device that performs intelligent management of the target energy device, such as an intelligent mobile terminal device. The target energy device is usually a device that uses new energy, such as an electric car. The current state parameters of the target energy device may include, for example, the remaining energy of the target energy device, the state of the energy storage module, the model of the energy storage module, and other parameters.
[0034] The first information recommendation node monitors the energy status of the target energy device. When it is detected that the target energy device needs energy replenishment, the first information recommendation node determines the current location of the target energy device, determines the environmental parameters within a certain range of the current location, and determines the current status parameters of the target energy device.
[0035] Step 102: Determine the trained reinforcement learning model.
[0036] Among them, the reinforcement learning model is obtained by updating the parameters of the reinforcement learning model stored in the blockchain node.
[0037] In the embodiment of the present application, when the reinforcement learning model stored in the blockchain node is updated, the first information recommendation node will be notified immediately to update the locally stored model, so as to synchronize the reinforcement learning model in the blockchain node and ensure that the locally stored model is the latest reinforcement learning model. The reinforcement learning model stored in the blockchain node can be obtained by the blockchain node after the first information recommendation node and at least one second information recommendation node perform model training locally and the obtained model is fed back to the blockchain node, and the blockchain node fuses the model.
[0038] Step 103: predict and process the environmental parameters and current state parameters through the reinforcement learning model to obtain a recommendation result of an energy supply device that provides energy replenishment for the target energy equipment.
[0039] In the embodiment of the present application, the environmental parameters and current state parameters of the target energy device are input into the reinforcement learning model, so that the input environmental parameters and current state parameters are predicted and processed by the reinforcement learning model to obtain the energy providing device recommendation result. Further, the energy providing device recommendation result can be displayed, so that the user can perform a selection operation based on the energy providing device recommendation result and select the energy providing device that provides energy for the target energy device.
[0040] The information recommendation method provided in the embodiment of the present application, if it is determined that the target energy device has an energy replenishment demand, the first information recommendation node obtains the environmental parameters of the area where the target energy device is located and the current state parameters of the target energy device, and determines the trained reinforcement learning model, and finally predicts and processes the environmental parameters and the current state parameters through the reinforcement learning model to obtain the recommended result of the energy providing device that provides energy replenishment for the target energy device. In this way, when the target energy device has an energy replenishment demand, the first information recommendation node uses the reinforcement learning model stored in the blockchain node to predict and process the environmental parameters of the area where the target energy device is located and the current state parameters of the target energy device locally, and obtains the recommended result of the energy providing device that provides energy replenishment for the target energy device, which solves the problem that the current charging pile recommendation service does not fully consider the individual needs of users, and proposes a method for recommending charging piles according to the real-time status of users, which fully considers the individual needs of users and ensures that the obtained prediction model has strong real-time performance.
[0041] The embodiment of the present application provides an information recommendation method, referring to Figure 2 As shown, the method is applied to the first information recommendation node, and the method includes the following steps:
[0042] Step 201: If it is determined that the target energy device has an energy replenishment demand, obtain the environmental parameters of the area where the target energy device is located and the current state parameters of the target energy device.
[0043] The environmental parameters at least include distribution information of energy providing devices that provide energy in the area where the target energy equipment is located.
[0044] In an embodiment of the present application, the target energy device is an electric vehicle, and the first information recommendation node is a central control device installed on the electric vehicle. When the central control device detects that the current remaining battery power of the electric vehicle is lower than the preset power, it is determined that the electric vehicle currently has an energy replenishment demand. The central control device obtains the environmental parameters of the area where the electric vehicle is currently located, including the charging stations in the current area and the conditions of each charging pile in each charging station, and obtains the current remaining power of the electric vehicle, as well as parameters such as the battery model of the electric vehicle, to obtain the current status parameters of the electric vehicle.
[0045] Step 202: Receive a first update parameter broadcasted by at least one second information recommendation node.
[0046] Among them, the first update parameter is the relevant parameter in the target prediction model broadcasted by the blockchain node to at least one second information recommendation node after the blockchain node determines that the reinforcement learning model is obtained.
[0047] In an embodiment of the present application, at least one second information recommendation node is an information recommendation node installed on other energy devices. In the application scenario, the roles of the first information recommendation node and the second information recommendation node can be interchanged. At least one second information recommendation node and the first information recommendation node perform model training on the reinforcement learning model before the update in the blockchain node, and then each trades the parameters of the corresponding part of the model obtained after training to the blockchain node. The blockchain node integrates and updates the parameters of the reinforcement learning model before the update with the parameters of the transaction of at least one second information recommendation node to obtain the reinforcement learning model, and then notifies at least one second information recommendation node to broadcast the parameters of their respective transactions to the blockchain node to obtain the first updated parameters. In some application scenarios, at least one second information recommendation node includes the first information recommendation node, or may not include the first information recommendation node. The specific situation is determined by the actual situation and is not limited here.
[0048] Step 203: Based on the first update parameter broadcast by at least one second information recommendation node, update the locally stored local prediction model to obtain a reinforcement learning model.
[0049] In an embodiment of the present application, the first information recommendation node uses the first update parameter broadcast by at least one second information recommendation node to update the parameters of the local prediction model locally stored in the first information recommendation node to obtain a reinforcement learning model.
[0050] Step 204: predict and process the environmental parameters and current state parameters through the reinforcement learning model to obtain a recommendation result of an energy supply device that provides energy replenishment for the target energy equipment.
[0051] In the embodiment of the present application, the central control device uses a reinforcement learning model to calculate the environmental parameter number and the current state parameter, implements the prediction processing process, and outputs the energy providing device recommendation result.
[0052] For example, after the central control device obtains the energy provider recommendation result, the central control device displays the energy provider recommendation result on its display screen, so that the user can choose the place where he wants to charge and the corresponding charging pile according to the prediction result. In some application scenarios, the prediction result can also be sent to the corresponding mobile terminal device through the central control device to be displayed on the user's mobile terminal device, which can be determined by the actual situation.
[0053] Based on the above embodiments, in other embodiments of the present application, refer to Figure 3 As shown, before the first information recommendation node executes step 202, it is also used to execute steps 205 to 209:
[0054] Step 205: When the number of samples in the experience replay pool exceeds the preset number of samples, the samples in the experience replay pool are sampled to obtain a first number of samples to be trained.
[0055] Among them, the preset number of samples is less than or equal to the maximum sample capacity value of the experience replay pool.
[0056] In an embodiment of the present application, the experience replay pool is used to store sample data obtained by the first information recommendation node, and the maximum number of samples can be an experience value obtained based on a large number of experiments or actual experience values. The preset number of samples can also be an experience value obtained based on a large number of experiments or actual experience values. The first number is an experience value obtained based on a large number of experiments or actual experience values.
[0057] The first information recommendation node detects the sample number of sample data stored in the experience replay pool set locally, and when it is determined that the sample number is greater than or equal to the preset sample number, it uses a preset sample sampling method to perform sample sampling processing on the sample data in the experience replay pool to obtain a first number of samples to be trained.
[0058] When the number of samples in the experience revisit pool does not exceed the preset number of samples, the first information recommendation node can directly use the local prediction model of the model to be trained to perform model prediction, obtain the corresponding prediction results, and display the prediction results to achieve the operation of recommending new energy providing devices to users.
[0059] When the central control device detects that the number of samples in the experience replay pool exceeds the preset number of samples, the central control device samples the samples in the experience replay pool according to the preset sampling method to obtain the first number of samples to be trained. The maximum sample capacity value of the experience replay pool set in the central control device can be automatically set according to the time period based on the different dynamic characteristics of each information recommendation node in the actual recommendation scenario.
[0060] Step 206: Determine the weight coefficient of each sample to be trained to obtain a first number of weight coefficients.
[0061] In the embodiment of the present application, the weight coefficient of each sample to be trained can be calculated by using a preset weight coefficient determination method. For example, it can be determined by presetting a policy function related to the sample data to be trained and a probability distribution function based on an action. In this way, the first number of samples to be trained are analyzed and calculated respectively using the preset weight coefficient determination method to obtain the weight coefficient of each sample to be trained.
[0062] Exemplarily, the central control device calculates the current policy value and behavior policy value of each sample to be trained, and calculates the corresponding weight coefficient of the sample to be trained according to the current policy value and behavior policy value of each sample to be trained. In this way, the weight coefficients of the first number of samples to be trained are calculated to obtain the first number of weight coefficients. Since the historical samples in the experience replay pool are not generated by the current blockchain network, a significant variance problem will be introduced, so the weight coefficient is introduced to improve the noise importance sampling method to solve this problem.
[0063] Step 207: Use a first number of samples to be trained and a first number of weight coefficients to perform model training on the local prediction model using a preset algorithm to obtain a reference prediction model.
[0064] In an embodiment of the present application, the local prediction model may be a prediction model that is notified on a blockchain node and that requires model training, and may be used to provide a recommendation of an energy group providing device for a target energy device before model training is required. The local prediction model is trained using a first number of samples to be trained and a first number of weight coefficients to obtain a reference prediction model. The energy providing device may be, for example, a charging pile or other device that provides users with new energy required by new energy devices.
[0065] The central control device inputs each sample to be trained and the weight coefficient corresponding to each sample to be trained into the local prediction model, and repeats the model training until a reference prediction model that meets the expectations is obtained.
[0066] Step 208: Divide the reference prediction model to obtain a second number of analysis blocks.
[0067] In an embodiment of the present application, the trained reference prediction model is divided into blocks according to a certain division rule to obtain a second number of analysis blocks.
[0068] The central control device divides the reference prediction model into blocks according to the model characteristics of the reference prediction model to obtain a second number of analysis blocks. The first information recommendation node, each second information recommendation node, and the blockchain node all use the same rule when dividing the corresponding prediction models, so that the divided blocks have a one-to-one correspondence.
[0069] Step 209: Based on the second number of analysis blocks, update the local prediction model in the blockchain node.
[0070] In the embodiment of the present application, after the first information recommendation node analyzes and processes the second number of analysis blocks, the local prediction model in the blockchain node can be updated based on the second number of analysis blocks according to the analysis results. It should be noted that the blockchain node will perform the fusion update operation of the local prediction model only after receiving a certain number of corresponding analysis blocks sent by the first information recommendation node.
[0071] Exemplarily, after dividing the second number of analysis blocks, the central control device analyzes and processes the second number of analysis blocks to obtain analysis results, and determines block information that can be uploaded to the blockchain node based on the analysis results, so that the blockchain node can update the local prediction model in the blockchain node based on the received block information.
[0072] In this way, by utilizing the decentralized and distributed characteristics of blockchain, an efficient, preferential, decentralized and personalized charging facility recommendation model solution is provided for electric vehicle users living in small and medium-sized cities.
[0073] Based on the above embodiment, in other embodiments of the present application, step 206 can be implemented by steps 206a to 206e:
[0074] Step 206a: Determine the current policy value of each sample to be trained.
[0075] In the embodiment of the present application, the central control device can use the current strategy function To determine the current policy value of each sample to be trained, in the embodiment of the present application, the current policy function can be set to: Among them, β is the noise extracted from the normal distribution, which is usually an empirical value obtained from a large number of experiments, and softmax() is the normalization function; is the corresponding selection result in the i-th sample to be trained, s t is the characteristic parameter of the new energy device corresponding to the i-th sample to be trained.
[0076] Step 206b: Determine the behavior strategy value of each sample to be trained.
[0077] In the embodiment of the present application, the behavior strategy function can be used To determine the behavior strategy value of each sample to be trained, Specifically, we can use the conditional probability distribution based on action To calculate it.
[0078] Step 206c: Calculate the sum of the current strategy value and the smoothing hyperparameter of each sample to be trained to obtain a first value.
[0079] In the embodiment of the present application, the smoothing hyperparameter μ may be an empirical value obtained from a large number of experiments, or an empirical value obtained from actual experience.
[0080] Step 206d: Calculate the sum of the behavior measurement value of each sample to be trained and the smoothing hyperparameter to obtain a second value.
[0081] Step 206e: determine the ratio of the first value to the second value of each sample to be trained, obtain the weight coefficient of each sample to be trained, and then obtain a first number of weight coefficients.
[0082] In the embodiment of the present application, the central control device can use the following formula to calculate the weight coefficient of each sample to be trained:
[0083] Based on the above embodiment, in other embodiments of the present application, step 209 can be implemented by steps 209a to 209c:
[0084] Step 209a: Determine the third quantity based on the performance parameters of the nodes recommended by the first information.
[0085] In an embodiment of the present application, the performance parameters of the first information recommendation node may include operating performance parameters of the information recommendation device running the first information recommendation node, such as remaining operating space of the central processing unit, remaining operating memory storage space and other parameters. The third quantity can be continuously adjusted according to the operating time.
[0086] Step 209b: Select a third number of blocks to be processed from the second number of analysis blocks.
[0087] In the embodiment of the present application, the third number of blocks to be processed may be selected from the second number of analysis blocks at random, or may be selected after calculating and ranking the contribution value of each analysis block.
[0088] Step 209c: Based on the third number of blocks to be processed, update the local prediction model in the blockchain node to obtain a reinforcement learning model.
[0089] In an embodiment of the present application, the central control device processes a third number of blocks to be processed, and then broadcasts the relevant information content of the third number of blocks to be processed. In this way, after the blockchain node receives the relevant information content of the third number of blocks to be processed, the relevant information content of the third number of blocks to be processed can be used to update the local prediction model in the blockchain node to obtain a reinforcement learning model.
[0090] Based on the above embodiment, in other embodiments of the present application, step 209c can be implemented by steps a11 to a17:
[0091] Step a11: Serialize the second number of analysis blocks to obtain a serialized block queue.
[0092] In the embodiment of the present application, a serialization method is used to perform serialization processing on the second number of analysis blocks to obtain a first serialized block.
[0093] Step a12: determine the sorting positions of the third number of blocks to be processed in the serialized block queue to obtain a target position set.
[0094] In the embodiment of the present application, the sorting positions of the selected third number of blocks to be processed are determined from the serialized block queue, and a target position set of position information about the third number of blocks to be processed is obtained.
[0095] Step a13: Determine the contribution value of each of the third number of blocks to be processed to obtain a third number of target contribution values.
[0096] In the embodiment of the present application, a preset contribution value calculation method is used to calculate the contribution value of each of the third number of blocks to be processed to obtain a third number of target contribution values.
[0097] Step a14, broadcasting the target position set, the third number of target contribution values, the third number and the maximum sample capacity value.
[0098] In an embodiment of the present application, the first information recommendation node packages the target position set of the third number of blocks to be processed, the third number of target contribution values, the third number and the maximum sample capacity value of the experience replay pool of the first information recommendation node, and then uses a broadcast communication method for broadcast processing. The target position set includes at least the identification information of the third number of blocks to be processed and the sorting position in the serialized block queue. The blockchain node receives the target position set, the third number of target contribution values, the third number and the maximum sample capacity value sent by the preset number of first information recommendation nodes, and compares and analyzes the content sent by the preset number of first information recommendation nodes to select the corresponding block in the optimal first information recommendation node corresponding to the same block to update the local prediction model in the blockchain node. The blockchain node selects the corresponding block to be fused from the information sent by the preset number of first information recommendation nodes, and broadcasts the block to be fused. It should be noted that the third number corresponding to different first information recommendation nodes may be different, which is specifically determined by the specific corresponding first information recommendation node.
[0099] Step a15: If the block to be merged broadcasted by the blockchain node belongs to the third number of blocks to be processed, determine the adjustable parameter value of the block to be merged in the third number of blocks to be processed.
[0100] In the embodiment of the present application, after receiving the block to be merged, the corresponding first information recommendation node determines whether the block to be merged is its own block. If it is its own block, it determines the adjustable parameter value of the block to be merged.
[0101] Step a16: Packaging the adjustable parameter values of the block to be merged to obtain broadcast data.
[0102] Step a17: broadcast the broadcast data.
[0103] Among them, the broadcast data is used to update the parameter values in the corresponding block in the local prediction model in the blockchain node.
[0104] Based on the above embodiment, in other embodiments of the present application, step a13 can be implemented by steps a131 to a132:
[0105] Step a131, obtaining a third number of reference blocks corresponding to the latest prediction model in the blockchain node.
[0106] The third number of reference blocks corresponds one-to-one to the third number of blocks to be processed.
[0107] In an embodiment of the present application, the latest prediction model is the latest model currently stored in the blockchain node. The latest prediction model is divided into blocks using the method of dividing the prediction model in the first information recommendation node, and a third number of reference blocks corresponding to the third number of blocks to be processed are selected from the blockchain node.
[0108] Step a132: Determine the contribution value of each block to be processed based on each block to be processed and the reference block corresponding to each block to be processed, so as to obtain a third number of target contribution values.
[0109] In an embodiment of the present application, the similarity between each block to be processed and the reference block corresponding to each block to be processed is calculated as the contribution value of each block to be processed, or the standard deviation between the parameters of each block to be processed and the reference block corresponding to each block to be processed is calculated as the contribution value of each block to be processed.
[0110] In this way, by using smart protocols to divide and serialize the local model parameters of each information recommendation node, and using bidding to merge the models, we can achieve decentralized model training and reduce the computational workload and communication costs of blockchain technology.
[0111] Based on the above embodiments, in other embodiments of the present application, refer to Figure 4 As shown, after the first information recommendation node executes step 204, it may also choose to execute steps 210 to 214:
[0112] Step 210: Display the energy supply device recommendation result.
[0113] In the embodiment of the present application, the energy providing device recommendation result is displayed.
[0114] Step 211: Determine a selection operation based on the energy provider recommendation result.
[0115] In an embodiment of the present application, the first information recommendation node determines a selection operation for the energy providing device recommendation result based on the displayed energy providing device recommendation result. The selection object of the selection operation does not necessarily belong to the energy providing device recommendation result. The selection object is the energy providing device selected by the user to ultimately provide energy for the target energy device.
[0116] Step 212: After performing the selection operation on the target energy device, determine the update state parameters of the target energy device.
[0117] In the embodiment of the present application, after the user performs a selection operation on the target energy device, such as selecting a corresponding energy providing device to replenish energy for the target energy device, the state parameters of the target energy device after the energy replenishment is completed are determined to obtain the updated state parameters of the target energy device. The updated state parameters may correspond one to one with the aforementioned current state parameters.
[0118] Exemplarily, for example, after the user selects the recommended charging pile information displayed, the user drives the new energy vehicle to the selected charging pile according to the navigation to charge the new energy vehicle. After the charging is completed, the update status parameters of the new energy vehicle are determined.
[0119] Step 213: Calculate the current state parameter, the updated state parameter and the selection operation through the reward function to obtain the reward value corresponding to the selection operation.
[0120] In the embodiment of the present application, the reward function can be expressed as and F are the corresponding parameter results obtained by analyzing the current state parameters and the updated state parameters based on the selection content of the selection operation. In this way, the selection operation is quantified through the reward function to achieve labeling processing.
[0121] Step 214: store the current state parameters, selected operations and reward values as sample data in the experience replay pool.
[0122] In an embodiment of the present application, the current state parameters, selection parameters and reward values are stored as sample data in the experience replay pool. While increasing the samples, the user's selection operations can also be considered in real time. In this way, in the subsequent model training process, the real-time nature of the samples is guaranteed and the user's usage personality is fully considered.
[0123] Based on the above embodiments, the present application embodiment provides a schematic diagram of an application scenario application architecture, such as Figure 5 As shown, it includes blockchain node A1, information recommendation node A2, information recommendation node A3, information recommendation node A4 and information recommendation node A5. Here, only four information recommendation nodes are used as examples. Information recommendation node A2, information recommendation node A3, information recommendation node A4 and information recommendation node A5 are respectively run on the central control devices of four new energy electric vehicles. Figure 5 The schematic diagram of the application scenario shown in the figure shows that the A3C algorithm is deployed in the information recommendation node A2, information recommendation node A3, information recommendation node A4 and information recommendation node A5 respectively. The Actor network in the A3C algorithm is a policy function used to fit the policy value function, and the Critic network is a value function used to fit the state value function. The two functions are fitted using a deep neural network. Since the sample data for model training of each information recommendation node is very limited, an experience replay pool is added to the A3C algorithm. Each information recommendation node does not set the size of the replay pool. According to the different dynamic characteristics of the information recommendation nodes in the actual recommendation scenario, the capacity of the replay pool is automatically set according to the time period.
[0124] The order of the charging pile finally selected by the user in the prediction result is O, the total charging time is T, the vehicle status after charging is S' and the cost is F to evaluate the quality of the recommendation result. In order to avoid the differences caused by heterogeneous data, the O, T, S' and F data are standardized to obtain and Then use the reward function: Calculate the reward value for each choice result. Among them: α+β+γ+δ=1, α, β, γ and δ are weights, usually experience values. The value function is also only related to the current state, and the value function is expressed as V π (s) = E π (R t+1 +rR t+2 +r 2 R t+3 +…|s=S t ).
[0125] Since the process of each node obtaining instant rewards and the latest status is very sparse, we model the user's sparse charging demand search operation as a sequential decision. Since other than the state of new energy vehicles, which can be observed, other user behaviors cannot be predicted and represented, the sequential process from the interaction between the information recommendation node and the recommendation list output by the prediction model to the final selection of a charging station to complete charging can be regarded as a Markov process, such as Figure 6As shown. Among them, the probability of taking action a in a certain state at the current moment is also only related to the current state, which can be expressed by the formula: π(a|s)=P(A t =a|s=S t ), where π(a|s) is the policy function of the process, which can be represented by a conditional probability distribution. At this time, the action with a large probability is selected with a high probability. For the value after taking the action, we use the value function v π (s) indicates that the value function is also only related to the current state and can be expressed as: V π (s) = E π (R t+1 +rR t+1 +r 2 R t+2 +……|s=S t ), the value function is a recursive function. After being converted into the Bellman equation, it can be solved by dynamic programming. For the automatic recommendation problem of charging facilities, it is impossible to determine the state transition probability, so reinforcement learning can be used to solve it.
[0126] Since the online training data for each child node is very limited, in this application we add an experience replay pool to the Actor-critic algorithm. Each node does not set the replay pool size. According to the different dynamic characteristics of each node in the actual recommendation scenario, the replay pool capacity is automatically set according to the time period.
[0127] Correspondingly, the process of using the Actor-critic algorithm to implement model training in smart cars can be referred to Figure 7 As shown, the following steps are included:
[0128] Step b11, start.
[0129] Step b12: synchronize the global neural network parameters to the smart car.
[0130] The smart car corresponds to the aforementioned first information recommendation node. The global neural network parameters are the parameters of the model used for prediction in the blockchain node. The smart car synchronizes the global neural network parameters and updates the parameters in the local prediction model.
[0131] Step b13: Obtain the state St of the smart car.
[0132] Step b14: Select action at based on the policy network.
[0133] Among them, the policy network is the local prediction model after updating the parameters.
[0134] Step b15, execute action at, calculate reward rt, determine the new state St+1 of the smart car, and the corresponding conditional probability π(a|s).
[0135] Among them, the action at, the reward rt, the new state St+1 of the smart car, and the corresponding conditional probability π(a|s) are stored in the experience revisit pool.
[0136] Step b16: determine whether the training conditions are met. If so, execute step b17; if not, execute step b13.
[0137] Among them, the training conditions can be, for example, that the number of samples in the experience revisit pool exceeds a certain number, or that the blockchain node instructs the smart car to perform model training.
[0138] Step b17: perform importance sampling (s1, s2, ..., st) from the experience revisit pool and calculate the sampling weight w'.
[0139] The sampling weight w' of each sample can be calculated by the following calculation formula: Among them, μ is the smoothing hyperparameter set, is the probability distribution of the behavior strategy function based on the action, which is equivalent to π(s t |a t ,δ′), where δ′ is the parameter of the strategy function at the previous moment, is the current policy function, which can be set to: β is the noise drawn from a normal distribution.
[0140] Step b18: Calculate Rq(a) based on the evaluation network i |S i ).
[0141] Step b19, i=k-1.
[0142] Step b20, i=i-1.
[0143] Step b21: Accumulate Actor network gradients.
[0144] Step b22: accumulate critic network gradients.
[0145] Step b23, determine whether i is 0, if it is 0, execute step b24, otherwise execute step b20.
[0146] Step b24: Update the local prediction model.
[0147] After completing local training, the idea of federated learning is mainly applied in the process of model fusion. The recommendation model recommends global updates of the model by uploading and fusing model parameters. This model fusion method does not require uploading user privacy data, but only fusing model parameters. The decentralized and distributed characteristics of blockchain are used to solve the problems caused by model fusion by the central server. The following is the workflow of parameter fusion of the smart protocol designed in this proposal:
[0148] 1) After the local training of the information recommendation node is completed, the local model is obtained, and the local model is divided into N different parameter blocks according to the intelligent protocol, which are recorded as
[0149] 2) The information recommendation node serializes each parameter block so that each parameter block can be uploaded, read and written independently and seamlessly. At the same time, the independence between blocks can be used to perform parallel updates from multiple devices at the same time, which can be recorded as:
[0150] 3) The information recommendation node divides each parameter into blocks Compared with the latest global model C on the blockchain node global The corresponding blocks [q1,q2,…,q N ] to compare and obtain the blocks and q j The similarity between (j=1,2,…,N,) or the standard deviation of the parameters is assigned to each block to obtain the contribution value P j i .
[0151] 4) The information recommendation node determines its maximum contribution rate, i.e., the maximum number of shared model blocks λ. λ is determined by the computing power of the information recommendation node, so as to ensure that the computing and communication bandwidth usage on each information recommendation node does not exceed the load of the information recommendation device running the information recommendation node. The information recommendation node randomly selects λ blocks according to the budget and records the block positions and contribution values of these λ blocks.
[0152] 5) The information recommendation node broadcasts the number of blocks λ, the block positions of the λ blocks, the contribution values of the λ blocks, and the replay pool capacity of the information recommendation node to all training nodes (i.e., other information recommendation nodes).
[0153] 6) If the smart protocol (i.e., blockchain node) does not receive the broadcast of the specified number of information recommendation nodes, the information recommendation node continues to perform local training.
[0154] 7) If the smart protocol receives a specified number of node applications, the smart protocol calculates the contribution value of the λ blocks uploaded by each information recommendation node and the playback pool capacity K of the corresponding information recommendation node. i Select blocks to update the global model. The process of selecting blocks can be: For the mth block among the λ blocks of node i, its score can be recorded as: score m =K i *P n i , n is the block sequence number of λ blocks.
[0155] 8) The smart protocol automatically selects the block with the highest score as the update block according to the score, and then broadcasts it to all online information recommendation nodes. The information recommendation node to which the block with the highest score belongs packages the blocks through the smart protocol, and broadcasts the block information of the block with the highest score to all online information recommendation nodes.
[0156] 9) All information recommendation nodes verify the received block information, and after verification, update the local global model.
[0157] In this way, a federated learning idea is adopted in the global training process. Each node only needs to transmit model parameters, avoiding the delay and privacy leakage caused by data transmission. A decentralized model parameter fusion method based on blockchain technology is proposed. Through an intelligent protocol suitable for decentralized upload and fusion of deep learning model parameters, all parameters of the deep neural network are avoided from being transmitted on the chain, reducing the computational workload and data transmission cost of each intelligent node. Furthermore, by simplifying and abstracting the charging facility recommendation into a timing strategy generation problem, and constructing a mathematical model, the optimal solution of the recommendation model is finally obtained by using the reinforcement learning method. When solving the reinforcement learning model, the A3C algorithm idea suitable for distributed machine learning is borrowed. The improved experience replay pool technology is added to each information recommendation node to solve the problem of too few local samples. Specifically, the problem of too large variance after adding the replay pool is solved by the improved noise importance sampling algorithm. The proposed reinforcement learning method realizes the independent and identically distributed data during the recommendation model training, and enhances the robustness of the final generated global recommendation model.
[0158] It should be noted that, for the description of the same steps and the same contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.
[0159] An information recommendation method provided by an embodiment of the present application, when the number of samples in an experience replay pool exceeds a preset number of samples, a first information recommendation node samples the samples in the experience replay pool, obtains a first number of samples to be trained, determines a weight coefficient for each sample to be trained, obtains a first number of weight coefficients, and uses the first number of samples to be trained and the first number of weight coefficients to train a local prediction model to obtain a reference prediction model, then divides the reference prediction model to obtain a second number of analysis blocks, and based on the second number of analysis blocks, updates the local prediction model in the blockchain node, finally receives a first update parameter broadcasted by at least one second information recommendation node, and updates the local prediction model based on the first update parameter to obtain a target prediction model, uses the target prediction model to predict the first current feature parameter, obtains a prediction result, and displays the prediction result. In this way, after the first information recommendation node performs model training at the node to obtain a reference prediction model, the reference prediction model is divided into blocks, and the local prediction model in the blockchain node is updated according to the analysis blocks obtained by the division, so as to realize model prediction localization, and the prediction model in the blockchain node is updated by the locally trained prediction model to realize decentralization. After the blockchain node updates the local prediction model, the blockchain node notifies at least one corresponding second information recommendation node to broadcast the first update parameter of the updated local prediction model. After receiving the first update parameter, the first information recommendation node updates the locally stored local prediction model to obtain the target prediction model, and uses the target prediction model to predict the first current characteristic parameter of the new energy device to obtain the prediction result. In this way, the problem of excessive centralization of services when recommending charging piles is solved, and a decentralized service provision solution is proposed, which ensures that when the central server fails, the new energy device can still make effective predictions based on its own characteristic parameters to recommend charging piles, thereby increasing the universality of the solution.
[0160] Based on the above embodiments, the embodiments of the present application provide an information recommendation system, which can be applied to Figures 1 to 4 In the information recommendation method provided in the corresponding embodiment, refer to Figure 8 As shown, the information recommendation system 3 may include: a first information recommendation node 31 and a blockchain node 32; wherein:
[0161] The first information recommendation node 31 is used to implement the following steps:
[0162] When the number of samples in the experience replay pool exceeds the preset number of samples, sampling the samples in the experience replay pool to obtain a first number of samples to be trained; wherein the preset number of samples is less than or equal to the maximum sample capacity value of the experience replay pool;
[0163] Determine a weight coefficient for each sample to be trained to obtain a first number of weight coefficients;
[0164] Using a first number of samples to be trained and a first number of weight coefficients, the local prediction model is trained to obtain a reference prediction model; wherein the reference prediction model is used to recommend a new energy providing device that provides new energy for the new energy equipment;
[0165] dividing the reference prediction model to obtain a second number of analysis blocks;
[0166] Based on the second number of analysis blocks, updating the local prediction model in the blockchain node;
[0167] Receiving a first update parameter broadcasted by at least one second information recommendation node; wherein the first update parameter is broadcasted by the blockchain node after the blockchain node updates the local prediction model and notifies the at least one second information recommendation node;
[0168] Based on the first update parameter, updating the local prediction model to obtain a target prediction model;
[0169] Obtaining a first current characteristic parameter of the new energy device;
[0170] Using the target prediction model to predict the first current feature parameter to obtain a prediction result;
[0171] Display the prediction results;
[0172] The blockchain node 32 is used to run the smart protocol. After receiving the analysis blocks broadcast by a preset number of first information recommendation nodes, it determines at least one block to be fused based on the analysis blocks broadcast by the preset number of first information recommendation nodes, updates the local prediction model based on the at least one block to be fused to obtain the target model, and broadcasts the identification information of at least one block to be fused, so that the corresponding first information recommendation node broadcasts the adjustable parameter value of at least one block to be fused.
[0173] In other embodiments of the present application, the first information recommendation node executes the step of sampling the samples in the experience replay pool when the number of samples in the experience replay pool exceeds the preset number of samples to obtain the first number of samples to be trained, and is also used to execute the following steps:
[0174] Obtain the second updated parameter of the current prediction model from the blockchain node;
[0175] Based on the second update parameter, the local prediction model is updated to obtain a local prediction model.
[0176] In other embodiments of the present application, when the first information recommendation node executes the step of determining the weight coefficient of each sample to be trained and obtains the first number of weight coefficients, it can be implemented by the following steps:
[0177] Determine the current policy value for each sample to be trained;
[0178] Determine the behavior strategy value of each sample to be trained;
[0179] Calculate the sum of the current strategy value and the smoothing hyperparameter of each sample to be trained to obtain a first value;
[0180] Calculate the sum of the behavior measurement value of each sample to be trained and the smoothing hyperparameter to obtain a second value;
[0181] The ratio of the first value to the second value of each sample to be trained is determined to obtain a weight coefficient of each sample to be trained, thereby obtaining a first number of weight coefficients.
[0182] In other embodiments of the present application, when the first information recommendation node executes the step of updating the local prediction model in the blockchain node based on the second number of analysis blocks, it can be implemented by the following steps:
[0183] Determining a third number based on the performance parameters of the recommended nodes of the first information;
[0184] Selecting a third number of blocks to be processed from the second number of analysis blocks;
[0185] Based on the third number of blocks to be processed, the local prediction model in the blockchain node is updated.
[0186] In other embodiments of the present application, when the first information recommendation node executes the step of updating the local prediction model in the blockchain node based on the third number of blocks to be processed, it can be implemented by the following steps:
[0187] Serializing the second number of analysis blocks to obtain a serialized block queue;
[0188] Determine the sorting positions of the third number of blocks to be processed in the serialized block queue to obtain a target position set;
[0189] Determine the contribution value of each of the third number of blocks to be processed to obtain a third number of target contribution values;
[0190] broadcasting a set of target positions, a third number of target contribution values, a third number, and a maximum sample capacity value;
[0191] If the block to be merged broadcasted by the blockchain node belongs to the third number of blocks to be processed, determine the adjustable parameter value of the block to be merged in the third number of blocks to be processed;
[0192] The adjustable parameter values of the block to be fused are packaged to obtain broadcast data;
[0193] Broadcast broadcast data; wherein the broadcast data is used to update the parameter values in the corresponding block in the local prediction model in the blockchain node.
[0194] In other embodiments of the present application, when the first information recommendation node executes the step of determining the contribution value of each of the third number of blocks to be processed and obtains the third number of target contribution values, it can be achieved by the following steps:
[0195] Obtaining a third number of reference blocks corresponding to the latest prediction model in the blockchain node; wherein the third number of reference blocks corresponds one-to-one to the third number of blocks to be processed;
[0196] Based on each block to be processed and a reference block corresponding to each block to be processed, a contribution value of each block to be processed is determined, thereby obtaining a third number of target contribution values.
[0197] In other embodiments of the present application, the first information recommendation node executes the step of sampling the samples in the experience replay pool when the number of samples in the experience replay pool exceeds the preset number of samples to obtain the first number of samples to be trained, and is also used to execute the following steps:
[0198] Acquire a second current characteristic parameter of the new energy device; wherein the second current characteristic parameter at least includes a performance parameter of an energy storage device of the new energy device, a state parameter of a new energy providing device in an area where the new energy device is located, and a location parameter of the new energy device;
[0199] Using the local prediction model to predict the second current feature parameter to obtain a prediction result;
[0200] Display the prediction results;
[0201] Based on the prediction results, determine the selection results;
[0202] The second current feature parameter and the selection result are stored as sample data in the experience replay pool.
[0203] It should be noted that the specific implementation process of the steps performed by the information recommendation system in this embodiment can refer to Figures 1 to 4 The implementation process of the information recommendation method provided in the corresponding embodiment will not be repeated here.
[0204] An information recommendation system provided by an embodiment of the present application, when the number of samples in an experience replay pool exceeds a preset number of samples, a first information recommendation node samples the samples in the experience replay pool, obtains a first number of samples to be trained, determines a weight coefficient for each sample to be trained, obtains a first number of weight coefficients, and uses the first number of samples to be trained and the first number of weight coefficients to train a local prediction model to obtain a reference prediction model, then divides the reference prediction model to obtain a second number of analysis blocks, and based on the second number of analysis blocks, updates the local prediction model in the blockchain node, finally receives a first update parameter broadcasted by at least one second information recommendation node, and updates the local prediction model based on the first update parameter to obtain a target prediction model, uses the target prediction model to predict the first current feature parameter, obtains a prediction result, and displays the prediction result. In this way, after the first information recommendation node performs model training at the node to obtain a reference prediction model, the reference prediction model is divided into blocks, and the local prediction model in the blockchain node is updated according to the analysis blocks obtained by the division, so as to realize model prediction localization, and the prediction model in the blockchain node is updated by the locally trained prediction model to realize decentralization. After the blockchain node updates the local prediction model, the blockchain node notifies at least one corresponding second information recommendation node to broadcast the first update parameter of the updated local prediction model. After receiving the first update parameter, the first information recommendation node updates the locally stored local prediction model to obtain the target prediction model, and uses the target prediction model to predict the first current characteristic parameter of the new energy device to obtain the prediction result. In this way, the problem of excessive centralization of services when recommending charging piles is solved, and a decentralized service provision solution is proposed, which ensures that when the central server fails, the new energy device can still make effective predictions based on its own characteristic parameters to recommend charging piles, thereby increasing the universality of the solution.
[0205] Based on the above embodiments, the embodiments of the present application provide an information recommendation device, which can be applied to Figures 1 to 4 In the information recommendation method provided in the corresponding embodiment, refer to Fig. 9 As shown, the information recommendation device 4 may include: a processor 41, a memory 42 and a communication bus 43, wherein:
[0206] A memory 42, for storing executable instructions;
[0207] A communication bus 43, used to realize the communication connection between the processor 41 and the memory 42;
[0208] The processor 41 is used to execute the information recommendation program stored in the memory 42 to achieve Figures 1 to 4 The implementation process of the information recommendation method provided in the corresponding embodiment will not be repeated here.
[0209] Based on the foregoing embodiments, the embodiments of the present application provide a computer-readable storage medium, referred to as a storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement reference Figures 1 to 4 The implementation process of the information recommendation method provided in the corresponding embodiment will not be repeated here.
[0210] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0211] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0212] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0213] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0214] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. An information recommendation method, characterized in that: The method is applied to a first information recommendation node, and the method includes: If it is determined that the target energy device has an energy replenishment demand, obtain the environmental parameters of the area where the target energy device is located and the current state parameters of the target energy device; wherein the environmental parameters include the distribution information of energy providing devices that provide energy in the area where the target energy device is located; Determine a trained reinforcement learning model; wherein the reinforcement learning model is obtained by updating the parameters of the reinforcement learning model stored in the blockchain node; The reinforcement learning model performs prediction processing according to the environmental parameters and the current state parameters to obtain a recommendation result of an energy providing device that provides energy replenishment for the target energy equipment.
2. The method according to claim 1, characterized in that The step of determining a trained reinforcement learning model comprises: Receiving a first update parameter broadcast by at least one second information recommendation node; wherein the first update parameter is a relevant parameter in the target prediction model broadcasted by the blockchain node to at least one second information recommendation node after determining to obtain the reinforcement learning model; Based on the first update parameter broadcast by at least one of the second information recommendation nodes, the locally stored local prediction model is updated to obtain the reinforcement learning model.
3. The method according to claim 2, characterized in that Before receiving the first update parameter broadcast by at least one second information recommendation node, the method further includes: When the number of samples in the experience replay pool exceeds the preset number of samples, sampling the samples in the experience replay pool to obtain a first number of samples to be trained; wherein the preset number of samples is less than or equal to the maximum sample capacity value of the experience replay pool; Determine a weight coefficient of each of the samples to be trained to obtain a first number of the weight coefficients; Using the first number of the samples to be trained and the first number of the weight coefficients, performing model training on the local prediction model using a preset algorithm to obtain a reference prediction model; Dividing the reference prediction model to obtain a second number of analysis blocks; Based on the second number of analysis blocks, the local prediction model in the blockchain node is updated.
4. The method according to claim 3, characterized in that The step of determining a weight coefficient of each of the samples to be trained to obtain a first number of the weight coefficients includes: Determining a current strategy value for each of the samples to be trained; Determining a behavior strategy value for each of the samples to be trained; Calculating the sum of the current strategy value and the smoothing hyperparameter of each of the samples to be trained to obtain a first value; Calculating the sum of the behavior measurement value and the smoothing hyperparameter of each of the samples to be trained to obtain a second value; The ratio of the first value to the second value of each of the samples to be trained is determined to obtain the weight coefficient of each of the samples to be trained, and then the first number of the weight coefficients is obtained.
5. The method according to claim 3, characterized in that: The updating of the local prediction model in the blockchain node based on the second number of analysis blocks includes: determining a third number based on performance parameters of the nodes recommended by the first information; Selecting the third number of blocks to be processed from the second number of analysis blocks; Based on the third number of the blocks to be processed, the local prediction model in the blockchain node is updated to obtain the reinforcement learning model.
6. The method according to claim 5, characterized in that The updating of the local prediction model in the blockchain node based on the third number of blocks to be processed to obtain the reinforcement learning model includes: Performing serialization processing on the second number of analysis blocks to obtain a serialized block queue; Determine the sorting positions of the third number of the to-be-processed blocks in the serialized block queue to obtain a target position set; Determine the contribution value of each of the blocks to be processed in the third number of the blocks to be processed to obtain the third number of target contribution values; Broadcasting the target position set, the third number of the target contribution values, the third number and the maximum sample capacity value; If the block to be merged broadcasted by the blockchain node belongs to the third number of the blocks to be processed, determine the adjustable parameter value of the block to be merged among the third number of the blocks to be processed; Packaging the adjustable parameter values of the blocks to be merged to obtain broadcast data; Broadcast the broadcast data; wherein the broadcast data is used to update the parameter values in the corresponding block in the local prediction model in the blockchain node to obtain the reinforcement learning model.
7. The method according to claim 6, characterized in that The determining the contribution value of each of the third number of the blocks to be processed to obtain the third number of target contribution values includes: Obtaining the third number of reference blocks corresponding to the latest prediction model in the blockchain node; wherein the third number of reference blocks corresponds one-to-one to the third number of blocks to be processed; Based on each of the blocks to be processed and the reference blocks corresponding to each of the blocks to be processed, the contribution value of each of the blocks to be processed is determined, thereby obtaining the third number of the target contribution values.
8. The method according to claim 1, characterized in that After the environmental parameters and the current state parameters are predicted and processed by the reinforcement learning model to obtain a recommendation result of an energy supply device that provides energy replenishment for the target energy equipment, the method further includes: Display the energy supply device recommendation result; Determining a selection operation based on the energy provider recommendation result; After performing the selection operation on the target energy device, determining an update state parameter of the target energy device; The current state parameter, the updated state parameter and the selection operation are calculated by a reward function to obtain a reward value corresponding to the selection operation; The current state parameter, the selection operation and the reward value are stored as sample data in an experience replay pool.
9. An information recommendation system, characterized in that: The system comprises: a first information recommendation node and a blockchain node; wherein: The first information recommendation node is used to implement the steps of the information recommendation method according to any one of claims 1 to 7; The blockchain node is used to run an intelligent protocol and store a reinforcement learning model; wherein the reinforcement learning model is used to recommend an energy providing device for energy replenishment for at least one energy device including a target energy device.
10. An information recommendation device, characterized in that: The information recommendation device is used to run the first information recommendation node, and the device at least includes: a memory, a processor and a communication bus; wherein: The memory is used to store executable instructions; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute the information recommendation program stored in the memory to implement the steps of the information recommendation method according to any one of claims 1 to 7.
11. A storage medium, characterized in that: The storage medium stores an information recommendation program, and when the information recommendation program is executed by the processor, the steps of the information recommendation method according to any one of claims 1 to 7 are implemented.