Data Push Method, Device, Electronic Device and Readable Storage Medium
By conducting online training of the original network model and the variant network model, selecting the one with high training accuracy as the new original network model, the problem of lower user interest in data push is solved, and the effect of data diversity and user preference mining is achieved.
Patent Information
- Application Number
- CN202210593749.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-05-27
AI Technical Summary
In the prior art, data push methods often push users' duplicate or similar content, resulting in a decrease in user interest and unable to effectively tap users' potential preferences.
The first data is pushed based on the user preference information output from the original network model, the user operation event information is obtained, and the original network model and the variant network model are trained online, and the one with high training accuracy is selected as the new original network model and the second data is pushed.
It improves the training accuracy of the network model, improves the diversity of push data, can explore and extend users' potential preferences, and increase user stickiness.
Smart Images

Figure CN114861821B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data push, and in particular to a data push method, apparatus, electronic device, and readable storage medium. Background Art
[0002] Nowadays, data push under big data is no longer blind push and analysis. Instead, based on a large amount of data and on the premise of analysis and mining technology, it conducts personalized analysis and precise delivery for the population within the coverage area, and at the same time provides a comprehensive pre-advertising effect prediction and post-advertising effect monitoring report for the company. It realizes intelligent and efficient precision marketing, reduces the enterprise promotion cost, and meets the personalized needs of users.
[0003] The inventors have long-term research and found that the data push methods in related technologies mainly push things with repetitive or similar content for users. On the one hand, it will reduce the user's interest in the same topic. On the other hand, this is not conducive to exploring and extending the user's potential preferences. Summary of the Invention
[0004] This application provides a data push method, apparatus, electronic device, and readable storage medium, which can improve the diversity of push data and can explore and extend the user's potential preferences.
[0005] In a first aspect, a data push method is provided. The method includes: pushing first data based on user preference information output by an original network model; obtaining operation event information of a user's operation on the first data; inputting the operation event information into the original network model and a mutated network model respectively, so that the original network model and the mutated network model are trained online; wherein, the mutated network model is obtained by structural mutation of the original network model; taking the one with higher training accuracy in the original network model and the mutated network model as the new original network model, and pushing second data based on the user preference information output by the original network model.
[0006] In a second aspect, a data push apparatus is provided. The data push apparatus includes: a push module, configured to push first data based on user preference information output by an original network model; an obtaining module, configured to obtain operation event information of a user's operation on the first data; a processing module, configured to input the operation event information into the original network model and a mutated network model respectively, so that the original network model and the mutated network model are trained online; wherein, the mutated network model is obtained by structural mutation of the original network model; and taking the one with higher training accuracy in the original network model and the mutated network model as the new original network model; the push module is further configured to push second data based on the user preference information output by the original network model.
[0007] In a third aspect, an electronic device is provided, which includes a processor and a memory coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the method provided in the first aspect as described above.
[0008] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the method provided in the first aspect as described above.
[0009] The beneficial effects of this application are as follows: Different from the prior art, this application provides a data push method, device, electronic device, and readable storage medium. The operation event information of the user's operation on the first data is used to train the original network model and the mutated network model obtained by mutating the original network model respectively, and the one with higher training accuracy is used as the new original network model. The user preference information output by the new original network model is used to push the second data. It is possible to change the network model by structural mutation to find a network model with higher training accuracy, thereby improving the training accuracy of the network model. On the other hand, the operation event information in real time is used to perform online training on the network model, so that the network model can better learn the preference changes of the user, improve the diversity of the pushed data, and be able to explore and extend the potential preferences of the user. Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:
[0011] Figure 1 is a schematic flowchart of an embodiment of the data push method provided by this application;
[0012] Figure 2 is a schematic flowchart of an embodiment of the structural mutation of the original network model provided by this application;
[0013] Figure 3 is a schematic flowchart of an embodiment of the original network model provided by this application;
[0014] Figure 4 is a schematic flowchart of an embodiment of the mutated network model provided by this application;
[0015] Figure 5 is a schematic flowchart of another embodiment of the mutated network model provided by this application;
[0016] Figure 6It is a schematic flowchart of another embodiment of the data push method provided by this application;
[0017] Figure 7 It is a schematic diagram of an application scenario of an embodiment of the data push method provided by this application;
[0018] Figure 8 It is another schematic diagram of an application scenario of an embodiment of the data push method provided by this application;
[0019] Figure 9 It is another schematic diagram of an application scenario of an embodiment of the data push method provided by this application;
[0020] Figure 10 It is a schematic flowchart of an embodiment of the data push device provided by this application;
[0021] Figure 11 It is a schematic structural diagram of another embodiment of the electronic device provided by this application;
[0022] Figure 12 It is a schematic structural diagram of another embodiment of the electronic device provided by this application;
[0023] Figure 13 It is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. It can be understood that the specific embodiments described herein are only used to explain this application, rather than limiting this application. Additionally, it should be noted that for the sake of description, only parts related to this application rather than all structures are shown in the accompanying drawings. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0025] Referring to "embodiment" herein means that the specific features, structures, or characteristics described in conjunction with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0026] Refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the data push method provided by this application. The method includes:
[0027] Step 11: Push the first data based on the user preference information output by the original network model.
[0028] In some embodiments, the original network model can be pre-trained using training samples.
[0029] Among them, the original network model is a deep reinforcement learning model. Among them, the deep reinforcement learning model can be a DQN (Deep Q-Learning Network) model or a Policy Network model.
[0030] In some embodiments, the original network model is applied to the advertising push scenario. After learning the preference information of many users, the first advertising data is pushed using the user preference information output by the original network model. Among them, the first advertising data can include multiple sub-advertising data.
[0031] In some embodiments, the original network model is applied to the commodity push scenario. After learning the preference information of many users, the first commodity data is pushed using the user preference information output by the original network model. Among them, the first commodity data can include multiple sub-commodity data.
[0032] Step 12: Obtain the operation event information of the user's operation on the first data.
[0033] Among them, the operation event information can include the operation duration and / or the number of operations of the user's operation on the first data within a preset time. For example, after pushing the first data, the operation duration and / or the number of operations of the user's operation on the first data within 1 hour, 1 day, or one week.
[0034] In some embodiments, the operation event information of multiple users' operations on the first data can be obtained. Subsequently, the user preference information of each user can be determined according to the operation event information of multiple users, and then data can be pushed according to the user preference information of each user.
[0035] In some embodiments, the data to be pushed can be used as a guide to determine which users have preferences for the target data, and then the target data is continuously pushed to this user, while the push to users who do not prefer the target data is stopped.
[0036] Step 13: Input the operation event information into the original network model and the mutated network model respectively, so that the original network model and the mutated network model are trained online; among them, the mutated network model is obtained by structurally mutating the original network model.
[0037] In some embodiments, a network model is generally formed by connecting different nodes. Among them, different nodes are interconnected to form different network layers. For example, an input layer, a convolutional layer, and an output layer can be formed. The network layers are also connected through the output nodes in one network layer and the input nodes in another network layer. The connection lines between nodes are called "edges" here.
[0038] After the original network model is offline-trained using sample data, the original network model can be used for data push.
[0039] However, since the user's preference information actually changes over time. For example, on the first day, user A operated data a in the first data, on the second day, user A operated data b in the first data, and on the third day, user A operated data c in the first data. And data a, data b, and data c have different types. Based on this, if the original network model maintains the relevant parameters of the offline training, then in fact the original network model cannot change based on real-time data, which will limit the diversity of data pushed by the original network model.
[0040] Therefore, the original network model can be structurally mutated to obtain a mutated network model, and the operation event information can be used to train the original network model and the mutated network model respectively.
[0041] Step **14**: Use the one with higher training accuracy among the original network model and the mutated network model as the new original network model, and push the second data based on the user preference information output by the original network model.
[0042] Of course, the mutated network model obtained through structural mutation does not necessarily mean that the mutated network model after mutation has better performance than the original network model. Therefore, it is necessary to determine the model performance between the two.
[0043] Since the same data is used to train network models with different structures, the training accuracy in the original network model and the mutated network model can be determined. Based on this, the one with higher training accuracy among the original network model and the mutated network model can be used as the new original network model.
[0044] Among them, the training accuracy in the original network model and the mutated network model can be determined by calculating evaluation metrics such as the accuracy, precision, recall, F1 score, and ROC curve (Receiver Operating Characteristic Curve) of each network model.
[0045] In some embodiments, after determining a model with high training accuracy, it is used as the original network model, and based on the user preference information output by the original network model, the first data is updated to obtain second data; and then the second data is pushed to the user.
[0046] The first data may include multiple sub-data, and the user may operate on any sub-data, or not operate on any sub-data, which will be recorded as the user's operation event information.
[0047] Inputting this operational event information into the original network model will output the corresponding user preference information. The first data is then updated based on the user preference information, for example, retaining some sub-data that the user prefers and replacing other sub-data that the user prefers less. This cycle allows the content of the pushed data to be dynamically updated based on the user's real-time operations.
[0048] In one application scenario, after offline training of the original network model is completed, first data is pushed based on user preference information output by the original network model. Then, operation event information indicating a user's operation on the first data is obtained within a preset time. The preset time can be 1 minute, 5 minutes, 20 minutes, 1 hour, or 2 hours. It is understood that the preset time can be set based on actual needs.
[0049] When the preset time is reached, the original network model undergoes structural mutation to obtain a variant network model. Then, the operation event information is input into the original network model and the variant network model respectively, so that the original network model and the variant network model are trained online based on the operation event information.
[0050] After the training is completed, the one with higher training accuracy between the original network model and the mutated network model is used as the new original network model, and the second data is pushed based on the user preference information output by the original network model.
[0051] Among them, when the one with higher training accuracy between the original network model and the mutated network model is used as the new original network model, the one with lower training accuracy between the original network model and the mutated network model is directly deleted to release storage space.
[0052] In the actual data push process, the above method can be used for looping. As the network model changes, the network model can improve the diversity of pushed data and mine and extend the user's potential preference information.
[0053] In one application scenario, see Figure 2 , to illustrate the model variation:
[0054] Step 21: Determine the weight matrix in the original network model.
[0055] In some embodiments, after the original network model is successfully trained offline, a weighted matrix is formed. The weighted matrix has weight values for the connection relationships between different nodes.
[0056] Step 22: Randomly transform the weight values in the weighted matrix.
[0057] For example, use a random function to randomly transform the weight values in the weighted matrix to obtain a new weighted matrix.
[0058] Step 23: Use the randomly transformed weighted matrix to perform structural mutation on the original network model to obtain a mutated network model.
[0059] Combined Figure 3 、 Figure 4 and Figure 5 are described as follows:
[0060] As Figure 3 shown, the original network model includes at least Node 1, Node 2, Node 3, Node 4, and Node 5. Among them, Node 1 is connected to Node 4 and Node 5, Node 2 is connected to Node 4, Node 3 is connected to Node 5, and Node 4 is connected to Node 5.
[0061] In some embodiments, structural mutation is performed on the original network model to obtain a mutated network model as Figure 4 shown. Among them, the mutated network model includes at least Node 1, Node 2, Node 3, Node 4, and Node 5. Among them, Node 1 is connected to Node 4 and Node 5, Node 2 is connected to Node 4, Node 3 is connected to Node 4 and Node 5, and Node 4 is connected to Node 5. That is, in this structural mutation, the connection relationship between Node 3 and the other nodes is changed.
[0062] In some embodiments, structural mutation is performed on the original network model to obtain a mutated network model as Figure 5 shown. Among them, the mutated network model includes at least Node 1, Node 2, Node 3, Node 4, Node 5, and Node 6. Among them, Node 1 is connected to Node 4 and Node 5, Node 2 is connected to Node 4, Node 3 is connected to Node 6, Node 4 is connected to Node 5, and Node 6 is connected to Node 5. That is, in this structural mutation, Node 6 is added, and the connection relationship between Node 3 and Node 5 is changed.
[0063] In other embodiments, when performing structural mutation on the original network model, nodes can be added and the connection relationships between nodes can be changed simultaneously. That is, the structural mutation includes two cases: node mutation and edge mutation. Since there is a mutual dependence relationship between the edges and nodes in the network model, the two mutation processes also affect each other. On the basis of this mutation, during the mutation process, by means of random numbers and the weighted matrix, slightly scaling some weight values in the weighted matrix can help to achieve the mutation process.
[0064] In this embodiment, the operation event information of the user's operation on the first data is used to train the original network model and the mutated network model obtained by mutating the original network model, respectively, so that the one with higher training accuracy is used as the new original network model, and the user preference information output by the new original network model is used to push the second data. The network model can be changed by utilizing structural variation to find a network model with higher training accuracy, thereby improving the training accuracy of the network model. On the other hand, the network model is trained online using real-time operation event information, so that the network model can better learn the changes in user preferences, thereby improving the diversity of pushed data and being able to explore and extend the user's potential preferences.
[0065] See Figure 6 , Figure 6 This is a flow chart of another embodiment of the data push method provided by this application. The method includes:
[0066] Step 61: Pushing first data based on the user preference information output by the original network model.
[0067] Step 62: Acquire operation event information of the user operating the first data.
[0068] Steps 61 to 62 have the same or similar technical solutions as those in the above embodiment and are not described in detail here.
[0069] Step 63: Input the operation event information into the original network model and the variant network model respectively, so that the original network model and the variant network model are trained online; wherein the variant network model is obtained by performing structural mutation on the original network model.
[0070] In some application scenarios, combined with Figure 7 Explain the online training of the original network model and the variant network model:
[0071] like Figure 7 As shown, at time t1, the operational event information at time t1 is input into the deep reinforcement learning model and the variant deep reinforcement learning model mutated by the genetic algorithm to perform online training on the deep reinforcement learning model and the variant deep reinforcement learning model. The operational event information at time t1 is then stored in a memory. The original network model corresponds to the deep reinforcement learning model, and the variant network model corresponds to the variant deep reinforcement learning model.
[0072] At time t2, the operational event information from time t1 to time t2 is input into the deep reinforcement learning model and the mutated deep reinforcement learning model mutated by the genetic algorithm to perform online training on the deep reinforcement learning model and the mutated deep reinforcement learning model. The operational event information from time t1 to time t2 is stored in a memory.
[0073] At time tn, the operation event information from time tn-1 to time tn is input into the deep reinforcement learning model and the mutated deep reinforcement learning model mutated by the genetic algorithm to perform online training on the deep reinforcement learning model and the mutated deep reinforcement learning model. And the operation event information from time tn-1 to time tn is stored in the memory.
[0074] When the preset condition is satisfied, the historical operation event information in the memory is respectively input into the deep reinforcement learning model and the mutated deep reinforcement learning model to enable online training of the deep reinforcement learning model and the mutated deep reinforcement learning model.
[0075] In some embodiments, the preset condition may be a preset time. For example, the preset time is set to 8 hours, 12 hours, 24 hours, etc. When the preset time is satisfied, the historical operation event information in the memory is respectively input into the deep reinforcement learning model and the mutated deep reinforcement learning model to enable online training of the deep reinforcement learning model and the mutated deep reinforcement learning model.
[0076] In some embodiments, the preset condition may be the number of times of online training. For example, when the number of times of online training is an integer multiple of 10, the historical operation event information in the memory is respectively input into the deep reinforcement learning model and the mutated deep reinforcement learning model to enable online training of the deep reinforcement learning model and the mutated deep reinforcement learning model.
[0077] Step 64: Obtain historical operation event information.
[0078] Among them, the historical operation event information may be all operation event information obtained from the start of online operation of the original network model to the current time.
[0079] Step 65: Determine the first training accuracy of the original network model and the second training accuracy of the mutated network model by using the historical operation event information.
[0080] In an application scenario, all operation event information received during the online operation of the original network model, that is, the historical operation event information, can be used to train the original network model and the mutated network model, and then determine the first training accuracy of the original network model after training is completed, and the second training accuracy of the mutated network model after training is completed.
[0081] Step 66: Use the network model corresponding to the higher one of the first prediction accuracy and the second training accuracy as the new original network model, and push the second data based on the user preference information output by the original network model.
[0082] In an application scenario, when the original network model is a deep reinforcement learning model, in combination with Figure 8Explanation:
[0083] The deep reinforcement learning model is a typical two - tower structure, divided into a user tower and a data tower. Among them, the input features of the user tower are user features and environmental features, and the input features of the data tower are all user, environmental, user - data cross - features, and data features. Here, user features include basic features of users such as age, gender, occupation, preference labels, etc., environmental features include active features of mobile APPs, etc., user - data cross - features include operations of users on data, such as clicks, purchases, active duration, etc., and data features include the specific content of data, such as promotional selling points and preferential information of products in the data.
[0084] In the framework of reinforcement learning, since the user - tower feature vector represents the current state of the user, it can also be regarded as a state vector. The data - tower feature vector represents the data that the system will select next. This process of selecting data is the "action" of the agent. Therefore, the data - tower feature vector is also called an action vector.
[0085] The two - tower model processes the state vector and the action vector respectively through MLP (Multilayer Perceptron). Among them, before performing MLP processing, vector embedding needs to be carried out through Embedding to form the corresponding state vector and action vector.
[0086] Then, the final action quality score Q(s, a) is generated by the interaction layer. It is through the level of this score that the agent decides which actions to take, that is, which data to push to the user. The formula is as follows:
[0087] y s,a = Q(s,a)=r immediate +γr future .
[0088] Among them, r immediate represents the reward in the current situation, and r future represents the future return.
[0089] Among them, the formula of the reward function r is as follows:
[0090] r = r0 + αr1 + βr2 + χr3.
[0091] Among them, r0 represents the reward for click and purchase behaviors, and r1, r2, r3 represent the rewards for active data such as the usage duration and frequency of data by the user in several different time periods after a single recommendation.
[0092] In an application scenario, combined with Figure 9 Explanation:
[0093] Deep reinforcement learning is a combination of deep learning and reinforcement learning. It utilizes the perception ability of deep learning to solve the modeling problems of policies and value functions, and then uses the error backpropagation algorithm to optimize the objective function. At the same time, it utilizes the decision-making ability of reinforcement learning to define problems and optimize objectives. To a certain extent, deep reinforcement learning has the general intelligence to solve complex problems and has achieved success in some fields.
[0094] As Figure 9 shown, in the offline stage, the deep reinforcement learning model is initialized. After the initialization is completed and it enters the online stage, it is roughly divided into five modules: the deep reinforcement learning model can serve as an agent, a feedback module, an environment module, an action module, and a state module.
[0095] In the online stage, the push results given by the deep reinforcement learning model can be continuously updated. Specifically, the agent outputs user preference information to the action module, and the action module pushes data to the environment module based on the user preference information.
[0096] The environment module collects the operation event information of the user on the pushed data and then sends it to the feedback module. The feedback module collects the operation event information of the pushed data and then sends it to the agent.
[0097] The agent uses the genetic algorithm to mutate the deep reinforcement learning model to obtain a mutated deep reinforcement learning model, and uses the operation event information to evaluate the training accuracy of the deep reinforcement learning model and the mutated deep reinforcement learning model, and retains the one with higher training accuracy.
[0098] Then, update the state of the retained model and output the next user preference information to the action module, and loop in this way to continuously update the model.
[0099] Among them, in the feedback module, in addition to click and purchase behaviors in the evaluation indicators, new evaluation indicators are added for the active data such as the usage duration and frequency of the pushed data by the user within several different time periods after a data push, in order to focus on long-term returns. For example, after a data push, the active data such as the usage duration and frequency of the pushed data by the user within 1 hour, 1 day, and 1 week are respectively recorded. In this way, when the online updated agent processes the dynamic changes of the push, it pays attention to both the short-term preference information of the user and the long-term preference information.
[0100] Furthermore, to a certain extent, it avoids the decrease in the user's interest in the same pushed data, thereby enabling the excavation of the user's potential preferences, accurately pushing data, continuously amplifying the user's potential preferences, and increasing user stickiness.
[0101] Refer to Figure 10 , Figure 10It is a schematic structural diagram of an embodiment of the data push device provided by this application. The data push device 100 includes: a push module 101, an acquisition module 102, and a processing module 103.
[0102] Among them, the push module 101 is used to push the first data based on the user preference information output by the original network model.
[0103] The acquisition module 102 is used to acquire the operation event information of the user's operation on the first data.
[0104] The processing module 103 is used to input the operation event information into the original network model and the mutated network model respectively, so that the original network model and the mutated network model are trained online; among them, the mutated network model is obtained by structurally mutating the original network model; and the one with the higher training accuracy in the original network model and the mutated network model is used as the new original network model.
[0105] The push module 101 is further used to push the second data based on the user preference information output by the original network model.
[0106] In some embodiments, the processing module 103 is further used to determine the weight matrix in the original network model; randomly transform the weight values in the weight matrix; and use the randomly transformed weight matrix to perform structural mutation on the original network model to obtain the mutated network model.
[0107] In some embodiments, the processing module 103 is further used to acquire historical operation event information; determine the first training accuracy of the original network model and the second training accuracy of the mutated network model by using the historical operation information; and use the network model corresponding to the higher one of the first prediction accuracy and the second training accuracy as the new original network model.
[0108] In some embodiments, the processing module 103 is further used to input the historical operation event information into the original network model and the mutated network model respectively when a preset condition is met, so that the original network model and the mutated network model are trained online.
[0109] In some embodiments, the push module 101 is further used to update the first data based on the user preference information output by the original network model to obtain the second data; and push the second data.
[0110] In other embodiments, the data push device 100 can also implement the method provided in any of the above embodiments.
[0111] Refer to Figure 11 , Figure 11It is a schematic structural diagram of an embodiment of an electronic device provided by the present application. The electronic device 110 includes a processor 111 and a memory 112 coupled to the processor 111; wherein, the memory 112 is used to store program data, and the processor 111 is used to execute the program data to implement the following method:
[0112] Push the first data based on the user preference information output by the original network model; obtain the operation event information of the user's operation on the first data; input the operation event information into the original network model and the mutated network model respectively to enable online training of the original network model and the mutated network model; wherein, the mutated network model is obtained by structural mutation of the original network model; use the one with higher training accuracy among the original network model and the mutated network model as the new original network model, and push the second data based on the user preference information output by the original network model.
[0113] It can be understood that the processor 111 is also used to execute the program data to implement the method of any of the above embodiments, which will not be elaborated here.
[0114] Refer to Figure 12 , Figure 12 It is a schematic structural diagram of another embodiment of an electronic device provided by the present application. The electronic device 110 may be, for example, a mobile electronic device. The electronic device 110 may include: a memory 112, a processor (Central Processing Unit, CPU) 111, a circuit board (not shown in the figure), a power supply circuit, and a microphone 123. The circuit board is arranged inside the space surrounded by the housing; the processor 111 and the memory 112 are arranged on the circuit board; the power supply circuit is used to supply power to each circuit or device of the electronic device; the memory 112 is used to store executable program codes; the processor 111 runs the computer program corresponding to the executable program codes by reading the executable program codes stored in the memory 112 to identify the above-mentioned identification information to implement the unlocking and waking-up functions.
[0115] The electronic device may further include: a peripheral interface 114, an RF (Radio Frequency) circuit 116, an audio circuit 117, a speaker 122, a power management chip 119, an input / output (I / O) subsystem 120, other input / control devices 121, a display 113, and an external port 115. These components communicate through one or more communication buses or signal lines 118.
[0116] The memory 112 can be accessed by the processor 111, the peripheral interface 114, etc. The memory 112 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other volatile solid-state storage devices. The peripheral interface 114 can connect the input and output peripherals of the device to the processor 111 and the memory 112.
[0117] The I / O subsystem 120 can connect the input and output peripherals on the device, such as the display 113 and other input / control devices 121, to the peripheral interface 114. The I / O subsystem 120 may include a display controller 1201 and one or more input controllers 1202 for controlling the other input / control devices 121. Among them, one or more input controllers 1202 receive electrical signals from the other input / control devices 121 or send electrical signals to the other input / control devices 121. The other input / control devices 121 may include physical buttons (press buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels. It should be noted that the input controller 1202 can be connected to any one of the following: a keyboard, an infrared port, a USB interface, and an indicating device such as a mouse.
[0118] The display 113 is an input interface and an output interface between the user electronic device and the user, and displays visual output to the user. The visual output may include graphics, text, icons, videos, etc.
[0119] The display controller 1201 in the I / O subsystem 120 receives electrical signals from the display 113 or sends electrical signals to the display 163. The display 113 detects contacts on the touch screen, and the display controller 1201 converts the detected contacts into interactions with the user interface objects displayed on the display 113, that is, realizes human-computer interaction. The user interface objects displayed on the display 113 may be icons for running games, icons for connecting to corresponding networks, etc.
[0120] The RF circuit 116 is mainly used to establish communication between the mobile phone and the wireless network (i.e., the network side), and to realize the reception and transmission of data between the mobile phone and the wireless network. For example, sending and receiving text messages, e-mails, etc. Specifically, the RF circuit 116 receives and transmits RF signals, which are also called electromagnetic signals. The RF circuit 116 converts electrical signals into electromagnetic signals or electromagnetic signals into electrical signals, and communicates with the communication network and other devices through the electromagnetic signals. The RF circuit 116 may include known circuits for performing these functions, including but not limited to an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC (COder-DECoder) chipset, a Subscriber Identity Module (SIM), and so on.
[0121] The audio circuit 117 is mainly used to receive audio data from the peripheral interface 114, convert the audio data into an electrical signal, and send the electrical signal to the speaker 122. The speaker 122 is used to restore the voice signal received by the mobile phone from the wireless network through the RF circuit 116 into sound and play the sound to the user. The power management chip 119 is used to supply power and perform power management for the hardware connected to the processor 111, the I / O subsystem 120, and the peripheral interface 114.
[0122] See Figure 13 , Figure 13 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. The computer-readable storage medium 130 stores a computer program 131. When the computer program 131 is executed by a processor, the following method is implemented:
[0123] Based on the user preference information output by the original network model, push the first data; obtain the operation event information of the user's operation on the first data; input the operation event information into the original network model and the mutated network model respectively, so that the original network model and the mutated network model are trained online; wherein, the mutated network model is obtained by mutating the structure of the original network model; use the one with higher training accuracy in the original network model and the mutated network model as the new original network model, and based on the user preference information output by the original network model, push the second data.
[0124] It can be understood that when the computer program 131 is executed by a processor, it is also used to implement the method of any of the above embodiments, which will not be elaborated here.
[0125] In summary, the present application provides a data push method, device, electronic device and readable storage medium, which uses the operation event information of the user's operation on the first data to train the original network model and the mutated network model obtained by mutating the original network model, respectively, so as to use the one with higher training accuracy as the new original network model, and use the user preference information output by the new original network model to push the second data. It can use structural variation to change the network model to find a network model with higher training accuracy, thereby improving the training accuracy of the network model. On the other hand, it uses real-time operation event information to perform online training on the network model, so that the network model can better learn the changes in user preferences, thereby improving the diversity of pushed data and being able to explore and extend the user's potential preferences.
[0126] Furthermore, by using historical operation event information to train the network model online again, the network model can learn the user's long-term preference information and short-term preference information at the same time. The output preference information can pay attention to the user's long-term preferences and short-term preferences at the same time, which helps to push data in line with user preferences and increase user stickiness.
[0127] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or ignoring or not implementing certain features.
[0128] If the integrated units in the above other embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0129] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A data push method, characterized in that, The method includes: Pushing the first data based on the user preference information output by the original network model; Obtaining the operation event information of the user's operation on the first data; Inputting the operation event information into the original network model and the mutated network model respectively, so that the original network model and the mutated network model are trained online; wherein, the mutated network model is obtained by structure mutation of the original network model; Taking the network model with higher training accuracy among the original network model and the mutated network model as the new original network model, and pushing the second data based on the user preference information output by the original network model; Wherein, the mutated network model is obtained by structure mutation of the original network model, including: Determining the weight matrix in the original network model; Randomly transforming the weight values in the weight matrix; Using the randomly transformed weight matrix to perform structure mutation on the original network model to obtain the mutated network model.
2. The method according to claim 1, wherein The taking the network model with higher training accuracy among the original network model and the mutated network model as the new original network model includes: Obtaining the historical operation event information; Using the historical operation information to determine the first training accuracy of the original network model and the second training accuracy of the mutated network model; Taking the network model corresponding to the higher one of the first prediction accuracy and the second training accuracy as the new original network model.
3. The method according to any one of claims 1-2, characterized in that, The operation event information includes the historical operation event information; the inputting the operation event information into the original network model and the mutated network model respectively, so that the original network model and the mutated network model are trained online, includes: When a preset condition is met, inputting the historical operation event information into the original network model and the mutated network model respectively, so that the original network model and the mutated network model are trained online.
4. The method according to any one of claims 1-2, characterized in that, The operation event information includes the operation duration and / or the number of operations of the user's operation on the first data within a preset time.
5. The method according to any one of claims 1-2, characterized in that, The original network model is a deep reinforcement learning model.
6. The method according to any one of claims 1-2, characterized in that, The pushing the second data based on the user preference information output by the original network model includes: Updating the first data based on the user preference information output by the original network model to obtain the second data; Pushing the second data.
7. A data push device, characterized in that, The data pushing device includes: A pushing module, configured to push the first data based on the user preference information output by the original network model; An obtaining module, configured to obtain the operation event information of the user's operation on the first data; A processing module, configured to input the operation event information into the original network model and the mutated network model respectively, so that the original network model and the mutated network model are trained online; wherein, the mutated network model is obtained by performing structural mutation on the original network model; and taking the one with higher training accuracy among the original network model and the mutated network model as the new original network model; and determining the weighted matrix in the original network model; randomly transforming the weight values in the weighted matrix; using the randomly transformed weighted matrix to perform structural mutation on the original network model to obtain a mutated network model. The pushing module is further configured to push second data based on the user preference information output by the original network model.
8. An electronic device, characterized in that, The electronic device includes a processor and a memory coupled to the processor. Wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Data processing method and device for network training, electronic equipment and storage medium
CN113240109A