Message button optimization method, device and computing device
By constructing a message button topology map and relationship graph convolutional neural network optimization model, 5G message buttons are automatically optimized, which solves the problem of relying on manual experience, increases button clicks and browsing time, and improves the delivery effect of 5G messages.
Patent Information
- Application Number
- CN202110949359.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-18
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-08-18
AI Technical Summary
The existing 5G message button optimization relies on manual experience, has the risk of errors, is inefficient, and makes it difficult to find the best optimization solution, which affects the user experience.
By obtaining user usage data from the MaaP platform, building a message button topology map, and using a graph convolutional neural network to optimize the model, buttons are optimized based on the number of clicks and browsing time. Messages are sent through chatbots to obtain rewards, and the model is tuned to select optimization actions that maximize the number of clicks and browsing time.
It has achieved automated and intelligent 5G message button optimization, increased the number of button clicks and browsing time, and improved the effectiveness of 5G message delivery.
Smart Images

Figure CN115842795B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of communication technology, and in particular to a message button optimization method, apparatus, and computing device. Background Art
[0002] The 5G messaging service, based on the native SMS interface of mobile devices, allows users to send and receive media content such as text, images, audio, video, location, and contacts. This includes point-to-point messaging, group messaging, group chat messaging, and messaging between peers and applications. Compared to the single-function SMS, 5G messaging not only broadens the scope of messaging, supporting users with multimedia content such as text, audio and video, cards, and location, but also extends the depth of the interactive experience. Users can search, discover, interact, and make payments within the messaging window, creating a one-stop information window for services. 5G messaging combines the characteristics of both 2C and 2B applications. It not only facilitates the exchange of rich media information such as voice, images, video, cards, and files between users, but also supports enterprises to provide interactive services through chatbots on the 5G messaging platform. 5G messaging is both a comprehensive upgrade to traditional SMS and provides a gateway to a rich array of upcoming 5G applications, accessible through the message button.
[0003] In the existing technology, 5G message button optimization needs to rely on manual experience. The existing technology has the following main disadvantages: 1) There is a risk of error in optimization based on manual experience; 2) Button optimization is not timely enough and inefficient, which easily affects user experience; 3) It is difficult to find the best button optimization solution. Summary of the Invention
[0004] In view of the above problems, embodiments of the present invention provide a message button optimization method, apparatus, and computing device, which overcome the above problems or at least partially solve the above problems.
[0005] According to one aspect of an embodiment of the present invention, a message button optimization method is provided, the method comprising: obtaining user usage data of each message button at the current moment from a MaaP platform, obtaining feature data of each message button based on the user usage data, wherein the user usage data includes at least the number of message button clicks and browsing time; processing the feature data of any message button through a message button optimization model and outputting a modification action of the message button; optimizing the message button based on the modification action, and obtaining a reward for the optimized message button for a preset time; returning the reward to the message button optimization model for model tuning, and selecting a button optimization action that maximizes the number of message button clicks and browsing time.
[0006] In an optional manner, obtaining user usage data of each message button at the current moment from the MaaP platform, and obtaining characteristic data of each message button based on the user data, includes: obtaining user usage data of each message button for the chat robot to send 5G messages at the current moment from the MaaP platform; constructing a message button topology map based on the user usage data of each message button; standardizing the numerical attributes of the user usage data of each message button in the message button topology map, and performing one-hot encoding on the logical relationship between the message buttons to obtain characteristic data of each message button.
[0007] In an optional manner, a message button topology map is constructed based on the user usage data of each message button, including: treating each message button as a separate intelligent entity as a node in the constructed message button topology map; in the constructed message button topology map, each node is a message button, each edge is a logical relationship between message buttons, and the characteristic of each node is the user usage data of the corresponding message button.
[0008] In an optional manner, the feature data of any of the message buttons is processed by a message button optimization model and then the modification action of the message button is output, comprising: passing the state value in the feature data of any of the message buttons through the relational graph convolution layer of the state encoder in the message button optimization model, fusing the state value with the message button agent with a neighbor relationship to generate a potential feature vector; and outputting the modification action of the message button according to the potential feature vector.
[0009] In an optional manner, the message button is optimized based on the modification action, and a reward for the optimized message button being implemented for a preset time is obtained, including: applying the 5G message open platform to optimize the message button according to the modification action; sending a 5G message to the terminal user based on the optimized message button through a chat robot; obtaining new user usage data sent by the terminal user based on the optimized message button being implemented for a preset time; and calculating the reward for the optimized message button being implemented for a preset time based on the new user usage data through a reward function.
[0010] In an optional manner, the reward is returned to the message button optimization model for model tuning, and after selecting the button optimization action that maximizes the number of message button clicks and browsing time, it includes: feeding back the new user usage data as the user usage data of the message button at the next moment to the MaaP platform.
[0011] In an optional manner, before outputting the modification action of the message button after processing the characteristic data of any of the message buttons by the message button optimization model, it includes: collecting the message button topology map of historical moments and the 5G message button modification actions of the manually labeled user usage data corresponding to each message button from the MaaP platform as a total data set; preprocessing the message button topology map of historical moments and the corresponding modification actions in the total data set; and applying the preprocessed total data set to train the message button optimization model.
[0012] According to another aspect of an embodiment of the present invention, a message button optimization device is provided, the device comprising: a data acquisition unit, for acquiring user usage data of each message button at the current moment from a MaaP platform, and acquiring feature data of each message button based on the user usage data, wherein the user usage data includes at least the number of message button clicks and browsing time; a modification action output unit, for processing the feature data of any message button through a message button optimization model and outputting a modification action of the message button; a reward acquisition unit, for optimizing the message button based on the modification action, and acquiring a reward for the optimized message button for a preset time; an action optimization unit, for returning the reward to the message button optimization model for model tuning, and selecting a button optimization action that maximizes the number of message button clicks and browsing time.
[0013] According to another aspect of an embodiment of the present invention, there is provided a computing device, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;
[0014] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of the above-mentioned message button optimization method.
[0015] According to another aspect of an embodiment of the present invention, a computer storage medium is provided, wherein the storage medium stores at least one executable instruction, and the executable instruction enables the processor to execute the steps of the above-mentioned message button optimization method.
[0016] An embodiment of the present invention obtains user usage data of each message button at the current moment from the MaaP platform, and obtains characteristic data of each message button based on the user usage data, wherein the user usage data at least includes the number of message button clicks and browsing time; the characteristic data of any message button is processed by a message button optimization model and the modification action of the message button is output; the message button is optimized based on the modification action, and a reward for the optimized message button to be implemented for a preset time is obtained; the reward is returned to the message button optimization model for model tuning, and the button optimization action that maximizes the number of message button clicks and browsing time is selected. The button optimization action that maximizes the number of message button clicks and browsing time can be selected to implement button optimization of 5G messages for industry chat robots, thereby maximizing the delivery effect of 5G messages.
[0017] The above description is only an overview of the technical results of the embodiments of the present invention. In order to more clearly understand the technical means of the embodiments of the present invention, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiments of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0019] Figure 1 A schematic diagram showing a flow chart of a message button optimization method provided by an embodiment of the present invention is shown;
[0020] Figure 2 A schematic diagram of a message button topology diagram of a message button optimization method provided by an embodiment of the present invention is shown;
[0021] Figure 3 A schematic structural diagram of a message button optimization model of a message button optimization method provided by an embodiment of the present invention is shown;
[0022] Figure 4 It shows a schematic structural diagram of a message button optimization device provided by an embodiment of the present invention;
[0023] Figure 5 A schematic structural diagram of a computing device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0024] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0025] For industry customers, 5G messaging provides enhanced messaging services between individuals and applications, realizing "messaging as a service" and introducing a new messaging interaction mode - chatbots. Users can intuitively and conveniently enjoy various 5G application services such as bill payment and top-up, ticket booking, hotel reservations, logistics inquiries, restaurant reservations, and takeout orders within the message window. Chatbots are a service provided by industry customers to end users in the form of conversations. This service is usually based on artificial intelligence software, simulating human intelligent conversations to provide users with specific service functions.
[0026] The MaaP system (including the MaaP platform management module and the MaaP platform) is the core network element of the industry's 5G messaging service. This network element will provide industry users with 5G commercial messaging service access and message uplink and downlink capabilities, and provide users with industry chatbot search, detail query, message uplink and downlink and other functions.
[0027] The 5G messaging application open platform can help industry customers implement application-to-person (A2P) communication in multiple scenarios on demand. Enterprises can use this 5G messaging application open platform to quickly deploy messaging applications without the need for complex code development, helping industry customers to easily and conveniently create their own 5G messaging applications.
[0028] Figure 1 The flow chart of the message button optimization method provided by the embodiment of the present invention is shown. The message button optimization method is applied to the server side, such as Figure 1 As shown, the message button optimization method includes:
[0029] Step S11: Obtain user usage data of each message button at the current moment from the MaaP platform, and obtain feature data of each message button based on the user usage data, wherein the user usage data at least includes the number of message button clicks and browsing time.
[0030] The MaaP system (including the MaaP platform management module and the MaaP platform) is the core network element of the industry's 5G messaging service. This network element will provide industry users with access to 5G commercial messaging services and message uplink and downlink capabilities, and provide users with functions such as industry chatbot search, detail query, and message uplink and downlink. In step S11, optionally, user usage data for each message button used by the chatbot to send 5G messages at the current moment is obtained from the MaaP platform; a message button topology map is constructed based on the user usage data for each message button; the numerical attributes of the user usage data for each message button in the message button topology map are standardized, and the logical relationships between the message buttons are one-hot encoded to obtain feature data for each message button.
[0031] In the embodiment of the present invention, after obtaining the user usage data of each message button, each message button is regarded as a separate agent and a node in the constructed message button topology map. Each message button is regarded as a separate agent, and multiple message buttons constitute a multi-agent environment. By constructing a message button topology map, the multi-agent environment is represented by a message button topology map. Figure 2 As shown in the constructed message button topology diagram, each node is a message button, each edge is a logical relationship between message buttons, and each node is characterized by the user usage data of the corresponding message button. The logical relationship between message buttons includes parallel relationships, primary and secondary relationships, etc.
[0032] Step S12: Processing the characteristic data of any of the message buttons through a message button optimization model and outputting the modification action of the message button.
[0033] The 5G messaging application open platform helps industry customers implement application-to-person (A2P) communication on demand across multiple scenarios. Enterprises can quickly deploy messaging applications through the platform without complex code development, making it easy and convenient for industry customers to create their own 5G messaging applications. Reinforcement learning is an important machine learning method that incorporates three elements: state, action, and reward. An agent takes actions based on its current state, receives rewards, and then refines these actions, ensuring it can perform even better the next time it encounters the same state. Q-Learning is a value-based reinforcement learning algorithm that focuses on training a critic. Q (s, a) represents the expected reward for taking action a (a∈A) at a given state s (s∈S). The environment responds with a corresponding reward r based on the agent's actions. DQN (Deep Q-network) combines Q-Learning with deep learning, using a deep neural network to represent the value function to predict Q values, and by continuously updating the neural network, it learns the optimal action path. The goal of DQN is to learn a policy π that maximizes the discounted cumulative reward over T time steps: The Q function can be defined as the expectation of the discounted cumulative reward, which is based on the current state and the selected action, and all subsequent actions are taken according to the policy π: We need to find a strategy that maximizes the Q function:
[0034]
[0035] DQN has two neural networks: target-net, which has relatively fixed parameters and is used to obtain the value of Q-target, and eval-net, which is used to obtain the value of Q-evaluation. The Q value is updated according to the following rules:
[0036]
[0037] The relational graph convolutional neural network provides a scoring method for self-encoding and decoding triples (subject-predicate-object). It is a neural network model optimized based on the graph convolution network (GCN). The essential purpose of GCN is to extract the spatial features of topological graphs. The original GCN processes single relations between datasets with topological structures, while R-GCN can handle multi-relational data features in knowledge bases.
[0038] The embodiment of the present invention introduces multi-agent topology graph learning based on the traditional DQN model, and regards each message button as a separate agent. Multiple message buttons constitute a multi-agent environment. In a multi-agent environment, if all agents are considered, it will take a lot of time and computing costs. Therefore, the embodiment of the present invention constructs a message button topology graph to represent the multi-agent environment with a message button topology graph. Each node in the topology graph is a message button, and each edge in the topology graph is the logical relationship between message buttons. The feature of each node is the user usage data of the message button. Therefore, the message button agent of the embodiment of the present invention only needs to consider the influence from neighboring agents, which improves the shortcomings of traditional multi-agents such as large computational complexity and information redundancy.
[0039] Before step S12, the message button topology diagram of historical moments and the manually annotated 5G message button modification actions of the user usage data corresponding to each message button are collected from the MaaP platform as a total data set. Each node in the message button topology diagram is a message button, and each edge in the message button topology diagram is a logical relationship between message buttons (including parallel, primary and secondary, etc.). The message button topology diagram can be represented as a directed, unweighted graph G = (V, E, R), where E is a set of edges. V is a set of nodes V = {V1, V2, V3, ..., V N}, e ij Represents node V i and node V j The connection relationship between ij =(v i ,l,v j ), l is the type of edge. The dataset includes: message button topology adjacency matrix, e ij Indicates the network node V i and networking node V j The connection relationship between ij =(v i ,l,v j ), l is the type of edge.
[0040] For the total data set, the message button topology graph of the historical moments in the total data set and the corresponding modification actions are preprocessed. The message button node feature vector is S = {s1, s2, ..., s N}: Each node is characterized by user usage data for the message button. User usage data includes the number of clicks on each button and the reading time after entering the button. All numerical attributes are normalized: (X-mean) / std. The calculation is performed separately for each dimension, and the data is subtracted from its mean by attribute (column-by-column) and divided by its variance. Standardization will improve the convergence speed and accuracy of the model. The manually labeled modification action (a) corresponding to the message button includes three automatic actions: button position, button style, and button name. All button positions, button styles, and button names need to be one-hot encoded in advance, and the value of the action is the corresponding encoding. The total dataset is divided into training data and test data, with 80% of the entire dataset used as training data and the remaining 20% as test data. The training set is used to train the message button optimization model, and the test set is used to evaluate and verify the message button optimization model.
[0041] Finally, the pre-processed total data set is used to train the message button optimization model. Figure 3 As shown in Figure 2, the message button optimization model is divided into three parts: state encoder (composed of relational graph convolutional layers), action encoder (composed of fully connected layers), and Q network. Utilizing the ability of relational graph convolutional networks (R-GCN) to process multi-relational topology data, the state value o of the message button agent i is t i After passing through the relationship graph convolution layer of the state encoder, it is fused with the message button agent with neighbor relationship to generate the potential feature vector z t i Finally, the Q network modifies the action and state output Q value according to the encoded message button, compares it with the target Q value and calculates the error, gradually making the model maximize the sum of the rewards of all message button agents.
[0042] The state encoder inputs the feature vector s of each message button node i , the encoder maps each node to a latent feature representation z of dimension d i For each node, the expression of each layer of the neural network is as follows:
[0043]
[0044] Among them, h i k is the node v in the kth layer of the neural network i The hidden state, r is the relationship type, W rk is the parameter matrix of the kth neural network layer with a specific relation type. The activation function is ReLU(·)=max(0,·), where ReLU is a rectified linear unit. The input of the first layer is the node feature vector x i =h i 0 , if K layers are stacked, the final output of the encoder is z i =h i K In ordinary GCN, D' -1 / 2 A'D' -1 / 2 It is the symmetric normalization of the adjacency matrix A, A'=A+I, D' is the node degree diagonal matrix of A', and for a single node in the embodiment of the present invention, normalization is done by dividing it by the degree of its node, so that the value of each adjacent edge information transmission is normalized, and the influence of the former node will not be greater than that of the latter node due to the fact that the former node has more edges and the latter node has fewer edges. Therefore, 1 / c i,r Equivalent to the normalization of the adjacency matrix in GCN, c i,r is a regularization constant, choose c i,r =|N ir |, N ir represents the neighbor set of node i under relation r.
[0045] The encoder includes: the first layer is the first input layer, which is used to input the feature vector s of node i corresponding to the moment message button i . The second layer is the first relational graph convolution layer (R-GCN), the number of convolution kernels is 128 (i.e., the dimension of the output), and the activation function is set to "relu". The third layer is the first random dropout layer (dropout), and the dropout probability is set to 0.2. The input neurons will be randomly disconnected with a certain probability (20%) each time the parameters are updated during training. The dropout layer is used to prevent overfitting. The fourth layer is the second relational graph convolution layer (R-GCN), the number of convolution kernels is 64, and the activation function is set to "lamda". The fourth layer is the second random dropout layer (dropout): the dropout probability is set to 0.2.
[0046] The second input layer of the action encoder is used to input the 5G message button setting modification action (a) corresponding to the message user usage data. After passing through the first fully connected layer (Dense) and the second fully connected layer, 16 and 8 neurons are set respectively, and the activation function is "relu".
[0047] The Q network uses a merge layer to merge the features s and a extracted from multiple connection layers, and passes them through the third fully connected layer (Dense) and the fourth connected layer, with 64 and 32 neurons respectively, and the activation function is "Relu". A random dropout layer (dropout) is set after each fully connected layer, with a dropout probability set to 0.2. During the training process, the input neurons will be randomly disconnected with a certain probability (20%) each time the parameters are updated to prevent overfitting. The output layer consists of 1 fully connected neuron, which outputs the Q value of the selected 5G message button setting modification action.
[0048] During the training process, (S, A, S', R, C) is stored in the experience replay buffer at each time step (moment). When training the neural network, a random small batch is selected from it to update instead of using the most recent one, thus solving the problem of correlation between samples and greatly improving the stability of the system. The environment state set is represented by S = {s1, s2, ..., s N}, where the action set is represented by A={a1,a2,…,a N}, where the next set of environmental states is represented by S' = {s'1, s'2, ..., s' N}, where the reward set is represented by R = {r1, r2, …, r N}, where the adjacency matrix set is represented as C = {C1, C2, ..., C N}, randomly take out a batch of data of size M from the experience replay, and minimize the error function:
[0049]
[0050] For each agent i, the input to the Q function is O i,c 、a i , output Q value, update the parameters of Q function neural network so that Q value is as close to target value y as possible, where target value y i for:
[0051] y i =r i +γmax a′ Q(O′ i,c ,a′ i ;θ′)
[0052] γ is the discount factor, θ is the parameter of the Q function. The agent performs action a i After that, the multi-agent topology is updated to the state of the next time step, and then the target network is updated by θ'=βθ+(1-β)θ', and the updated Q function neural network weights are assigned to Q target =Q.
[0053] After obtaining the trained message button optimization model, in step S12, optionally, the state value in the feature data of any of the message buttons is passed through the relationship graph convolution layer of the state encoder in the message button optimization model, and the state value is fused with the message button agent with a neighbor relationship to generate a potential feature vector; the modification action of the message button is output according to the potential feature vector.
[0054] Step S13: Optimizing the message button based on the modification action, and obtaining a reward for the optimized message button being implemented for a preset time.
[0055] In an embodiment of the present invention, optionally, the 5G message open platform is applied to optimize the message button according to the modification action; a 5G message is sent to the terminal user based on the optimized message button through a chat robot; new user usage data sent by the terminal user based on the optimized message button for a preset time is obtained; and a reward function is used to calculate a reward for the optimized message button for a preset time based on the new user usage data.
[0056] Step S14: returning the feedback to the message button optimization model for model tuning, and selecting a button optimization action that maximizes the number of message button clicks and browsing time.
[0057] The method for returning the reward to the message button optimization model for model tuning is similar to the training method of the message button optimization model and will not be repeated here. Finally, the button optimization action that maximizes the number of message button clicks and the viewing time is selected.
[0058] After step S14, the new user usage data is fed back to the MaaP platform as the user usage data of the message button at the next moment. That is, the 5G message terminal user feeds back the 5G message button usage data sent by the chatbot at time step t+1 to the MaaP platform.
[0059] This embodiment of the present invention introduces multi-agent topology graph learning based on the traditional DQN model. It represents the multi-agent environment consisting of multiple message buttons in 5G messaging using a message button topology graph. Each message button is considered a separate agent, and each node in the topology graph is a message button. Each edge in the topology graph represents the logical relationship between message buttons, and each node is characterized by the user usage data for that message button. The message button optimization model is divided into three parts: a state encoder (composed of a relationship graph convolution layer), an action encoder (composed of a fully connected layer), and a Q network. Taking advantage of the ability of the relational graph convolutional network (R-GCN) to process multi-relational topological data, the state value of the message button agent i passes through the relational graph convolution layer of the state encoder and is fused with the message button agents with neighboring relationships to generate a potential feature vector. Finally, the Q network modifies the action and state output Q value according to the encoded message button, compares it with the target Q value and calculates the error, gradually making the message button optimization model maximize the sum of the rewards of all message button agents, that is, it can select the button optimization action that maximizes the number of terminal users' message button clicks and browsing time, thereby realizing 5G message button optimization for industry chatbots and maximizing the delivery effect of 5G messages.
[0060] An embodiment of the present invention obtains user usage data of each message button at the current moment from the MaaP platform, and obtains characteristic data of each message button based on the user usage data, wherein the user usage data at least includes the number of message button clicks and browsing time; the characteristic data of any message button is processed by a message button optimization model and the modification action of the message button is output; the message button is optimized based on the modification action, and a reward for the optimized message button to be implemented for a preset time is obtained; the reward is returned to the message button optimization model for model tuning, and the button optimization action that maximizes the number of message button clicks and browsing time is selected. The button optimization action that maximizes the number of message button clicks and browsing time can be selected to implement button optimization of 5G messages for industry chat robots, thereby maximizing the delivery effect of 5G messages.
[0061] Figure 4 FIG. 1 shows a schematic diagram of the structure of the message button optimization device according to an embodiment of the present invention. Figure 4 As shown, the message button optimization device includes: a data acquisition unit 401, a modification action output unit 402, a reward acquisition unit 403, an action optimization unit 404 and a model training unit 405.
[0062] The data acquisition unit 401 obtains the user usage data of each message button at the current moment from the MaaP platform, and obtains the characteristic data of each message button based on the user usage data, wherein the user usage data at least includes the number of message button clicks and the browsing time; the modification action output unit 402 is used to output the modification action of the message button after processing the characteristic data of any message button through the message button optimization model; the reward acquisition unit 403 is used to optimize the message button based on the modification action, and obtain the reward of the optimized message button for a preset time; the action optimization unit 404 is used to return the reward to the message button optimization model for model tuning, and select the button optimization action that maximizes the number of message button clicks and browsing time.
[0063] In an optional manner, the data acquisition unit 401 is used to: obtain user usage data of each message button used by the chat robot to send 5G messages at the current moment from the MaaP platform; construct a message button topology map based on the user usage data of each message button; standardize the numerical attributes of the user usage data of each message button in the message button topology map, and perform one-hot encoding on the logical relationship between the message buttons to obtain feature data of each message button.
[0064] In an optional manner, the data acquisition unit 401 is used to: regard each message button as a separate intelligent entity, as a node in the constructed message button topology graph; in the constructed message button topology graph, each node is a message button, each edge is a logical relationship between message buttons, and the characteristic of each node is the user usage data of the corresponding message button.
[0065] In an optional manner, the modification action output unit 402 is used to: pass the state value in the feature data of any of the message buttons through the relationship graph convolution layer of the state encoder in the message button optimization model, fuse the state value with the message button agent with a neighbor relationship to generate a potential feature vector; and output the modification action of the message button according to the potential feature vector.
[0066] In an optional manner, the reward acquisition unit 403 is used to optimize the message button based on the modification action, and obtain the reward for the optimized message button to implement the preset time for use in: applying the 5G message open platform to optimize the message button according to the modification action; sending 5G messages to terminal users based on the optimized message button through a chat robot; obtaining new user usage data sent by the terminal user based on the optimized message button to implement the preset time; and calculating the reward for the optimized message button to implement the preset time based on the new user usage data through a reward function.
[0067] In an optional manner, the action optimization unit 404 is further configured to: feed back the new user usage data as user usage data of the message button at a next moment to the MaaP platform.
[0068] In an optional manner, the model training unit 405 includes: collecting message button topology maps of historical moments and manually annotated 5G message button modification actions of user usage data corresponding to each message button from the MaaP platform as a total data set; preprocessing the message button topology maps of historical moments and the corresponding modification actions in the total data set; and applying the preprocessed total data set to train the message button optimization model.
[0069] An embodiment of the present invention obtains user usage data of each message button at the current moment from the MaaP platform, and obtains characteristic data of each message button based on the user usage data, wherein the user usage data at least includes the number of message button clicks and browsing time; the characteristic data of any message button is processed by a message button optimization model and the modification action of the message button is output; the message button is optimized based on the modification action, and a reward for the optimized message button to be implemented for a preset time is obtained; the reward is returned to the message button optimization model for model tuning, and the button optimization action that maximizes the number of message button clicks and browsing time is selected. The button optimization action that maximizes the number of message button clicks and browsing time can be selected to implement button optimization of 5G messages for industry chat robots, thereby maximizing the delivery effect of 5G messages.
[0070] An embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores at least one executable instruction, and the computer executable instruction can execute the message button optimization method in any of the above method embodiments.
[0071] The executable instructions can be used to cause the processor to perform the following operations:
[0072] Obtaining user usage data of each message button at the current moment from the MaaP platform, and obtaining feature data of each message button based on the user usage data, wherein the user usage data includes at least the number of clicks and browsing duration of the message button;
[0073] Processing the characteristic data of any of the message buttons through a message button optimization model and outputting a modification action for the message button;
[0074] Optimizing the message button based on the modification action, and obtaining a reward for the optimized message button being implemented for a preset time;
[0075] The reward is returned to the message button optimization model for model tuning, and a button optimization action that maximizes the number of message button clicks and browsing time is selected.
[0076] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0077] Obtain user usage data for each message button used by the chatbot to send 5G messages at the current moment from the MaaP platform;
[0078] Constructing a message button topology map based on the user usage data of each message button;
[0079] The numerical attributes of the user usage data of each message button in the message button topology are standardized, and the logical relationships between the message buttons are one-hot encoded to obtain feature data of each message button.
[0080] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0081] Treat each message button as a separate agent and as a node in the constructed message button topology graph;
[0082] In the constructed message button topology diagram, each node is a message button, each edge is a logical relationship between message buttons, and the feature of each node is the user usage data of the corresponding message button.
[0083] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0084] Passing the state value in the feature data of any of the message buttons through the relational graph convolution layer of the state encoder in the message button optimization model, fusing the state value with the message button agents that have a neighbor relationship to generate a potential feature vector;
[0085] Outputting a modification action of the message button according to the latent feature vector.
[0086] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0087] Applying the 5G message open platform to optimize the message button according to the modification action;
[0088] Sending a 5G message to the terminal user through the chat robot based on the optimized message button;
[0089] Acquire new user usage data sent by the terminal user based on the optimized message button implementation preset time;
[0090] The reward function is used to calculate the reward of the optimized message button for the preset time according to the new user usage data.
[0091] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0092] The new user usage data is fed back to the MaaP platform as the user usage data of the message button at the next moment.
[0093] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0094] The total dataset is collected from the MaaP platform, including the message button topology map at historical moments and the manually labeled user usage data of 5G message button modification actions corresponding to each message button.
[0095] Preprocessing the message button topology graphs at historical moments and the corresponding modification actions in the total data set;
[0096] The preprocessed total data set is used to train the message button optimization model.
[0097] An embodiment of the present invention obtains user usage data of each message button at the current moment from the MaaP platform, and obtains characteristic data of each message button based on the user usage data, wherein the user usage data at least includes the number of message button clicks and browsing time; the characteristic data of any message button is processed by a message button optimization model and the modification action of the message button is output; the message button is optimized based on the modification action, and a reward for the optimized message button to be implemented for a preset time is obtained; the reward is returned to the message button optimization model for model tuning, and the button optimization action that maximizes the number of message button clicks and browsing time is selected. The button optimization action that maximizes the number of message button clicks and browsing time can be selected to implement button optimization of 5G messages for industry chat robots, thereby maximizing the delivery effect of 5G messages.
[0098] An embodiment of the present invention provides a computer program product, which includes a computer program stored on a computer storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the message button optimization method in any of the above method embodiments.
[0099] The executable instructions can be used to cause the processor to perform the following operations:
[0100] Obtaining user usage data of each message button at the current moment from the MaaP platform, and obtaining feature data of each message button based on the user usage data, wherein the user usage data includes at least the number of clicks and browsing duration of the message button;
[0101] Processing the characteristic data of any of the message buttons through a message button optimization model and outputting a modification action for the message button;
[0102] Optimizing the message button based on the modification action, and obtaining a reward for the optimized message button being implemented for a preset time;
[0103] The reward is returned to the message button optimization model for model tuning, and a button optimization action that maximizes the number of message button clicks and browsing time is selected.
[0104] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0105] Obtain user usage data for each message button used by the chatbot to send 5G messages at the current moment from the MaaP platform;
[0106] Constructing a message button topology map based on the user usage data of each message button;
[0107] The numerical attributes of the user usage data of each message button in the message button topology are standardized, and the logical relationships between the message buttons are one-hot encoded to obtain feature data of each message button.
[0108] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0109] Treat each message button as a separate agent and as a node in the constructed message button topology graph;
[0110] In the constructed message button topology diagram, each node is a message button, each edge is a logical relationship between message buttons, and the feature of each node is the user usage data of the corresponding message button.
[0111] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0112] Passing the state value in the feature data of any of the message buttons through the relational graph convolution layer of the state encoder in the message button optimization model, fusing the state value with the message button agents that have a neighbor relationship to generate a potential feature vector;
[0113] Outputting a modification action of the message button according to the latent feature vector.
[0114] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0115] Applying the 5G message open platform to optimize the message button according to the modification action;
[0116] Sending a 5G message to the terminal user through the chat robot based on the optimized message button;
[0117] Acquire new user usage data sent by the terminal user based on the optimized message button implementation preset time;
[0118] The reward function is used to calculate the reward of the optimized message button for the preset time according to the new user usage data.
[0119] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0120] The new user usage data is fed back to the MaaP platform as the user usage data of the message button at the next moment.
[0121] In an optional manner, the executable instruction causes the processor to perform the following operations:
[0122] The total dataset is collected from the MaaP platform, including the message button topology map at historical moments and the manually labeled user usage data of 5G message button modification actions corresponding to each message button.
[0123] Preprocessing the message button topology graphs at historical moments and the corresponding modification actions in the total data set;
[0124] The preprocessed total data set is used to train the message button optimization model.
[0125] An embodiment of the present invention obtains user usage data of each message button at the current moment from the MaaP platform, and obtains characteristic data of each message button based on the user usage data, wherein the user usage data at least includes the number of message button clicks and browsing time; the characteristic data of any message button is processed by a message button optimization model and the modification action of the message button is output; the message button is optimized based on the modification action, and a reward for the optimized message button to be implemented for a preset time is obtained; the reward is returned to the message button optimization model for model tuning, and the button optimization action that maximizes the number of message button clicks and browsing time is selected. The button optimization action that maximizes the number of message button clicks and browsing time can be selected to implement button optimization of 5G messages for industry chat robots, thereby maximizing the delivery effect of 5G messages.
[0126] Figure 5 The schematic diagram of the structure of the computing device provided by the embodiment of the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the device.
[0127] like Figure 5 As shown, the computing device may include: a processor (processor) 502 , a communications interface (Communications Interface) 504 , a memory (memory) 506 , and a communication bus 508 .
[0128] Processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508. Communication interface 504 is used to communicate with other devices, such as clients or other server network elements. Processor 502 is used to execute program 510, which may specifically perform the steps described in the above-mentioned message button optimization method embodiment.
[0129] Specifically, the program 510 may include program codes, which include computer operation instructions.
[0130] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement an embodiment of the present invention. The one or more processors included in the device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0131] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0132] The program 510 may be specifically configured to enable the processor 502 to perform the following operations:
[0133] Obtaining user usage data of each message button at the current moment from the MaaP platform, and obtaining feature data of each message button based on the user usage data, wherein the user usage data includes at least the number of clicks and browsing duration of the message button;
[0134] Processing the characteristic data of any of the message buttons through a message button optimization model and outputting a modification action for the message button;
[0135] Optimizing the message button based on the modification action, and obtaining a reward for the optimized message button being implemented for a preset time;
[0136] The reward is returned to the message button optimization model for model tuning, and a button optimization action that maximizes the number of message button clicks and browsing time is selected.
[0137] In an optional manner, the program 510 enables the processor to perform the following operations:
[0138] Obtain user usage data for each message button used by the chatbot to send 5G messages at the current moment from the MaaP platform;
[0139] Constructing a message button topology map based on the user usage data of each message button;
[0140] The numerical attributes of the user usage data of each message button in the message button topology are standardized, and the logical relationships between the message buttons are one-hot encoded to obtain feature data of each message button.
[0141] In an optional manner, the program 510 enables the processor to perform the following operations:
[0142] Treat each message button as a separate agent and as a node in the constructed message button topology graph;
[0143] In the constructed message button topology diagram, each node is a message button, each edge is a logical relationship between message buttons, and the feature of each node is the user usage data of the corresponding message button.
[0144] In an optional manner, the program 510 enables the processor to perform the following operations:
[0145] Passing the state value in the feature data of any of the message buttons through the relational graph convolution layer of the state encoder in the message button optimization model, fusing the state value with the message button agents that have a neighbor relationship to generate a potential feature vector;
[0146] Outputting a modification action of the message button according to the latent feature vector.
[0147] In an optional manner, the program 510 enables the processor to perform the following operations:
[0148] Applying the 5G message open platform to optimize the message button according to the modification action;
[0149] Sending a 5G message to the terminal user through the chat robot based on the optimized message button;
[0150] Acquire new user usage data sent by the terminal user based on the optimized message button implementation preset time;
[0151] The reward function is used to calculate the reward of the optimized message button for the preset time according to the new user usage data.
[0152] In an optional manner, the program 510 enables the processor to perform the following operations:
[0153] The new user usage data is fed back to the MaaP platform as the user usage data of the message button at the next moment.
[0154] In an optional manner, the program 510 enables the processor to perform the following operations:
[0155] The total dataset is collected from the MaaP platform, including the message button topology map at historical moments and the manually labeled user usage data of 5G message button modification actions corresponding to each message button.
[0156] Preprocessing the message button topology graphs at historical moments and the corresponding modification actions in the total data set;
[0157] The preprocessed total data set is used to train the message button optimization model.
[0158] An embodiment of the present invention obtains user usage data of each message button at the current moment from the MaaP platform, and obtains characteristic data of each message button based on the user usage data, wherein the user usage data at least includes the number of message button clicks and browsing time; the characteristic data of any message button is processed by a message button optimization model and the modification action of the message button is output; the message button is optimized based on the modification action, and a reward for the optimized message button to be implemented for a preset time is obtained; the reward is returned to the message button optimization model for model tuning, and the button optimization action that maximizes the number of message button clicks and browsing time is selected. The button optimization action that maximizes the number of message button clicks and browsing time can be selected to implement button optimization of 5G messages for industry chat robots, thereby maximizing the delivery effect of 5G messages.
[0159] The algorithm or demonstration provided herein are not inherently relevant to any particular computer, virtual system or other equipment. Various general-purpose systems may also be used together with the teachings based on this. According to the above description, it is apparent that the structure required for constructing this type of system. In addition, the embodiment of the present invention is not directed to any specific programming language yet. It should be understood that various programming languages can be utilized to realize the content of the present invention described herein, and the above description of specific languages is for the purpose of disclosing the best mode of the present invention.
[0160] In the description provided herein, numerous specific details are described. However, it is understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0161] Similarly, it should be understood that in order to streamline the present invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the invention, various features of the embodiments of the invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the claims below, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Accordingly, the claims that follow the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the invention.
[0162] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition may be divided into multiple submodules or subunits or subcomponents. All features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed herein may be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0163] Furthermore, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of the present invention and to form different embodiments. For example, in the claims below, any of the claimed embodiments may be used in any combination.
[0164] It should be noted that the above embodiments illustrate rather than limit the invention, and that alternative embodiments may be devised by a person skilled in the art without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments should not be understood as limiting the order of execution unless otherwise specified.
Claims
1. A message button optimization method, characterized in that: The method comprises: Obtaining user usage data of each message button at the current moment from the MaaP platform, wherein the user usage data includes at least the number of clicks and browsing duration of the message button; A message button topology graph is constructed based on the user usage data of each message button, wherein each message button is regarded as a separate intelligent entity and serves as a node in the constructed message button topology graph, each edge is a logical relationship between message buttons, and each node is characterized by the user usage data of the corresponding message button. The logical relationship between message buttons includes a parallel relationship and a primary-secondary relationship; Normalizing the numerical attributes of the user usage data of each message button in the message button topology diagram, and performing one-hot encoding on the logical relationships between the message buttons to obtain feature data of each message button; Processing the feature data of any of the message buttons through a message button optimization model and outputting a modification action for the message button, wherein a state value in the feature data of any of the message buttons is passed through a relational graph convolution layer of a state encoder in the message button optimization model, and the state value is fused with message button agents having a neighbor relationship to generate a latent feature vector; outputting the modification action for the message button based on the latent feature vector, the modification action including: button position, button style, and button name; Optimizing the message button based on the modification action, and obtaining a reward for the optimized message button being implemented for a preset time; The reward is returned to the message button optimization model for model tuning, and a button optimization action that maximizes the number of message button clicks and browsing time is selected.
2. The method according to claim 1, characterized in that The optimizing the message button based on the modification action and obtaining a reward for the optimized message button for a preset time period includes: Applying the 5G message open platform to optimize the message button according to the modification action; Sending a 5G message to the terminal user through the chat robot based on the optimized message button; Acquire new user usage data sent by the terminal user based on the optimized message button implementation preset time; The reward function is used to calculate the reward of the optimized message button for the preset time according to the new user usage data.
3. The method according to claim 2, characterized in that The step of returning the reward to the message button optimization model for model tuning and selecting the button optimization action that maximizes the number of message button clicks and browsing time includes: The new user usage data is fed back to the MaaP platform as the user usage data of the message button at the next moment.
4. The method according to claim 1, wherein Before the feature data of any of the message buttons is processed by the message button optimization model and the modification action of the message button is output, the method includes: The total dataset is collected from the MaaP platform, including the message button topology map at historical moments and the manually labeled user usage data of 5G message button modification actions corresponding to each message button. Preprocessing the message button topology graphs at historical moments and the corresponding modification actions in the total data set; The preprocessed total data set is used to train the message button optimization model.
5. A message button optimization device, characterized in that: The device comprises: A data acquisition unit is used to obtain user usage data of each message button at the current moment from the MaaP platform, wherein the user usage data includes at least the number of message button clicks and browsing time; construct a message button topology map based on the user usage data of each message button, wherein each message button is regarded as a separate intelligent entity and serves as a node in the constructed message button topology map, each edge is a logical relationship between message buttons, and the feature of each node is the user usage data of the corresponding message button, and the logical relationship between message buttons includes a parallel relationship and a primary-secondary relationship; the numerical attributes of the user usage data of each message button in the message button topology map are standardized, and the logical relationship between message buttons is one-hot encoded to obtain the feature data of each message button; A modification action output unit is configured to process the feature data of any of the message buttons through a message button optimization model and output a modification action for the message button, wherein a state value in the feature data of any of the message buttons is passed through a relational graph convolution layer of a state encoder in the message button optimization model, and the state value is fused with message button agents having a neighbor relationship to generate a potential feature vector; and output the modification action of the message button based on the potential feature vector, the modification action including: button position, button style, and button name; a reward obtaining unit, configured to optimize the message button based on the modification action and obtain a reward for the optimized message button after it is used for a preset time; The action optimization unit is used to return the feedback to the message button optimization model for model tuning, and select the button optimization action that maximizes the number of message button clicks and browsing time.
6. A computing device comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of the message button optimization method according to any one of claims 1-4.
7. A computer storage medium, wherein the storage medium stores at least one executable instruction, wherein the executable instruction enables a processor to execute the steps of the message button optimization method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Video recommendation method and system
CN107515909A