Recommendation Information Determination Method and Apparatus, Electronic Device, and Readable Storage Medium
By constructing the Markov decision model and the backpropagation neural network model, combined with the strong connectivity component algorithm, multiple data sequences are processed and fused, the problem of inconsistent data information from different sources is solved, and the data fusion efficiency and the accuracy of recommended information are improved.
Patent Information
- Application Number
- CN202210659254.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-10
AI Technical Summary
The information expressed by data from different sources over time is inconsistent, which makes it impossible to guarantee the accuracy of recommended information.
By constructing a Markov decision model and a backpropagation neural network model, combining a strong connectivity component algorithm, multiple data sequences are processed and fused, the target data sequence is determined, and the recommended information is determined based on the fused data processing results.
It improves the efficiency of data fusion and the accuracy of recommended information, real-time fusion of multi-source data and digestion and processing of conflicting data.
Smart Images

Figure CN115048561B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of data processing and finance, and more particularly, to a method and apparatus for determining recommendation information, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the development of science and technology, how to comprehensively consider the influence of different factors such as a user's family, work, or market to perform data processing is an urgent problem to be solved.
[0003] In the process of implementing the concept of the present disclosure, the inventors found that at least the following problems exist in the related art: The information expressed by data from different sources may be inconsistent over time, resulting in the inability to guarantee the accuracy of recommendation information. Summary of the Invention
[0004] In view of this, the present disclosure provides a method and apparatus for determining recommendation information, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] According to one aspect of the present disclosure, there is provided a method for determining recommendation information, including:
[0006] Determining a state sequence according to M sequences of data to be processed associated with a target user, where each sequence of data to be processed in the M sequences of data to be processed includes a channel identifier and N data to be processed arranged in chronological order, the channel identifier is used to characterize the source of the N data to be processed, the state sequence includes N target states arranged in chronological order, M is a positive integer greater than 1, and N is a positive integer;
[0007] Processing the M sequences of data to be processed respectively according to the channel identifier and the N target states to obtain P sequences of target data, where P is a positive integer and P is less than or equal to M;
[0008] Performing data fusion processing on the P sequences of target data to obtain a fused data processing result; and
[0009] Determining recommendation information for the target user according to the fused data processing result.
[0010] According to an embodiment of the present disclosure, the determining a state sequence according to M sequences of data to be processed associated with a target user includes:
[0011] Inputting the M sequences of data to be processed into a Markov decision model and outputting the N target states.
[0012] According to an embodiment of the present disclosure, the above method further includes, before determining the state sequence according to the M to-be-processed data sequences associated with the target user:
[0013] Determine a sample action set and a sample state set according to the above channel identifier and the sample data sequence corresponding to the above channel identifier, wherein the above sample action set includes a retention operation and a deletion operation, the above sample state set includes multiple sample states, and each sample state in the above multiple sample states is respectively associated with the above retention operation and the above deletion operation;
[0014] Perform dimensionality reduction processing on the above sample state set and the above sample action set respectively to obtain a dimensionality-reduced sample state set and a dimensionality-reduced sample action set;
[0015] Determine a reward function based on the above dimensionality-reduced sample state set and the above dimensionality-reduced sample action set; and
[0016] Construct the above Markov decision model based on the above dimensionality-reduced sample state set, the above dimensionality-reduced sample action set, and the above reward function.
[0017] According to an embodiment of the present disclosure, the above processing the M to-be-processed data sequences respectively according to the above channel identifier and the above N target states to obtain P target data sequences includes:
[0018] Input the above M to-be-processed data sequences and the above state sequence into a trained backpropagation neural network model, and output the above P target data sequences.
[0019] According to an embodiment of the present disclosure, the above inputting the above M to-be-processed data sequences and the above state sequence into a trained backpropagation neural network model and outputting the above P target data sequences includes:
[0020] Determine an action sequence according to the above channel identifier of each of the above M to-be-processed data sequences and the above state sequence, wherein the above action sequence includes N target actions arranged in chronological order, and the above N target actions correspond to the above channel identifier; and
[0021] Process the to-be-processed data sequence corresponding to the above channel identifier according to the above N target actions to obtain the above P target data sequences.
[0022] According to an embodiment of the present disclosure, the above N target actions include a retention operation or a deletion operation;
[0023] The above processing the to-be-processed data sequence corresponding to the above channel identifier according to the above N target actions to obtain the above P target data sequences includes:
[0024] When the above N target actions include a retention operation, determine the channel identifier to be retained and the time to be retained corresponding to the above retention operation;
[0025] According to the above channel identifier to be retained, determine the data sequence to be retained in the above M data sequences to be processed;
[0026] According to the above time to be retained, retain the above N data to be processed in the above data sequence to be retained;
[0027] When the above N target actions include a deletion operation, determine the channel identifier to be deleted and the time to be deleted corresponding to the above deletion operation;
[0028] According to the above channel identifier to be deleted, determine the data sequence to be deleted in the above M data sequences to be processed; and
[0029] According to the above time to be deleted, delete the above N data to be processed in the above data sequence to be deleted.
[0030] According to an embodiment of the present disclosure, the training method of the above trained backpropagation neural network includes:
[0031] According to the above channel identifier and the sample data sequence corresponding to the above channel identifier, determine the sample state, sample action, and sample reward function;
[0032] Input the above sample data sequence, the above sample state, and the above sample reward function corresponding to the above channel identifier into the backpropagation neural network model to be trained, and output a predicted action;
[0033] Input the above predicted action and the above sample action corresponding to the above channel identifier into a loss function, and output a loss result;
[0034] Adjust the network parameters of the above backpropagation neural network model to be trained according to the above loss result until the above loss function or the number of iterations meets a preset condition; and
[0035] Use the model obtained when the above loss function or the number of iterations meets the preset condition as the above trained backpropagation neural network model.
[0036] According to an embodiment of the present disclosure, the above data fusion processing of the P target data sequences to obtain a fused data processing result includes:
[0037] Based on a strongly connected component algorithm, perform data fusion processing on the above P target data sequences to obtain the above fused data processing result.
[0038] According to an embodiment of the present disclosure, the data fusion processing of the above P target data sequences based on the strongly connected component algorithm to obtain the above data processing result after fusion includes:
[0039] Determine a data sequence set, a data sequence related combination, and a superset of the above data sequence related combination according to the above P target data sequences;
[0040] When any data sequence combination in the above superset belongs to the above data sequence related combination, add the above superset to the candidate data sequence related combination;
[0041] Calculate the relevant information entropy of any data sequence combination in the above candidate data sequence related combination; and
[0042] According to the above relevant information entropy, perform data fusion processing on the above P target data sequences to obtain the above data processing result after fusion.
[0043] According to an embodiment of the present disclosure, the determination of the recommendation information for the above target user based on the above data processing result after fusion includes:
[0044] Determine the risk preference value of the above target user according to the above data processing result after fusion;
[0045] Determine the risk type of the above target user according to the above risk preference value; and
[0046] Determine the above recommendation information for the above target user according to the above risk type.
[0047] According to another aspect of the present disclosure, a recommendation information determination device is provided, including:
[0048] A first determination module, configured to determine a state sequence according to M to-be-processed data sequences associated with a target user, where each to-be-processed data sequence in the above M to-be-processed data sequences includes a channel identifier and N to-be-processed data arranged in chronological order, the above channel identifier is used to characterize the source of the above N to-be-processed data, the state sequence includes N target states arranged in chronological order, M is a positive integer greater than 1, and N is a positive integer;
[0049] A first processing module, configured to process each of the above M to-be-processed data sequences respectively according to the above channel identifier and the above N target states to obtain P target data sequences, where P is a positive integer and P is less than or equal to M;
[0050] A second processing module, configured to perform data fusion processing on the above P target data sequences to obtain a data processing result after fusion; and
[0051] A second determination module, configured to determine recommendation information for the target user according to the processed result of the fused data described above.
[0052] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0053] One or more processors;
[0054] A memory, configured to store one or more instructions,
[0055] wherein, when the one or more instructions are executed by the one or more processors, the one or more processors are caused to implement the method as described in the present disclosure.
[0056] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, on which executable instructions are stored, and when the executable instructions are executed by a processor, the processor is caused to implement the method as described in the present disclosure.
[0057] According to another aspect of the present disclosure, there is provided a computer program product, the computer program product includes computer-executable instructions, and the computer-executable instructions are used to implement the method as described in the present disclosure when being executed.
[0058] According to the embodiments of the present disclosure, since the state sequence is determined according to a plurality of data sequences to be processed associated with the target user, and the target data sequence obtained by processing is subjected to data fusion to obtain the processed result of the fused data, at least partially, the technical problem in the related art that real-time fusion of data cannot be achieved is overcome, thereby improving the efficiency of data fusion. In addition, since the recommendation information for the target user is determined according to the processed result of the fused data, the efficiency and accuracy of determining the recommendation information are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become clearer. In the drawings:
[0060] Figure 1 Schematically shows a system architecture to which the recommendation information determination method according to the embodiments of the present disclosure can be applied;
[0061] Figure 2 Schematically shows a flowchart of the recommendation information determination method according to the embodiments of the present disclosure;
[0062] Figure 3 Schematically shows an exemplary diagram of the process of constructing a Markov decision model according to the embodiments of the present disclosure;
[0063] Figure 4Schematically shows an example diagram of the training process of a backpropagation neural network model according to an embodiment of the present disclosure;
[0064] Figure 5 Schematically shows an example diagram of obtaining a target data sequence according to an embodiment of the present disclosure;
[0065] Figure 6 Schematically shows an example diagram of obtaining a processed result of the fused data according to an embodiment of the present disclosure;
[0066] Figure 7 Schematically shows a block diagram of a recommendation information determination device according to an embodiment of the present disclosure; and
[0067] Figure 8 Schematically shows a block diagram of an electronic device suitable for implementing the recommendation information determination method according to an embodiment of the present disclosure. Detailed implementation manners
[0068] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0069] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0070] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0071] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0072] In the technical solutions of the present disclosure, the acquisition, storage, application, etc. of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0073] In the technical solutions of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user has been obtained.
[0074] With the development of science and technology, how to comprehensively consider the influence of different factors such as the user's family, work, or market, etc. to perform data processing is an urgent problem to be solved.
[0075] Data fusion technology refers to a data processing technology that uses a computer to analyze, synthesize, and combine data from multiple sources obtained in chronological order under certain criteria to complete the required decision-making and evaluation tasks. Through data fusion, the original scattered and independent multiple data can be fused together, thereby discovering the laws and trends of the data to enhance the data value.
[0076] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in the related art: The information expressed by data from different sources may be inconsistent over time, resulting in the inability to achieve real-time fusion of the data.
[0077] To at least partially solve the technical problems existing in the related art, the present disclosure provides a method and apparatus for determining recommended information, an electronic device, and a readable storage medium, which can be applied to the technical field of data processing and the financial field. The method for determining recommended information includes: determining a state sequence according to M sequences of data to be processed associated with a target user, where each sequence of data to be processed in the M sequences of data to be processed includes a channel identifier and N pieces of data to be processed arranged in chronological order, the channel identifier is used to represent the source of the N pieces of data to be processed, the state sequence includes N target states arranged in chronological order, M is a positive integer greater than 1, and N is a positive integer; processing each of the M sequences of data to be processed according to the channel identifier and the N target states to obtain P sequences of target data, where P is a positive integer and P is less than or equal to M; performing data fusion processing on the P sequences of target data to obtain a data processing result after fusion; and determining recommended information for the target user according to the data processing result after fusion.
[0078] It should be noted that the method and apparatus for determining recommended information provided in the embodiments of the present disclosure can be used in the technical field of data processing and the financial field. For example, it can be applied to recommending financial products to a target user. The method and apparatus for determining recommended information provided in the embodiments of the present disclosure can also be used in any field other than the technical field of data processing and the financial field, such as data fusion. The application fields of the method and apparatus for determining recommended information provided in the embodiments of the present disclosure are not limited.
[0079] Figure 1 Schematically shows a system architecture to which the method for determining recommended information according to an embodiment of the present disclosure can be applied. It should be noted that Figure 1 The illustration is only an example of a system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0080] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0081] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0082] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, and so on.
[0083] The server 105 can be a server that provides various services. For example, it can be a background management server (only for illustration) that supports the websites browsed by users using the terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0084] It should be noted that the method for determining recommended information provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the device for determining recommended information provided by the embodiments of the present disclosure can generally be set in the server 105. The method for determining recommended information provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the device for determining recommended information provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Or, the method for determining recommended information provided by the embodiments of the present disclosure can also be executed by the terminal devices 101, 102, or 103, or can also be executed by other terminal devices different from the terminal devices 101, 102, or 103. Correspondingly, the device for determining recommended information provided by the embodiments of the present disclosure can also be set in the terminal devices 101, 102, or 103, or set in other terminal devices different from the terminal devices 101, 102, or 103.
[0085] For example, the data sequence to be processed can be originally stored in any one of the terminal devices 101, 102, or 103 (for example, the terminal device 101, but not limited to this), or stored on an external storage device and can be imported into the terminal device 101. Then, the terminal device 101 can execute the method for determining recommended information provided by the embodiments of the present disclosure locally, or send the data sequence to be processed to other terminal devices, servers, or server clusters, and the other terminal devices, servers, or server clusters that receive the data sequence to be processed execute the method for determining recommended information provided by the embodiments of the present disclosure.
[0086] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0087] Figure 2 A flowchart schematically showing a method for determining recommendation information according to an embodiment of the present disclosure is shown.
[0088] As Figure 2 shown, the recommendation information determination method 200 may include operations S210 to S240.
[0089] In operation S210, a state sequence is determined according to M sequences of data to be processed associated with a target user. Each sequence of data to be processed in the M sequences of data to be processed includes a channel identifier and N data to be processed arranged in chronological order. The channel identifier is used to characterize the source of the N data to be processed. The state sequence includes N target states arranged in chronological order. M is a positive integer greater than 1, and N is a positive integer.
[0090] In operation S220, according to the channel identifier and the N target states, the M sequences of data to be processed are respectively processed to obtain P sequences of target data. P is a positive integer, and P is less than or equal to M.
[0091] In operation S230, data fusion processing is performed on the P sequences of target data to obtain a data processing result after fusion.
[0092] In operation S240, recommendation information for the target user is determined according to the data processing result after fusion.
[0093] According to an embodiment of the present disclosure, the data to be processed may include data related to the risk preference of the target user. Multi-source data from multiple data sources from different channels may be obtained as the data to be processed. The multi-source data may include, for example, income, expenditure, assets, and credit records. The source of the data to be processed may be selected and set by the user according to the data fusion requirement.
[0094] According to an embodiment of the present disclosure, the sequence of data to be processed may include a channel identifier and N data to be processed arranged in chronological order. The M sequences of data to be processed may be used to characterize data from different data sources, and may be expressed as: {D1, D2,..., D i}, D i represents the data of the i-th source (i is greater than or equal to 2).
[0095] According to an embodiment of the present disclosure, the state sequence may include a channel identifier and N target states arranged in chronological order. The target state may be used to characterize the expected data fusion result at that moment. Processing actions to be taken for the data to be processed for different channels and different moments may be determined based on the channel identifier and the N target states. The data to be processed in the M data sequences to be processed may be respectively processed according to the processing actions to obtain P target data sequences.
[0096] According to an embodiment of the present disclosure, data fusion processing may include data layer fusion, feature layer fusion, and decision layer fusion. Data layer fusion may directly perform fusion on the collected raw data layer, and comprehensive analysis and processing of the data may be carried out before the data to be processed from different channels is preprocessed. Feature layer fusion may first perform feature extraction on the data to be processed from different channels, and then perform comprehensive analysis and processing on the feature information.
[0097] According to an embodiment of the present disclosure, multi-source data with different channels and features may be used to analyze, sort, and fuse the data, and recommendation information for the target user may be determined based on the processed result of the fused data. The recommendation information may include, for example, link information and product attribute information. The link information may be used to jump to the display interface of the commodity, and the product attribute information includes information related to products with a relatively high purchase rate for the target user. The link information may include, but is not limited to: web links, application links, or specific location links, etc.
[0098] According to an embodiment of the present disclosure, since the state sequence is determined based on multiple data sequences to be processed associated with the target user, data fusion is performed on the obtained target data sequences, and the processed result of the fused data is obtained, at least partially overcoming the technical problem in the related art that real-time fusion of data cannot be achieved, thereby improving the efficiency of data fusion. In addition, since the recommendation information for the target user is determined based on the processed result of the fused data, the efficiency and accuracy of determining the recommendation information are improved.
[0099] The following refers to Figures 3 - 6 , and further describes the method shown in Figure 2 in combination with specific embodiments.
[0100] According to an embodiment of the present disclosure, operation S240 may include the following operations.
[0101] Determine the risk preference value of the target user according to the processed result of the fused data. Determine the risk type of the target user according to the risk preference value. Determine the recommendation information for the target user according to the risk type.
[0102] According to an embodiment of the present disclosure, relevant data of a target user can be fused online in real time. Based on the processing result of the fused data, a risk preference of the target user with higher credibility can be obtained, and the risk type of the target user can be determined, so as to determine recommendation information for the target user.
[0103] Figure 3 FIG. schematically shows an example diagram of a process of constructing a Markov decision model according to an embodiment of the present disclosure.
[0104] According to an embodiment of the present disclosure, operation S210 may include the following operations.
[0105] Input M sequences of data to be processed into the Markov decision model, and output N target states.
[0106] According to an embodiment of the present disclosure, a Markov decision process (MDP) is a mathematical model for sequential decision-making, which can be used to simulate the random strategy and reward achieved by an agent in an environment where the system state has the Markov property. Based on the channel identifier and the sample data sequence corresponding to the channel identifier, a Markov decision process can be defined by defining a sample action set, a sample state set, and a reward function, and a Markov decision model can be established on the basis of defining the Markov decision process.
[0107] According to an embodiment of the present disclosure, before operation S210, the recommendation information determination method 200 may further include the following operations.
[0108] Determine a sample action set and a sample state set according to the channel identifier and the sample data sequence corresponding to the channel identifier. The sample action set includes a retention operation and a deletion operation, and the sample state set includes a plurality of sample states, and each sample state in the plurality of sample states is respectively associated with the retention operation and the deletion operation. Perform dimensionality reduction processing on the sample state set and the sample action set respectively to obtain a reduced-dimensional sample state set and a reduced-dimensional sample action set. Based on the reduced-dimensional sample state set and the reduced-dimensional sample action set, determine a reward function. Based on the reduced-dimensional sample state set, the reduced-dimensional sample action set, and the reward function, construct a Markov decision model.
[0109] According to an embodiment of the present disclosure, since the data information amounts of multiple data sources from different channels are different, when processing data from different sources, different action selections need to be made, so that when there is conflicting information, the conflicting information can be resolved through different action selections to ensure the effectiveness of the fusion result. The sample action set can be defined as: A = {a1, a2}, where a1 represents a retention operation and a2 represents a deletion operation, and a retention operation or a deletion operation can be taken for data from different sources according to the actual situation.
[0110] According to an embodiment of the present disclosure, since the data fusion result changes after a certain sample action is taken on data from different sources at different times, the fusion result at this moment is used as the sample state at this moment. The sample state set can be defined as: S = {s1, s2,..., s t ,...}, where a t represents the action taken at time t, m t represents the fusion result at time t + 1 obtained after retaining a certain data at time t, and n t represents the fusion result at time t + 1 obtained after deleting a certain data at time t.
[0111] According to an embodiment of the present disclosure, a reward function R can be defined. The reward function represents the reward value or penalty value in the case of a certain action a1 or a2 and a certain state s t .
[0112] According to an embodiment of the present disclosure, based on the trained Markov decision model, the Markov decision model can be used to predict the target state. For example, M data sequences to be processed can be input into the Markov decision model to predict the target state at the current moment.
[0113] As Figure 3 shown, the sample action set 303 and the sample state set 304 can be determined according to the channel identifier 301 and the sample data sequence 302 corresponding to the channel identifier. The sample action set 303 is dimensionally reduced to obtain the dimensionally reduced sample action set 305, and the sample state set is dimensionally reduced to obtain the dimensionally reduced sample state set 306. Based on the dimensionally reduced sample action set 305 and the dimensionally reduced sample state set 306, the reward function 307 can be determined. Based on the dimensionally reduced sample action set 305, the dimensionally reduced sample state set 306, and the reward function 307, a Markov decision model 308 can be constructed.
[0114] According to an embodiment of the present disclosure, since a Markov decision model is pre-constructed based on the sample action set, the sample state set, and the reward function, and the data sequence to be processed is input into the pre-constructed Markov decision model to output the target state, the state prediction for the data sequence to be processed is realized, and the efficiency and accuracy of data processing are improved.
[0115] Figure 4 An example diagram schematically shows the training process of the backpropagation neural network model according to an embodiment of the present disclosure.
[0116] According to an embodiment of the present disclosure, the training method of the trained backpropagation neural network may include the following operations.
[0117] Determine a sample state, a sample action, and a sample reward function according to a channel identifier and a sample data sequence corresponding to the channel identifier. Input the sample data sequence, the sample state, and the sample reward function corresponding to the channel identifier into the backpropagation neural network model to be trained, and output a predicted action. Input the predicted action and the sample action corresponding to the channel identifier into a loss function, and output a loss result. Adjust the network parameters of the backpropagation neural network model to be trained according to the loss result until the loss function or the number of iterations meets a preset condition. Use the model obtained when the loss function or the number of iterations meets the preset condition as the trained backpropagation neural network model.
[0118] As Figure 4 shown, the sample data sequence 401, the sample state 402, and the sample reward function 403 corresponding to the channel identifier can be input into the backpropagation neural network model 404 to be trained, and a predicted action 405 is output. Input the predicted action 405 and the sample action 406 corresponding to the channel identifier into the loss function 407, and output a loss result 408. Adjust the network parameters of the backpropagation neural network model 404 to be trained according to the loss result 408 until the loss function 407 or the number of iterations meets a preset condition. Use the model obtained when the loss function 407 or the number of iterations meets the preset condition as the trained backpropagation neural network model.
[0119] Figure 5 Schematically shows an example schematic diagram of obtaining a target data sequence according to an embodiment of the present disclosure.
[0120] According to an embodiment of the present disclosure, operation S220 may include the following operations.
[0121] Input M to-be-processed data sequences and a state sequence into the trained backpropagation neural network model, and output P target data sequences.
[0122] According to an embodiment of the present disclosure, inputting M to-be-processed data sequences and a state sequence into the trained backpropagation neural network model and outputting P target data sequences may include the following operations.
[0123] Determine an action sequence according to the channel identifier and the state sequence of each of the M to-be-processed data sequences. The action sequence includes N target actions arranged in chronological order, and the N target actions correspond to the channel identifier. Process the to-be-processed data sequence corresponding to the channel identifier according to the N target actions to obtain P target data sequences.
[0124] According to an embodiment of the present disclosure, the N target actions include a retention operation or a deletion operation.
[0125] According to an embodiment of the present disclosure, processing the data sequence to be processed corresponding to the channel identifier according to N target actions to obtain P target data sequences may include the following operations.
[0126] When the N target actions include a retention operation, determine the channel identifier to be retained and the time to be retained corresponding to the retention operation. According to the channel identifier to be retained, determine the data sequence to be retained among the M data sequences to be processed. According to the time to be retained, retain N data to be processed in the data sequence to be retained. When the N target actions include a deletion operation, determine the channel identifier to be deleted and the time to be deleted corresponding to the deletion operation. According to the channel identifier to be deleted, determine the data sequence to be deleted among the M data sequences to be processed. And according to the time to be deleted, delete N data to be processed in the data sequence to be deleted.
[0127] According to an embodiment of the present disclosure, after obtaining M data sequences to be processed and the state sequence, the M data sequences to be processed and the state sequence may be processed to obtain P target data sequences. The model obtained by training the backpropagation neural network model to be trained using the sample data sequence, the sample state, and the sample reward function may be used to process the M data sequences to be processed and the state sequence to obtain P target data sequences.
[0128] According to an embodiment of the present disclosure, the optimal policy may be determined through the backpropagation neural network model, so as to implement different actions for data from different sources and make the optimal policy selection in the presence of conflicting data. For example, at time t, the system state is s t , for the data sequence to be processed with the channel identifier Q1, the predicted action to be executed at the current moment may be obtained according to the policy calculated by the trained backpropagation neural network model, and after performing the predicted action on the data to be processed, it is transferred to the next target state. After finally processing all source data, P target data sequences are obtained.
[0129] As Figure 5 shown, the action sequence 503 may be determined according to the respective channel identifiers 501 and state sequences 502 of the M data sequences to be processed. The action sequence 503 includes N target actions arranged in chronological order, and the N target actions correspond to the channel identifier 501. In step S510, each target action among the N target actions may be determined to be a retention operation or a deletion operation.
[0130] When N target actions include a retention operation, the to-be-retained channel identifier 504 and the to-be-retained moment 505 corresponding to the retention operation can be determined. According to the to-be-retained channel identifier 504, the to-be-retained data sequence 508 is determined from M to-be-processed data sequences. According to the to-be-retained moment 505, N to-be-processed data are retained in the to-be-retained data sequence 508. When N target actions include a deletion operation, the to-be-deleted channel identifier 506 and the to-be-deleted moment 507 corresponding to the deletion operation are determined. According to the to-be-deleted channel identifier 506, the to-be-deleted data sequence 509 is determined from M to-be-processed data sequences. And according to the to-be-deleted moment 507, N to-be-processed data are deleted from the to-be-deleted data sequence. After processing the to-be-processed data sequence corresponding to the channel identifier 501, P target data sequences 510 can be obtained.
[0131] According to an embodiment of the present disclosure, since the to-be-trained backpropagation neural network model is trained based on the sample data sequence, the sample state, and the sample reward function corresponding to the channel identifier, and the to-be-processed data sequence and the state sequence are input into the trained backpropagation neural network model to output the target data sequence, real-time interaction of actual data is realized, and real-time resolution processing of conflicting data from multiple sources can be performed.
[0132] Figure 6 Schematically shows an example schematic diagram of obtaining the data processing result after fusion according to an embodiment of the present disclosure.
[0133] According to an embodiment of the present disclosure, operation S230 may include the following operations.
[0134] Based on the strongly connected component algorithm, data fusion processing is performed on P target data sequences to obtain the data processing result after fusion.
[0135] According to an embodiment of the present disclosure, based on the strongly connected component algorithm, performing data fusion processing on P target data sequences to obtain the data processing result after fusion may include the following operations.
[0136] According to P target data sequences, a data sequence set, a data sequence-related combination, and a superset of the data sequence-related combination are determined. When any data sequence combination in the superset belongs to the data sequence-related combination, the superset is added to the candidate data sequence-related combination. The relevant information entropy of any data sequence combination in the candidate data sequence-related combination is calculated. According to the relevant information entropy, data fusion processing is performed on P target data sequences to obtain the data processing result after fusion.
[0137] According to an embodiment of the present disclosure, for example, i and j represent the channel identifiers of two to-be-processed data sequences associated with the target user, and the fusion equation can be expressed by the following formulas (1) and (2).
[0138]
[0139]
[0140] where m = i, j, represents the source data, P m represents the covariance matrix, P represents the data processing result after fusion, and the fusion equation still holds when m is n.
[0141] For example Figure 6 As shown, based on the P target data sequences 601, the data sequence set 602 and the data sequence correlation combination 603 can be determined. The superset 604 of the data sequence correlation combination can be determined according to the data sequence correlation combination 603. In operation S610, it can be determined whether any data sequence combination in the superset 604 of the data sequence correlation combination belongs to the data sequence correlation combination.
[0142] When it is determined that any data sequence combination in the superset 604 of the data sequence correlation combination belongs to the data sequence correlation combination, the superset 604 of the data sequence correlation combination can be added to the candidate data sequence correlation combination 605. Calculate the correlation information entropy 606 of any data sequence combination in the candidate data sequence correlation combination 605. Based on the correlation information entropy 606, data fusion processing can be performed on the P target data sequences 601 to obtain the data processing result 607 after fusion.
[0143] According to the embodiments of the present disclosure, since the data processing result after fusion is obtained by performing data fusion processing on the target data sequences based on the strongly connected component algorithm, it can effectively utilize the data of different information sources to estimate the data processing result and improve the fusion performance of multi-source data.
[0144] Figure 7 Schematically shows a block diagram of a recommendation information determination device according to an embodiment of the present disclosure.
[0145] For example Figure 7 As shown, the recommendation information determination device 700 includes a first determination module 701, a first processing module 702, a second processing module 703, and a second determination module 704.
[0146] The first determination module 701 is configured to determine a state sequence according to M to-be-processed data sequences associated with a target user. Each to-be-processed data sequence among the M to-be-processed data sequences includes a channel identifier and N to-be-processed data arranged in chronological order. The channel identifier is used to characterize the source of the N to-be-processed data. The state sequence includes N target states arranged in chronological order. M is a positive integer greater than 1, and N is a positive integer.
[0147] The first processing module 702 is configured to process M to-be-processed data sequences respectively according to the channel identifier and N target states, so as to obtain P target data sequences, where P is a positive integer and P is less than or equal to M.
[0148] The second processing module 703 is configured to perform data fusion processing on the P target data sequences to obtain a fused data processing result.
[0149] The second determination module 704 is configured to determine recommendation information for the target user according to the fused data processing result.
[0150] According to an embodiment of the present disclosure, the first determination module 701 may include a first input sub-module.
[0151] The first input sub-module is configured to input the M to-be-processed data sequences into a Markov decision model and output N target states.
[0152] According to an embodiment of the present disclosure, the recommendation information determination device 700 may further include a third determination module, a dimensionality reduction processing module, a fourth determination module, and a construction module.
[0153] The third determination module is configured to determine a sample action set and a sample state set according to the channel identifier and the sample data sequence corresponding to the channel identifier. The sample action set includes a retention operation and a deletion operation, and the sample state set includes a plurality of sample states, and each sample state in the plurality of sample states is respectively associated with the retention operation and the deletion operation.
[0154] The dimensionality reduction processing module is configured to perform dimensionality reduction processing on the sample state set and the sample action set respectively to obtain a dimensionality-reduced sample state set and a dimensionality-reduced sample action set.
[0155] The fourth determination module is configured to determine a reward function based on the dimensionality-reduced sample state set and the dimensionality-reduced sample action set.
[0156] The construction module is configured to construct a Markov decision model based on the dimensionality-reduced sample state set, the dimensionality-reduced sample action set, and the reward function.
[0157] According to an embodiment of the present disclosure, the first processing module 702 may include a second input sub-module.
[0158] The second input sub-module is configured to input the M to-be-processed data sequences and the state sequence into a trained backpropagation neural network model and output P target data sequences.
[0159] According to an embodiment of the present disclosure, the second input sub-module may include a first determination unit and a first processing unit.
[0160] The first determination unit is configured to determine an action sequence according to the channel identifier and the status sequence of each of the M data sequences to be processed. The action sequence includes N target actions arranged in chronological order, and the N target actions correspond to the channel identifier.
[0161] The first processing unit is configured to process the data sequence to be processed corresponding to the channel identifier according to the N target actions, and obtain P target data sequences.
[0162] According to an embodiment of the present disclosure, the N target actions include a retention operation or a deletion operation.
[0163] According to the N target actions, the first processing unit may include a first determination subunit, a second determination subunit, a retention subunit, a third determination subunit, a fourth determination subunit, and a deletion subunit.
[0164] The first determination subunit is configured to determine the channel identifier to be retained and the moment to be retained corresponding to the retention operation when the N target actions include the retention operation.
[0165] The second determination subunit is configured to determine the data sequence to be retained in the M data sequences to be processed according to the channel identifier to be retained.
[0166] The retention subunit is configured to retain N data to be processed in the data sequence to be retained according to the moment to be retained.
[0167] The third determination subunit is configured to determine the channel identifier to be deleted and the moment to be deleted corresponding to the deletion operation when the N target actions include the deletion operation.
[0168] The fourth determination subunit is configured to determine the data sequence to be deleted in the M data sequences to be processed according to the channel identifier to be deleted.
[0169] The deletion subunit is configured to delete N data to be processed in the data sequence to be deleted according to the moment to be deleted.
[0170] According to an embodiment of the present disclosure, the training device of the trained backpropagation neural network may include a fifth determination module, a first output module, a second input module, an adjustment module, and a sixth determination module.
[0171] The fifth determination module is configured to determine a sample state, a sample action, and a sample reward function according to the channel identifier and the sample data sequence corresponding to the channel identifier.
[0172] The first output module is configured to input the sample data sequence, the sample state, and the sample reward function corresponding to the channel identifier into the backpropagation neural network model to be trained, and output a predicted action.
[0173] A second input module, configured to input a predicted action and a sample action corresponding to a channel identifier into a loss function, and output a loss result.
[0174] An adjustment module, configured to adjust network parameters of a backpropagation neural network model to be trained according to the loss result until the loss function or the number of iterations meets a preset condition.
[0175] A sixth determination module, configured to use the model obtained when the loss function or the number of iterations meets the preset condition as the trained backpropagation neural network model.
[0176] According to an embodiment of the present disclosure, the second processing module 703 may include a processing sub-module.
[0177] The processing sub-module is configured to perform data fusion processing on P target data sequences based on a strongly connected component algorithm to obtain a fused data processing result.
[0178] According to an embodiment of the present disclosure, the processing sub-module may include a second determination unit, an addition unit, a calculation unit, and a fusion processing unit.
[0179] The second determination unit is configured to determine a data sequence set, a data sequence related combination, and a superset of the data sequence related combination according to the P target data sequences.
[0180] The addition unit is configured to add the superset to a candidate data sequence related combination when any data sequence combination in the superset belongs to the data sequence related combination.
[0181] The calculation unit is configured to calculate the relative information entropy of any data sequence combination in the candidate data sequence related combination.
[0182] The fusion processing unit is configured to perform data fusion processing on the P target data sequences according to the relative information entropy to obtain a fused data processing result.
[0183] According to an embodiment of the present disclosure, the second determination module 704 may include a first determination sub-module, a second determination sub-module, and a third determination sub-module.
[0184] The first determination sub-module is configured to determine a risk preference value of a target user according to the fused data processing result.
[0185] The second determination sub-module is configured to determine a risk type of the target user according to the risk preference value.
[0186] The third determination sub-module is configured to determine recommendation information for the target user according to the risk type.
[0187] Any of a plurality of modules, sub-modules, units, and sub-units according to embodiments of the present disclosure, or at least part of the functions of any of them, may be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable manner of integrating or packaging circuits, in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0188] For example, any of the first determination module 701, the first processing module 702, the second processing module 703, and the second determination module 704 may be combined and implemented in one module / unit / sub-unit, or any one of the module / unit / sub-unit may be split into multiple module / unit / sub-units. Alternatively, at least part of the functions of one or more of these module / unit / sub-units may be combined with at least part of the functions of other module / unit / sub-units and implemented in one module / unit / sub-unit. According to embodiments of the present disclosure, at least one of the first determination module 701, the first processing module 702, the second processing module 703, and the second determination module 704 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable manner of integrating or packaging circuits, in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the first determination module 701, the first processing module 702, the second processing module 703, and the second determination module 704 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0189] It should be noted that the part of the recommendation information determination device in the embodiments of the present disclosure corresponds to the part of the recommendation information determination method in the embodiments of the present disclosure. For the description of the part of the recommendation information determination device, please specifically refer to the part of the recommendation information determination method, and details are not described herein again.
[0190] Figure 8 FIG. 0 schematically shows a block diagram of an electronic device suitable for implementing a method for determining recommendation information according to an embodiment of the present disclosure. Figure 8 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0191] As Figure 8 shown, the computer electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), and so on. The processor 801 may also include on-board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0192] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to an embodiment of the present disclosure by executing the program in the ROM 802 and / or the RAM 803. It should be noted that the program may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the program stored in the one or more memories.
[0193] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0194] According to an embodiment of the present disclosure, the method flow according to the embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0195] The present disclosure also provides a computer-readable storage medium, which can be included in the device / device / system described in the above embodiment; or can exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0196] According to an embodiment of the present disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. For example, it can include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device.
[0197] For example, according to an embodiment of the present disclosure, the computer-readable storage medium can include the above-described ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803.
[0198] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program includes program code for executing the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the recommended information determination method provided by the embodiment of the present disclosure.
[0199] When the computer program is executed by the processor 801, the above functions defined in the system / device of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, modules, units, etc. can be implemented by computer program modules.
[0200] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication section 809, and / or installed from the removable medium 811. The program code included in the computer program may be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0201] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0202] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0203] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A method for determining recommendation information, comprising: Determining a state sequence according to M to-be-processed data sequences associated with a target user, where each of the M to-be-processed data sequences includes a channel identifier and N to-be-processed data arranged in chronological order, the channel identifier is used to represent the source of the N to-be-processed data, the state sequence includes N target states arranged in chronological order, M is a positive integer greater than 1, and N is a positive integer; Processing each of the M to-be-processed data sequences according to the channel identifier and the N target states to obtain P target data sequences, where processing each of the M to-be-processed data sequences according to the channel identifier and the N target states to obtain P target data sequences includes: determining an action sequence according to the channel identifiers and the state sequence of the M to-be-processed data sequences respectively, so as to resolve conflicting information existing in the M to-be-processed data sequences from different sources through different action selections, the action sequence includes N target actions arranged in chronological order, and the N target actions correspond to the channel identifier; processing the to-be-processed data sequence corresponding to the channel identifier according to the N target actions to obtain the P target data sequences, P is a positive integer, and P is less than or equal to M; Performing data fusion processing on the P target data sequences to obtain a fused data processing result; and Determining recommendation information for the target user according to the fused data processing result.
2. The method according to claim 1, wherein The determining a state sequence according to M to-be-processed data sequences associated with a target user includes: Inputting the M to-be-processed data sequences into a Markov decision model and outputting the N target states.
3. The method according to claim 2, further comprising, before determining the state sequence according to M to-be-processed data sequences associated with a target user: Determine a set of sample actions and a set of sample states according to the channel identifier and the sample data sequence corresponding to the channel identifier, where The sample action set includes a retention operation and a deletion operation, the sample state set includes multiple sample states, and each sample state in the multiple sample states is respectively associated with the retention operation and the deletion operation; Performing dimensionality reduction processing on the sample state set and the sample action set respectively to obtain a dimensionality-reduced sample state set and a dimensionality-reduced sample action set; Determining a reward function based on the dimensionality-reduced sample state set and the dimensionality-reduced sample action set; and And Constructing the Markov decision model based on the dimensionality-reduced sample state set, the dimensionality-reduced sample action set, and the reward function.
4. The method according to claim 1, wherein, The processing each of the M to-be-processed data sequences according to the channel identifier and the N target states to obtain P target data sequences includes: Inputting the M to-be-processed data sequences and the state sequence into a trained backpropagation neural network model and outputting the P target data sequences.
5. The method according to claim 1, wherein The N target actions include a retention operation or a deletion operation; Processing the to-be-processed data sequence corresponding to the channel identifier according to the N target actions to obtain the P target data sequences includes: When the N target actions include a retention operation, determining a to-be-retained channel identifier and a to-be-retained time corresponding to the retention operation; Determining a to-be-retained data sequence from the M to-be-processed data sequences according to the to-be-retained channel identifier; Retaining the N to-be-processed data in the to-be-retained data sequence according to the to-be-retained time; When the N target actions include a deletion operation, determining a to-be-deleted channel identifier and a to-be-deleted time corresponding to the deletion operation; Determining a to-be-deleted data sequence from the M to-be-processed data sequences according to the to-be-deleted channel identifier; and Deleting the N to-be-processed data in the to-be-deleted data sequence according to the to-be-deleted time.
6. The method according to claim 4, wherein, The training method of the trained backpropagation neural network includes: Determining a sample state, a sample action, and a sample reward function according to the channel identifier and the sample data sequence corresponding to the channel identifier; Inputting the sample data sequence, the sample state, and the sample reward function corresponding to the channel identifier into a to-be-trained backpropagation neural network model to output a predicted action; Inputting the predicted action and the sample action corresponding to the channel identifier into a loss function to output a loss result; Adjusting the network parameters of the to-be-trained backpropagation neural network model according to the loss result until the loss function or the number of iterations meets a preset condition; and Taking the model obtained when the loss function or the number of iterations meets the preset condition as the trained backpropagation neural network model.
7. The method according to claim 1, wherein The data fusion processing of the P target data sequences to obtain a fused data processing result includes: Based on the strongly connected component algorithm, performing data fusion processing on the P target data sequences to obtain the fused data processing result.
8. The method according to claim 7, wherein The performing data fusion processing on the P target data sequences based on the strongly connected component algorithm to obtain the fused data processing result includes: Determining a data sequence set, a data sequence related combination, and a superset of the data sequence related combination according to the P target data sequences; When any data sequence combination in the superset belongs to the data sequence related combination, adding the superset to the candidate data sequence related combination; Calculating the relevant information entropy of any data sequence combination in the candidate data sequence related combination; and Performing data fusion processing on the P target data sequences according to the relevant information entropy to obtain the fused data processing result.
9. The method according to any one of claims 1 to 8, wherein Determining recommended information for the target user according to the fused data processing result includes: Determining a risk preference value of the target user according to the fused data processing result; Determining a risk type of the target user according to the risk preference value; and Determining the recommended information for the target user according to the risk type.
10. A recommended information determination device includes: A first determination module, configured to determine a state sequence according to M to-be-processed data sequences associated with a target user, where each of the M to-be-processed data sequences includes a channel identifier and N to-be-processed data arranged in chronological order, the channel identifier is used to characterize the source of the N to-be-processed data, the state sequence includes N target states arranged in chronological order, M is a positive integer greater than 1, and N is a positive integer; A first processing module, configured to process each of the M to-be-processed data sequences according to the channel identifier and the N target states to obtain P target data sequences, where processing each of the M to-be-processed data sequences according to the channel identifier and the N target states to obtain P target data sequences includes: determining an action sequence according to the channel identifier of each of the M to-be-processed data sequences and the state sequence, so as to resolve conflict information existing in the M to-be-processed data sequences from different sources through different action selections, the action sequence includes N target actions arranged in chronological order, and the N target actions correspond to the channel identifier; processing the to-be-processed data sequence corresponding to the channel identifier according to the N target actions to obtain the P target data sequences, P is a positive integer, and P is less than or equal to M; A second processing module, configured to perform data fusion processing on the P target data sequences to obtain a fused data processing result; and A second determination module, configured to determine recommendation information for the target user according to the fused data processing result.
11. An electronic device, comprising: One or more processors; A memory, configured to store one or more instructions, wherein when the one or more instructions are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, on which executable instructions are stored, and when the executable instructions are executed by a processor, the processor implements the method according to any one of claims 1 to 9.
13. A computer program product, the computer program product includes computer-executable instructions, and the computer-executable instructions are used to implement the method according to any one of claims 1 to 9 when executed.
Citation Information
Patent Citations
Recommendation model training method and system
CN111311384A
Information recommendation method and device, electronic equipment and storage medium
CN114329173A