Information processing system, information processing method, and computer program

The information processing system generates feature vectors from user online actions to evaluate and compare behavior across platforms, addressing the privacy concerns related to third-party cookies and enabling effective user behavior analysis and targeted advertising.

JP7832406B1Active Publication Date: 2026-03-17HAKUHODO TECHNOLOGIES INC +1
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

The increasing awareness of privacy has led to the abolition of third-party cookies, necessitating a new method to evaluate and compare user online behavior across multiple online spaces without relying on them.

Method used

An information processing system that generates feature vectors from descriptive data of user online actions, performs mathematical processing to create user feature data, and allows for the comparison and evaluation of online behavior across platforms, using a vector generation and processing unit.

Benefits of technology

Enables accurate evaluation and comparison of user online behavior across multiple platforms, facilitating targeted advertising and user group identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007832406000001_ABST
    Figure 0007832406000001_ABST
Patent Text Reader

Abstract

This provides novel technology for evaluating users' online behavior. [Solution] An information processing method relating to one aspect of this disclosure includes, for each of a plurality of users, acquiring descriptive data that describes one or more online actions of a corresponding user. The information processing method includes, for each user, generating feature vectors based on one or more linguistic pieces of information that describe one or more online actions of the corresponding user included in the descriptive data, thereby generating one or more feature vectors corresponding to one or more online actions. The information processing method also includes, for each user, generating user feature data that statistically describes the online actions of the corresponding user by performing mathematical processing on the one or more feature vectors generated by the vector generation unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing system and an information processing method.

Background Art

[0002] Conventionally, a technique for tracking a user's online behavior using third - party cookies has been known. The tracked user's online behavior is used, for example, to estimate characteristics such as the user's interests and preferences. The estimated user characteristics are used, for example, for advertising delivery to the user (see, for example, Patent Document 1). In addition, an example is known in which an input sentence of a user is vectorized to create teacher data, and a prediction model regarding the user is constructed based on the teacher data (see, for example, Patent Document 2).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, due to the increasing awareness of privacy, the use of third - party cookies has tended to be abolished. Therefore, it is preferable to be able to compare and evaluate the online behavior of a user in multiple online spaces without using third - party cookies.

[0005] Therefore, according to one aspect of the present disclosure, it is desirable to be able to provide a new technique for evaluating a user's online behavior.

Means for Solving the Problems

[0006] According to one aspect of this disclosure, an information processing system is provided. The information processing system comprises an acquisition unit, a vector generation unit, and a vector processing unit. The acquisition unit is configured to acquire descriptive data for each of multiple users, describing one or more online actions of the corresponding user.

[0007] The vector generation unit is configured to generate one or more feature vectors corresponding to one or more online actions for each user by generating feature vectors based on one or more linguistic pieces of information that describe one or more online actions of the corresponding user included in the explanatory data.

[0008] The vector processing unit is configured to generate user feature data that statistically explains the online behavior of a corresponding user by performing mathematical processing on one or more feature vectors generated by the vector generation unit for each user.

[0009] According to the information processing system configured in this way, a user's online behavior can be appropriately evaluated in the feature space by mathematical processing of the one or more feature vectors that use one or more linguistic pieces of information that describe one or more online behaviors, and user feature data that statistically represents the characteristics of the user's online behavior well can be generated.

[0010] Using such user characteristic data, the online behavior of multiple users across multiple online spaces can be appropriately compared and evaluated. Therefore, according to one aspect of this disclosure, a meaningful technology for evaluating user online behavior can be provided.

[0011] According to one aspect of this disclosure, the descriptive data may include, for each online behavior, a behavioral description text which is a string of characters that describes the content of the corresponding online behavior, and type data which describes the type of the corresponding online behavior.

[0012] The vector generation unit can be configured to generate one or more feature vectors corresponding to one or more online actions by generating a feature vector based on the action description text and type data for each online action.

[0013] By generating user feature data in this way, it is possible to appropriately evaluate users' online behavior in the feature space, taking into account the type of behavior, and to generate user feature data that statistically represents the characteristics of users' online behavior well.

[0014] According to one aspect of this disclosure, a feature vector based on behavioral description text and type data may be a feature vector that is a higher-dimensional representation of a text vector generated by vectorizing the behavioral description text, based on type data. In this case, the higher-dimensional representation can be achieved by adding a vector element representing the type of the corresponding online behavior to the text vector.

[0015] According to one aspect of this disclosure, the descriptive data may include, for each online action, descriptive text which is a string of characters that describes the content of the corresponding online action and the type of the corresponding online action for each online action.

[0016] In this case, the vector generation unit can be configured to generate one or more feature vectors corresponding to one or more online actions by vectorizing the descriptive text for each online action.

[0017] According to one aspect of this disclosure, the behavioral description text or description text may describe the content of the corresponding online behavior as a string of characters the input from the corresponding user in the corresponding online behavior.

[0018] By generating user feature data based on feature vectors that take user input into account, it is possible to generate user feature data that accurately represents user interests and other characteristics.

[0019] According to one aspect of the present disclosure, the vector processing unit may be configured to generate, as user feature data, one user feature vector representing one or more feature vectors by performing mathematical processing on the one or more feature vectors.

[0020] According to one aspect of the present disclosure, the information processing system may further include an extraction unit configured to extract, from a plurality of users, a user having online behavior characteristics similar to the characteristics of the online behavior described by the reference vector based on the comparison between the user feature vector for each user and the reference vector. The information processing system configured in this way is useful, for example, for extracting distribution targets from user groups in a plurality of online spaces, although not limited thereto.

[0021] According to one aspect of the present disclosure, the information processing system may further include a model construction unit configured to construct an estimation model for estimating a predetermined feature of a user by using the user feature vector for each user as teacher data. The estimation model may be an estimation model having at least a part of the vector elements of the user feature vector as explanatory variables, and may be an estimation model for estimating a predetermined feature of a user explained by the explanatory variables.

[0022] The information processing system configured in this way is useful, for example, for estimating various features of a user based on the user's online behavior, although not limited thereto.

[0023] According to one aspect of the present disclosure, the vector processing unit may be configured to generate, as user feature data, user feature data that probabilistically explains the online behavior of the corresponding user, or user feature data that explains the predicted online behavior of the corresponding user, by performing mathematical processing on one or more feature vectors.

[0024] The information processing system configured in this way can effectively utilize the user feature data, for example, for pre-emptive advertisement distribution based on the user's online behavior, although not limited thereto.

[0025] According to one aspect of the present disclosure, one or more online actions may be a plurality of online actions, and one or more feature vectors may be a plurality of feature vectors corresponding to the plurality of online actions. The plurality of feature vectors may include feature vectors for each online action. The explanatory data may explain a plurality of online actions of the corresponding user in a time series. The explanatory data may include, for each online action, language information explaining the corresponding online action.

[0026] In this case, the vector generation unit may generate time-series data of a plurality of feature vectors corresponding to the plurality of online actions by generating a feature vector based on the language information for each online action.

[0027] The vector processing unit may be configured to generate user feature data that probabilistically explains the online actions of the corresponding user or user feature data that explains the predicted online actions of the corresponding user as user feature data by inputting the time-series data into a time-series prediction model.

[0028] <00001-01>According to one aspect of the present disclosure, the explanatory data may include, for each online action, action explanation text that is a character string explaining the content of the corresponding online action, and type data that explains the type of the corresponding online action. In this case, the vector generation unit may generate, as a feature vector for each online action, a first feature vector obtained by vectorizing the action explanation text and a second feature vector obtained by vectorizing the type data for each online action.

[0029] The vector processing unit may be configured to generate user feature data that probabilistically explains the online actions of the corresponding user or user feature data that explains the predicted online actions of the corresponding user as user feature data by inputting time-series data in which the first feature vector and the second feature vector are alternately arranged into a time-series prediction model.

[0030] According to one aspect of this disclosure, each of one or more pieces of linguistic information may include keywords that describe a corresponding online behavior among one or more online behaviors. Multiple users may include a first user group, a second user group, and so on.

[0031] The information processing system may further include a clustering unit configured to cluster specific types of groups within a first user group into multiple clusters based on the frequency of occurrence of multiple keywords in the online behavior of each user belonging to that specific type of group.

[0032] The information processing system may further include a centroid calculation unit configured to calculate the centroid of the user feature vectors of one or more users belonging to the corresponding cluster, as the cluster centroid for each cluster.

[0033] The information processing system may further include a correction unit configured to correct the cluster centroid for each cluster based on the user feature vectors of each user in the first user group and the user feature vectors of each user in the second user group.

[0034] The information processing system may further include a discrimination unit configured to distinguish between specific types of groups within a second user group based on a corrected cluster centroid.

[0035] According to the information processing system configured in this way, it is possible to appropriately identify similar groups existing in multiple different online spaces (e.g., digital platforms) based on user characteristic data, although this is not limited to such systems.

[0036] According to one aspect of this disclosure, the correction unit may be configured to determine the keyword set that appears in the corresponding cluster for each cluster. The correction unit may be configured to calculate the centroid of the user feature vector for the first user set as the first centroid for each cluster. The first user set is the set of users from the first user group that use the keyword set that appears in the corresponding cluster.

[0037] The correction unit may be configured to calculate the centroid of the user feature vector of the second user set as the second double center for each cluster. The second user set is the set of users from the second user group that use the keyword set that appears in the corresponding cluster.

[0038] The correction unit may be configured to correct the cluster centroid of a corresponding cluster based on the difference between the first centroid and the second centroid for each cluster. With such a configuration, it is possible to appropriately correct the cluster centroid.

[0039] According to one aspect of this disclosure, a computer program may be provided to enable a computer to at least partially implement the functions of the information processing system described above as an acquisition unit, a vector generation unit, and a vector processing unit. The computer program may be recorded on a computer-readable, non-temporary recording medium.

[0040] According to one aspect of this disclosure, an information processing method corresponding to the information processing system described above may be provided. The information processing method may be executed by a computer.

[0041] According to one aspect of this disclosure, the information processing method may include, for each of multiple users, obtaining descriptive data that describes one or more online actions of the corresponding user.

[0042] The information processing method may further include generating feature vectors corresponding to one or more online actions for each user by generating feature vectors based on one or more linguistic pieces of information that describe one or more online actions of the corresponding user included in the explanatory data.

[0043] The information processing method may further include generating user feature data that statistically explains the online behavior of a corresponding user by performing mathematical processing on one or more feature vectors generated by a vector generation unit for each user. An information processing method configured in this way will have the same effect as the information processing system described above. [Brief explanation of the drawing]

[0044] [Figure 1] This is a block diagram representing the configuration of an information processing system. [Figure 2] This is a flowchart illustrating the analysis-related processing performed by the processor in the first embodiment. [Figure 3] Figure 3A shows the structure of online behavioral data, and Figure 3B shows the structure of behavioral feature vectors. [Figure 4] This is a flowchart representing the vector generation process performed by the processor in the first embodiment. [Figure 5] This is a flowchart illustrating the extraction-related processes performed by the processor. [Figure 6] This is a flowchart representing the estimation-related processes performed by the processor. [Figure 7] Figure 7A is a diagram showing the structure of online behavioral data in the second embodiment, and Figure 7B is a diagram showing the structure of behavioral feature vectors in the second embodiment. [Figure 8] This is a flowchart representing the vector generation process performed by the processor in the second embodiment. [Figure 9] This is a flowchart illustrating the prediction-related processing performed by the processor in the third embodiment. [Figure 10] This is a block diagram showing the configuration of the time series prediction model in the third embodiment. [Figure 11] This is a block diagram showing the configuration of the time series prediction model in the fourth embodiment. [Figure 12] This is a block diagram showing the configuration of the time series forecasting model in a modified example. [Figure 13] This is a block diagram showing the configuration of the time series prediction model in the fifth embodiment. [Figure 14] This is a flowchart illustrating the correction search process performed by the processor in the sixth embodiment. [Figure 15] This is a flowchart showing the correction vector calculation process performed by the processor in the sixth embodiment. [Modes for carrying out the invention]

[0045] Exemplary embodiments of the present disclosure are described below with reference to the drawings. [First Embodiment] The information processing system 1 of this embodiment, shown in Figure 1, is a system for analyzing the online behavior of multiple users collected across multiple online spaces, specifically, multiple digital platform DPs. This embodiment does not presuppose the use of third-party cookies, nor does it presuppose an environment where the online behavior of users using multiple digital platform DPs can be tracked across those platforms. Regardless of whether or not third-party cookies are used, the information processing system 1 of this embodiment can appropriately analyze the online behavior of multiple users across multiple digital platform DPs.

[0046] In this embodiment, the information processing system 1 generates a user feature vector that accurately represents the online behavior characteristics of a corresponding user, based on online behavior data obtained from the corresponding digital platform DP for each digital platform DP. This user feature vector is used as a common indicator of user characteristics across multiple digital platform DPs. The information processing system 1 databases the user feature vectors of each user utilizing the digital platform DP for use in advertising distribution and behavioral analysis.

[0047] The information processing system 1 comprises a processor 11, memory 13, storage 15, user interface 17, and communication interface 19. The processor 11 executes processing according to the computer program stored in the storage 15.

[0048] Memory 13 is the main memory and is used as working memory when the processor 11 executes processing. Storage 15 is an auxiliary storage device such as a hard disk drive and a solid-state drive, and stores computer programs as well as various data used for processing executed by the processor 11.

[0049] The user interface 17 comprises a display unit 17A and an operation unit 17B. The display unit 17A is controlled by the processor 11 and displays information for the operator operating the information processing system 1. Examples of the display unit 17A include liquid crystal displays and organic EL displays. The operation unit 17B is configured to input operation signals from the operator to the processor 11. The operation unit 17B consists of one or more operating devices, such as a keyboard and a mouse.

[0050] The communication interface 19 is configured to communicate with external devices connected to a wide-area network. The information processing system 1 is configured to acquire online behavioral data of each of the aforementioned digital platforms DP through the communication interface 19.

[0051] According to this embodiment, the processor 11 executes the analysis-related processing shown in Figure 2 based on instructions from the operator input through the user interface 17. When the analysis-related processing is started, the processor 11 retrieves a set of online behavioral data provided from the digital platform DP specified by the operator from the storage 15 (S110). The retrieved set of online behavioral data includes user-specific online behavioral data for multiple users in the corresponding digital platform DP.

[0052] User-specific online behavior data provided from each digital platform DP is stored in storage 15. The processor 11 can acquire user-specific online behavior data from each digital platform DP and store it in storage 15 by executing periodic processing (not shown).

[0053] In another example, the processor 11 may communicate with a designated digital platform DP based on acquisition instructions from the operator input through the user interface 17 to acquire user-specific online behavior data on the corresponding digital platform DP.

[0054] In another example, the information processing system 1 may include a connection interface such as a USB (Universal Serial Bus) interface (not shown). Based on instructions from the operator input through the user interface 17, the processor 11 can acquire user-specific online behavior data stored in a USB device connected to the connection interface.

[0055] Among multiple digital platform DPs, the same user may exist between any first digital platform DP and any second digital platform DP that is different from the first digital platform DP. In this case, the online behavior data of the same user will be provided from both the first and second digital platform DPs as online behavior data of an individual user whose identity is unknown.

[0056] When processor 11 acquires a set of online behavior data in S110, in the following S120, it selects a target user from the group of users corresponding to the set of online behavior data, which is the user to be set as the target of processing from S130 onwards.

[0057] In the subsequent S130, the processor 11 references the target user's online behavior data and generates a feature vector for each online behavior. The feature vector generated here will be referred to as the behavior feature vector below.

[0058] As shown in Figure 3A, online behavior data contains a record for each online behavior performed by a corresponding user, for one or more online behaviors. Each record describes the content and type of the corresponding online behavior using linguistic information. In Figure 3A, each row indicated by a row number can be understood as a single record.

[0059] Specifically, each record has an action description text and type data. The action description text is a string that describes the content of the corresponding online action. The action description text describes the input from or output to the corresponding user in the online action. The type data is a string that describes the type of online action. The action feature vector for each online action is generated based on the action description text and type data of the corresponding online action.

[0060] Examples of online behavior include web search behavior through search sites provided by the digital platform DP, purchasing behavior through e-commerce (EC) sites provided by the digital platform DP, and web page browsing behavior (hereinafter referred to as "web browsing behavior").

[0061] If the online behavior is a web search, the behavior description text is the search string entered by the user during the web search. The type data describes that the online behavior is a web search and / or that the behavior description text is a search string using the string "search".

[0062] If the online behavior is a purchasing behavior, the behavior description text is a string representing the type of product purchased. The type data explains that the online behavior is a purchasing behavior and / or that the behavior description text describes the purchased item using the string "purchase".

[0063] If the online behavior is web browsing, the behavior description text is the output of the viewed webpage, specifically keywords that accurately represent the content of the webpage. The type data describes that the online behavior is web browsing and / or that the behavior description text is keywords contained in the webpage using the string "URL".

[0064] In S130, the processor 11 can execute the vector generation process shown in Figure 4. In the vector generation process, the processor 11 selects one of the records included in the target user's online behavior data as the record to be processed (S211).

[0065] In the subsequent S212, the processor 11 converts the action description text contained in the record to be processed into a feature vector. Word2Vec is a known technique for converting text into feature vectors. The feature vector generated by Word2Vec is a vector representation of text that expresses the similarity and semantic relationships between texts in a vector space and / or feature space. Examples of text here include words, combinations of words, and phrases.

[0066] In S212, the processor 11 converts the behavioral description text into a feature vector of a predetermined dimension using a Word2Vec conversion model. This conversion model, i.e., the text-to-feature vector conversion model, may be a conversion model specifically learned for the information processing system 1 of this embodiment, or it may be a publicly available pre-trained model. Hereinafter, the feature vector generated in S212 will be referred to as the "text vector".

[0067] In the subsequent S213, the processor 11 converts the text vector generated in S212 into a higher-dimensional feature vector by adding a vector element based on type data (see Figure 3B). By increasing the dimensionality of the text vector based on type data, the processor 11 generates an action feature vector corresponding to the record being processed. The feature vector generated by increasing the dimensionality of the text vector in S213 is the action feature vector described above.

[0068] When the text vector is an N1-dimensional vector, the behavioral feature vector generated in S213 is an (N1+N2)-dimensional vector. Here, the value N2 is the total number of types represented by the type data, in other words, the number of possible string types for the type data.

[0069] Each of the N2 additional vector elements (hereinafter referred to as additional vector elements) added to the text vector corresponds to one of the N2 types. Each additional vector element represents, with a binary value of "1" or "0", whether the online action corresponding to the text vector belongs to the corresponding type.

[0070] For example, if the online action corresponding to the text vector is a web search action, or in other words, if the action description text corresponding to the text vector is a search string, then N2 vector elements are added to the text vector, where the value of the additional vector element corresponding to type "search" is "1", and the values ​​of the remaining (N2-1) additional vector elements are "0".

[0071] Thus, the behavioral feature vector generated in S213 represents the type of online behavior corresponding to the text vector, using N2 additional vector elements.

[0072] When the processor 11 generates behavioral feature vectors corresponding to the records to be processed, in the subsequent S214, it determines whether or not it has generated behavioral feature vectors for all records included in the online behavior data of the target user.

[0073] If it determines that all vectors have been generated (Yes in S214), the processor 11 terminates the vector generation process shown in Figure 4. On the other hand, if it determines that not all vectors have been generated (No in S214), the processor 11 selects one record from the online behavior data of the target user that has not yet generated a behavior feature vector as the record to be processed (S211), and executes the processing from S212 onward.

[0074] In this way, the processor 11 generates behavioral feature vectors for each online behavior by generating a behavioral feature vector for each record for all records included in the target user's online behavior data. That is, the processor 11 generates one or more behavioral feature vectors corresponding to one or more online behaviors described in the target user's online behavior data.

[0075] In the subsequent S140, the processor 11 generates a user feature vector that statistically explains the target user's online behavior by performing mathematical processing on one or more behavioral feature vectors of the target user generated in S130. Through this mathematical processing, one or more behavioral feature vectors of the target user are transformed into a single representative user feature vector.

[0076] Specifically, the processor 11 generates a user feature vector that statistically represents one or more behavioral feature vectors of the target user by performing statistical processing on one or more behavioral feature vectors of the target user. This user feature vector is specifically the mean vector of one or more behavioral feature vectors of the target user. The value of each vector element in the user feature vector corresponds to the arithmetic mean of the corresponding vector elements in one or more behavioral feature vectors of the target user.

[0077] Subsequently, processor 11 determines whether or not user feature vectors have been calculated for all users (S150). If the determination is negative (No in S150), processor 11 returns to the process in S120 and sets one of the users who was not selected as a target for processing from S130 onwards from the group of users corresponding to the group of online behavior data acquired in S110 as the new target user, and then executes the processing from S130 onwards.

[0078] Through repeated processing, the processor 11 generates a user feature vector for each user that statistically best represents the online behavior of the corresponding user, based on the corresponding user's online behavior data.

[0079] When the processor 11 generates user feature vectors for all users (Yes in S150), it databases each user feature vector and stores them in the storage 15 (S160). The database of user feature vectors stored in the storage 15 describes the user feature vector for each user of the corresponding digital platform DP, associated with the identification code of the corresponding user.

[0080] After completing the processing in S160, processor 11 terminates the analysis-related processing shown in Figure 2. Following instructions from the user, processor 11 executes the analysis-related processing shown in Figure 2 for each digital platform DP and saves the user feature vector database to storage 15. Hereinafter, this database will be referred to as the user database.

[0081] The operator of Information Processing System 1 can use this user database to extract, for example, users from the second digital platform DP whose online behavioral characteristics are similar to those of users on the first digital platform DP.

[0082] The user database may be associated with other user information obtained from the corresponding digital platform DP. For example, the information provided by the first digital platform DP may include the responses to a survey conducted for each user. In this case, the user database of the first digital platform DP may be configured to store a user feature vector for each user, associated with the corresponding user's survey responses.

[0083] In this case, the processor 11 of the information processing system 1 can extract a set of users (i.e., one or more users) with specific interests or awareness from the group of users of the first digital platform DP registered in the user database, based on the responses to the questionnaire.

[0084] The processor 11 can calculate a representative user feature vector for this user set from the user feature vector of the first digital platform DP. This representative user feature vector will be referred to as the reference vector below.

[0085] Next, the processor 11 can extract a set of users of the second digital platform DP that have a user feature vector similar to the reference vector from the user group of the second digital platform DP by performing the extraction-related processing shown in Figure 5.

[0086] In the extraction-related processing, the processor 11 sets a representative user feature vector for the user to be extracted (S310). That is, in S310, a reference vector is set. In the subsequent S320, the processor 11 refers to the user database of the second digital platform DP, i.e., the digital platform DP from which users are to be extracted, and compares the reference vector with the user feature vector of each user.

[0087] As a result, the processor 11 extracts a group of users from the user group of the digital platform DP whose user feature vector has a similarity (e.g., cosine similarity) to the reference vector that is equal to or greater than a predetermined value. The extracted user group is a group of users whose online behavior characteristics are similar to those described by the reference vector. If the reference vector is based on the user feature vector of a set of users who gave a specific answer in a survey in the first digital platform DP, then the extracted user group can also be said to be a set of users who are estimated to be highly likely to give a specific answer if a survey were conducted.

[0088] In the subsequent S330, the processor 11 outputs a list of users extracted in S320. The user list may be data describing the identification code of each extracted user. The user list may be stored in the storage 15 as a data file, for example. This list may also be displayed to the operator through the display unit 17A. The processor 11 then terminates the extraction-related processing shown in Figure 5.

[0089] According to this extraction process, the operator can extract user sets from the second digital platform DP that have a high degree of identity or similarity to the user set extracted from the user group of the first digital platform DP through the information processing system 1, and appropriately select targets for advertising distribution and various measures from the user group of the second digital platform DP.

[0090] In the example above, the reference vector was generated from the user feature vector of the user set of interest. However, the reference vector may be generated in other ways. For example, in the context of product appeals to users, the reference vector may be generated by vectorizing the keywords of the product being promoted. The reference vector may be generated in a similar manner to the behavior feature vector, based on the keywords of the product being promoted and the types of online behaviors suitable for promoting the product.

[0091] Alternatively, the user group of the first digital platform DP may be clustered, and an estimation model may be constructed to estimate the clusters corresponding to the user feature vectors. The information processing system 1 may be configured to apply this estimation model to the user group of the second digital platform DP to estimate the cluster to which each user of the second digital platform DP belongs. The estimation model may be, for example, a classification model or a regression model.

[0092] For example, the information processing system 1 can cluster the user groups of the first digital platform DP based on each user's web browsing history obtained from the first digital platform DP. For example, the information processing system 1 can cluster the user groups into multiple clusters in terms of their awareness of a specific product (e.g., men's cosmetics) based on their browsing history of web pages related to that product.

[0093] Subsequently, the information processing system 1 identifies the user feature vector for each user, and constructs an estimation model using machine learning, with some or all of the multiple vector elements included in the user feature vector as explanatory variables and the label value, which is the identification value of the affiliated cluster, as the dependent variable, using combinations of user feature vectors and affiliated clusters for the user group of the first digital platform DP as training data.

[0094] Information processing system 1 can apply this estimation model to the user group of the second digital platform DP to estimate the cluster to which each user belongs in the second digital platform DP.

[0095] To this end, the processor 11 can perform the estimation-related processing shown in Figure 6 in accordance with instructions from the operator. When the processor 11 starts the estimation-related processing, it acquires classification data about the user group of the first digital platform DP (S410).

[0096] The classification data indicates the cluster to which each user belongs in the first digital platform DP using label values. As in the example above, the classification data is created based on each user's web browsing history obtained from the first digital platform DP.

[0097] The processor 11 can perform a clustering process (not shown) to generate classification data, thereby clustering the user group of the first digital platform DP into multiple clusters.

[0098] Subsequently, the processor 11 refers to the user database of the first digital platform DP and generates training data for each user in the user group of the first digital platform DP, showing the relationship between the corresponding user's label value identified from the classification data and the corresponding user's feature vector identified from the user database (S420).

[0099] In the subsequent S430, the processor 11 sets the vector element specified by the operator from among the multiple vector elements of the user feature vector as an explanatory variable, and sets the label value as the target variable.

[0100] In the subsequent S440, the processor 11 constructs an estimation model for estimating the value of the target variable from the values ​​of the set explanatory variables, using machine learning based on the user-specific training data generated in S420.

[0101] The estimation model may be a classification model having defined explanatory variables and a target variable. This estimation model is one in which at least some of the vector elements of the user feature vector are used as explanatory variables. The estimation model is configured to estimate the user's belonging cluster from the user's characteristics explained by the explanatory variables. For example, the estimation model is configured to estimate the characteristics (belonging cluster) related to awareness of a specific product category in a corresponding user.

[0102] In the subsequent S450, the processor 11 applies an estimation model to each user of the second digital platform DP and calculates each user's label value from the user feature vector of each user of the second digital platform DP.

[0103] In the subsequent S460, the processor 11 expands the user database of the second digital platform DP by adding the label values ​​for each user calculated in S450 to the user database. After that, the processor 11 terminates the estimation-related processing.

[0104] According to this estimation-related processing, the information processing system 1 can estimate the other characteristics of each user in the second digital platform DP based on the characteristics of each user's online behavior in the first digital platform DP, represented by a user feature vector, and other characteristics possessed by each user in the first digital platform DP. Therefore, according to this information processing system 1, meaningful analysis regarding multiple users utilizing multiple digital platform DPs can be performed across multiple digital platform DPs.

[0105] Furthermore, if the user database is associated with other user information besides the user feature vectors, some of this user information may be set as additional explanatory variables to construct an estimation model. By constructing such an estimation model, user characteristics can be estimated with high accuracy.

[0106] [Second Embodiment] Next, the information processing system 1 of the second embodiment will be described. However, the information processing system 1 of the second embodiment differs from the first embodiment only in the configuration of online behavioral data and the method of generating behavioral feature vectors.

[0107] Therefore, in the following description of the second embodiment, we will selectively describe the configuration of the information processing system 1 of the second embodiment that differs from that of the first embodiment. Specifically, we will describe the configuration of online behavioral data (see Figure 7A) and the details of the vector generation process shown in Figure 8, which the processor 11 executes in place of the vector generation process shown in Figure 4.

[0108] The configuration of the information processing system 1 of the second embodiment, which is not mentioned below, may be understood to be the same as that of the first embodiment. The same reference numerals are used for the components of the information processing system 1 of the second embodiment that are the same as those of the information processing system 1 of the first embodiment, and their descriptions are omitted.

[0109] In this embodiment, the user-specific online behavior data obtained from each digital platform DP (see Figure 7A) has a record for each online behavior for one or more online behaviors. Each record includes a series of strings of descriptive text that describes the content of the corresponding online behavior and the type of the corresponding online behavior.

[0110] In response to the configuration of such online behavioral data, processor 11 executes the process shown in Figure 8 as a vector generation process in S130. When the vector generation process shown in Figure 8 is started, the processor 11 selects one of the records included in the target user's online behavior data as the record to be processed (S221).

[0111] In the subsequent S222, the processor 11 generates an action feature vector for the record to be processed by converting the descriptive text contained in the record to be processed into a feature vector. Similar to the processing in S212 in the first embodiment, the processor 11 uses a Word2Vec conversion model to convert the descriptive text, which represents the content and type of online action as a string, into a feature vector of a predetermined dimension.

[0112] As can be understood from the above description, unlike the first embodiment, the processor 11 treats the entire string representing the content and type of online behavior as a series of strings relating to the corresponding online behavior. The processor 11 generates a behavior feature vector (see Figure 7B) by converting this series of strings into a feature vector.

[0113] When the processor 11 generates behavioral feature vectors corresponding to the records to be processed, in the subsequent S223, it determines whether or not it has generated behavioral feature vectors for all records included in the online behavior data of the target user.

[0114] If it determines that all vectors have been generated (Yes in S223), the processor 11 terminates the vector generation process shown in Figure 8. On the other hand, if it determines that not all vectors have been generated (No in S223), in S221, the processor 11 selects one record from the online behavior data of the target user that has not yet generated a behavior feature vector as the record to be processed, and executes the processing from S222 onwards.

[0115] In this way, the processor 11 generates a behavioral feature vector for each record for all records included in the target user's online behavioral data. As a result, the processor 11 generates a behavioral feature vector for each online behavior.

[0116] In the subsequent S140 (see Figure 2), the processor 11 generates a user feature vector that statistically explains the target user's online behavior by performing mathematical processing on one or more behavioral feature vectors of the target user generated in S130. Through this mathematical processing, one or more behavioral feature vectors of the target user are transformed into a single representative user feature vector.

[0117] Through repeated processing, the processor 11 generates a user feature vector for each user that statistically best represents the online behavior of the corresponding user, based on the corresponding user's online behavior data.

[0118] The configuration of the information processing system 1 of the second embodiment has been described above. In this embodiment, since the content and type of online behavior are vectorized together, it is possible to generate user feature vectors related to online behavior that strongly consider the differences in type. That is, even if the content (keywords) is the same, each descriptive text can be represented as a vector so that online behaviors of different types are placed separately in the feature space of the user feature vector.

[0119] [Third Embodiment] Next, the information processing system 1 of the third embodiment will be described. The information processing system 1 of the third embodiment is configured to input the descriptive text of online behavior included in the online behavior data of the second embodiment as time series data into the time series prediction model M100, thereby probabilistically predicting the next online behavior that a corresponding user will perform, and generating user characteristic data that probabilistically describes the online behavior of the corresponding user, or user characteristic data that describes the predicted online behavior.

[0120] The hardware configuration of the information processing system 1 in this embodiment is the same as that of the second embodiment. Therefore, in the following description, we will selectively describe the configurations in the information processing system 1 of the third embodiment that differ from the second embodiment, and the same reference numerals will be used for the configurations common to the second embodiment, and their descriptions will be omitted.

[0121] In this embodiment, the processor 11 executes the prediction-related processing shown in Figure 9 based on instructions from the operator. When the prediction-related processing is started, the processor 11 retrieves a set of online behavioral data provided from the digital platform DP specified by the operator from the storage 15 (S510). The set of online behavioral data includes user-specific online behavioral data in the corresponding digital platform DP.

[0122] In the subsequent S520, the processor 11 selects a target user from among the user group corresponding to a set of online behavioral data, which is the user to be set as the target for processing from S530 onwards.

[0123] In the subsequent S530, the processor 11 references the target user's online behavior data and generates input data for the time series prediction model M100. The input data generated here is data in which a predetermined number (two or more) of descriptive texts for online actions, including the most recent online action performed by the target user, are arranged in time series.

[0124] When the most recent online action is an online action performed at time T[N], the processor 11 generates input data in which the descriptive text E_T[n] (n=1,2, ...N) for each of the N online actions from the online action performed at time T[1] to the online action performed at time T[N] is arranged in chronological order.

[0125] The expression E_T[n] means that the corresponding description text is a string (description text) that represents the content and type of online action performed at time T[n]. Here, n is an integer value that indicates the order in which online actions occurred, with the most recent online action being the Nth online action.

[0126] In other words, the input data consists of descriptive texts E_T[n] (n=1,2, ...,N) arranged in chronological order, which are strings that describe the content and type of online behavior at each time point T[n] from time point T[1] to time point T[N]. The input data is composed of descriptive texts E_T[1], E_T[2], ..., E_T[N] arranged in order.

[0127] For the generation of this input data, the online behavior data provided from each digital platform DP is structured as a chronological description of multiple online actions performed by the corresponding user. That is, the online behavior data is structured as a chronological arrangement of the aforementioned records corresponding to each of the multiple online actions performed by the corresponding user. The online behavior data may also be structured so that time information is attached to each online action.

[0128] In the subsequent S540, the processor 11 inputs the input data generated in S530 into the pre-trained time series prediction model M100, thereby generating user characteristic data that probabilistically represents the content and type of the next (i.e., N+1) online action that the target user will perform, through the time series prediction model M100.

[0129] The time series forecasting model M100 can be configured to output data that explains the content and type of each online action predicted to be performed by the target user at the next time point T[N+1], along with its probability of occurrence, based on the input data.

[0130] This time series prediction model M100 is pre-trained using ground truth data. In this embodiment, the time series prediction model M100 is, for example, an LSTM (Long Short-Term Memory) model. An LSTM model is a type of neural network. A conceptual diagram of the time series prediction model M100 is shown in Figure 10.

[0131] As shown in Figure 10, the time series forecasting model M100 comprises multiple embedding layers M110, an LSTM network layer M120, and an output layer M130. The LSTM network layer M120 comprises multiple LSTM layers M121.

[0132] Each of the multiple embedding layers M110 corresponds to a time point T[n] and receives the descriptive text E_T[n] contained in the input data as input, converting the descriptive text E_T[n] into an embedding representation. The embedding representation is the feature vector of the descriptive text E_T[n]. Each descriptive text E_T[n] (n=1,…,N) is input as an embedding representation to the corresponding LSTM layer M121 through the corresponding embedding layer M110.

[0133] The LSTM network layer M120 probabilistically predicts the next online action a target user will take, based on the time-series data of feature vectors of the explanatory text input as an embedded representation.

[0134] In S540, processor 11 functions as an embedding layer M110 by converting each descriptive text in the input data into a behavioral feature vector (embedding representation) and generating time-series data of multiple behavioral feature vectors corresponding to the input data. Processor 11 inputs this time-series data of behavioral feature vectors to the LSTM network layer M120. By functioning as the LSTM network layer M120, processor 11 probabilistically predicts the next online action that the target user will take.

[0135] The output of the LSTM network layer M120 is decoded through the output layer M130. The output layer M130 converts the output of the LSTM network layer M120 into data that explains the probability of each of the multiple online actions that the target user may perform at time T[N+1], and outputs it.

[0136] Based on the output of the output layer M130, the processor 11 generates user characteristic data that probabilistically describes the content and type of the next (i.e., N+1) online action that the target user will perform, or user characteristic data that describes the content and type of the next online action that the target user is predicted to perform.

[0137] One example of user feature data is a user feature vector, which represents the content and type of online behavior with the highest probability of execution, in a form similar to that of a behavior feature vector.

[0138] The second example of user characteristic data is data that, for each online behavior whose execution probability is above a predetermined value, represents the content and type of the corresponding online behavior in a form similar to the behavioral feature vector. In other words, the second example of user characteristic data is data that, for each online behavior whose execution probability is above a predetermined value, describes the corresponding behavioral feature vector as a user characteristic vector along with the execution probability.

[0139] The third example of user characteristic data is data that, for each online action whose execution probability is above a predetermined value, describes the content and type of the corresponding online action as a string, along with the execution probability.

[0140] The processing in S540 corresponds to the processor 11 generating user feature data that statistically explains the target user's online behavior through mathematical processing of time-series data of multiple behavioral feature vectors of the target user.

[0141] After completing the processing in S540, the processor 11 determines whether or not user characteristic data has been calculated for all users (S550). If the determination is negative (No in S550), the processor 11 returns to the processing in S520 and sets one of the users who was not selected as a target for processing from S530 onwards from the group of users corresponding to the group of online behavior data acquired in S510 as the new target user, and then executes the processing from S530 onwards.

[0142] Through the repetition of this process, the processor 11 generates user characteristic data for each user that statistically best represents the online behavior of the corresponding user, based on the corresponding user's online behavior data. Specifically, it generates user characteristic data that represents the predicted online behavior of the corresponding user.

[0143] The processor 11 generates user characteristic data for all users, then stores each user characteristic data in a database in the storage 15 (S560). The user database stored in the storage 15 describes the user characteristic data for each user of the corresponding digital platform DP, associated with the identification code of the corresponding user.

[0144] The processor 11 can perform the prediction-related processing shown in Figure 9 for each digital platform DP according to the operator's instructions.

[0145] According to the third embodiment, the information processing system 1 calculates and databases future predicted values ​​regarding online behavior, rather than statistically representative values ​​regarding past user online behavior. Therefore, appropriate advertising delivery and various measures based on predictions of future user behavior can be executed across multiple digital platform DPs based on the user databases of each digital platform DP.

[0146] [Fourth Embodiment] Next, the information processing system 1 of the fourth embodiment will be described. Similar to the third embodiment, the information processing system 1 of the fourth embodiment is configured to use a time-series prediction model M200 to probabilistically predict the next online action a user will take and to generate user characteristic data that explains the predicted online action.

[0147] However, the information processing system 1 of the fourth embodiment differs from the third embodiment in that it handles online behavioral data comprising behavioral description text and type data as described in the first embodiment, rather than online behavioral data comprising the description text as described in the second embodiment. Due to this difference, the time series forecasting model M200 of this embodiment has a different configuration from the time series forecasting model M100 of the third embodiment.

[0148] In the following, we will selectively describe the configuration of the information processing system 1 of the fourth embodiment, which differs from that of the third embodiment. Configurations common to the third embodiment are denoted by the same reference numerals as in the third embodiment, and their descriptions will be omitted.

[0149] The time series prediction model M200 used in this embodiment comprises a plurality of embedding layers M210, an LSTM network layer M120, and an output layer M130, as shown in Figure 11. The configuration of the LSTM network layer M120 and the output layer M130 is the same as in the third embodiment. On the other hand, the embedding layer M210 in this embodiment differs from that of the third embodiment.

[0150] According to this embodiment, in the processing of S530 (see Figure 9), data is generated as input data, which consists of action description texts and type data for a predetermined number (N) online actions, including the most recent online action performed by the target user, arranged in chronological order.

[0151] Each of the multiple embedding layers M210 (see Figure 11) is configured to receive inputs of action description text E1_T[n] and type data E2_T[n] included in the input data. Here, the expression "action description text E1_T[n]" means that the corresponding action description text is a string (action description text) that represents the content of the online action performed at time T[n]. The expression "type data E2_T[n]" means that the corresponding type data is a string (type data) that represents the type of online action performed at time T[n].

[0152] Each of the multiple embedding layers M210 comprises a first embedding layer M211, a second embedding layer M212, and a composite layer M213. The first embedding layer M211 is input with the behavior description text E1_T[n], and the second embedding layer M212 is input with the type data E2_T[n].

[0153] The synthesis layer M213 combines the embedded representation of the behavior description text E1_T[n] generated by the first embedding layer M211 with the embedded representation of the type data E2_T[n] generated by the second embedding layer M212. The synthesis is achieved by adding the embedded representation of the behavior description text E1_T[n] and the embedded representation of the type data E2_T[n].

[0154] In this embodiment, the synthesized embedding representation is input to the corresponding LSTM layer M121. The output layer M130 outputs data that describes the probability of each of the multiple online actions that the target user may perform at time T[N+1], similar to the third embodiment.

[0155] The processor 11 uses the time series forecasting model M200 configured in this way to perform forecasting-related processing, similar to the third embodiment. Therefore, the information processing system 1 of this embodiment achieves the same effects as the third embodiment.

[0156] As a variation of the fourth embodiment, the time series forecasting model M300 shown in Figure 12 may be used. The only difference between the time series forecasting model M300 shown in Figure 12 and the time series forecasting model M200 shown in Figure 11 is the configuration of the embedding layer M310.

[0157] The modified time series prediction model M300 comprises multiple embedding layers M310, an LSTM network layer M120, and an output layer M130. Each of the multiple embedding layers M310 comprises a main embedding layer M311 and an additional layer M312. The main embedding layer M311 is input to the action description text E1_T[n]. The embedded representation of the action description text E1_T[n] generated by the main embedding layer M311 is input to the additional layer M312.

[0158] The additional layer M312 receives not only the embedding representation of the behavior description text E1_T[n] but also the type data E2_T[n]. Based on the type data E2_T[n], the additional layer M312 extends the embedding representation of the behavior description text E1_T[n] by adding a vector element representing the type of online behavior to the embedding representation of the behavior description text E1_T[n]. In the modified example, the extended embedding representation is then input to the corresponding LSTM layer M121. The modified information processing system 1 also achieves the same effects as the third embodiment.

[0159] [Fifth Embodiment] Next, the information processing system 1 of the fifth embodiment will be described. Similar to the third and fourth embodiments, the information processing system 1 of the fifth embodiment is configured to use a time-series prediction model M400 to probabilistically predict the next online action a user will take and to generate user characteristic data that explains the predicted online action.

[0160] However, the information processing system 1 of the fifth embodiment does not perform processing on the embedded representation of the behavior description text by the synthesis layer M213 or the additional layer M312, as in the fourth embodiment. The information processing system 1 of the fifth embodiment is configured to input time-series data of embedded representations, in which the embedded representation of the behavior description text E1_T[n] and the embedded representation of the type data E2_T[n] are arranged alternately, into the LSTM network layer M120.

[0161] Figure 13 shows the time series forecasting model M400 of this embodiment. This time series forecasting model M400 comprises a plurality of embedding layers M410, an LSTM network layer M420, and an output layer M130.

[0162] As shown in the figure, the time series prediction model M400 has 2N embedding layers M410, which is twice the number of embedding layers M410, to handle input data in which the behavioral description text and type data of N online behaviors are arranged. The LSTM network layer M420 has a number of LSTM layers M421 corresponding to the number of embedding layers M410.

[0163] The array in the embedded layer M410 is input to time-series data of behavioral description text E1_T[n] and type data E2_T[n], in which behavioral description text E1_T[n] and type data E2_T[n] are arranged alternately.

[0164] As a result, in the embedding layer M410, for each online action, an embedding representation of the action description text E1_T[n] and an embedding representation of the type data E2_T[n] are generated as feature vectors. The LSTM network layer M420 is input time-series data in which the embedding representations of the action description text E1_T[n] and the embedding representations of the type data E2_T[n] are arranged alternately.

[0165] In other words, in S540, the processor 11, in order to function as the embedding layer M410, vectorizes the action description text E1_T[n] and type data E2_T[n] for each online action contained in the input data and converts them into embedding representations, generating time-series data in which the embedding representations of the action description text E1_T[n] and the embedding representations of the type data E2_T[n] are arranged alternately. The processor 11 inputs this time-series data to the LSTM network layer M420.

[0166] In response to such input, the output layer M130 outputs data that explains the probability of each of the multiple online actions that the target user may perform at time T[N+1], similar to the third embodiment.

[0167] Based on the output from the output layer M130, the processor 11 generates, in S540, user characteristic data that probabilistically represents the content and type of the next (i.e., N+1) online action that the target user will perform, or user characteristic data that describes the predicted online actions of the target user. The information processing system 1 of this embodiment also achieves the same effects as the third and fourth embodiments.

[0168] [Sixth Embodiment] Next, the information processing system 1 of the sixth embodiment will be described. The information processing system 1 of the sixth embodiment generates user feature vectors and generates a user database, similar to the first or second embodiment. On the other hand, the information processing system 1 of the sixth embodiment differs from the first or second embodiment in that it has a function to correct the user feature vectors, taking into account the possibility that there may be differences in the meaning of keywords (search strings, etc.) in online behavior between digital platforms DP.

[0169] In the following description of the sixth embodiment, we will selectively describe the configuration related to the user feature vector correction function that is added to the information processing system 1 of the first or second embodiment. The configuration of the information processing system 1 of the sixth embodiment that is not mentioned below may be understood to be the same as that of the first or second embodiment.

[0170] The correction function is used to identify a specific group of users on the first digital platform DP from a group of users on a second digital platform DP that is different from the first digital platform DP. This identification function is realized when the processor 11 executes the correction search process shown in Figure 14 based on instructions from the operator.

[0171] When the correction search process is started, the processor 11 extracts a specific type of group from the user group of the first digital platform DP specified by the operator (S610). For example, as described in the first embodiment, if the user database of the first digital platform DP is accompanied by questionnaire responses, the processor 11 can refer to the questionnaire responses and extract one or more users who give specific answers from the user group of the first digital platform DP as a specific type of group. Hereinafter, the specific type of group (i.e., user set) extracted from the user group of the first digital platform DP will be referred to as the first extracted user group.

[0172] In the subsequent S620, the processor 11 clusters the first extracted user group into multiple clusters based on the frequency of occurrence of multiple keywords in each user's online behavior. Each of the multiple keywords can be understood as a string appearing in the behavior description text. The frequency of occurrence can be understood as the number of times the keyword appears in the set of online behavior records over a predetermined period described by the online behavior data. The multiple keywords can be understood as a predetermined number of keywords, listed in descending order of frequency of occurrence from the keyword with the highest frequency of occurrence throughout the entire first extracted user group. Each cluster corresponds to a set of users whose combinations of the frequency of occurrence of each of the multiple keywords are similar.

[0173] In the subsequent S630, the processor 11 calculates the cluster centroids of the user feature vectors of one or more users belonging to the corresponding cluster from the first extracted user group for each cluster, and generates a first cluster centroid table that describes the cluster centroids for each cluster for the first digital platform DP.

[0174] In the subsequent S640, the processor 11 calculates a correction vector to correct the cluster centroid of each cluster in a specific type of group in the first digital platform DP to the cluster centroid of each cluster in a specific type of group in the second digital platform DP. In S640, the processor 11 executes the correction vector calculation process shown in Figure 15.

[0175] When the correction vector calculation process is started, the processor 11 identifies a group of keywords whose frequency of occurrence in the corresponding cluster is above a threshold, for each cluster defined by the clustering in S620.

[0176] In other words, for each cluster, the processor 11 identifies a group of keywords (i.e., strings) from a set of behavioral description texts belonging to the user group in the corresponding cluster that have an occurrence frequency above a threshold. For each cluster, the processor 11 generates a keyword list listing the keywords identified above (S710).

[0177] In the subsequent S720, the processor 11 extracts, for each cluster, a group of users from the entire user base of the first digital platform DP whose keyword frequency listed in the keyword list is above a predetermined level. At this time, the processor 11 can refer to a set of online behavioral data provided from the first digital platform DP.

[0178] In the subsequent S730, the processor 11 refers to the user database of the first digital platform DP and calculates the average vector of the user feature vectors of the extracted user groups for each cluster, thereby calculating the centroid of the user feature vectors for each cluster with respect to all users of the first digital platform DP. The centroid calculated here for each cluster is referred to as the first centroid.

[0179] In the subsequent S740, the processor 11 applies the cluster-specific keyword list generated in S710 to the user group of the second digital platform DP to calculate the centroid of the cluster-specific user feature vector for all users of the first digital platform DP. The cluster-specific centroid calculated here is referred to as the second centroid.

[0180] In other words, in S740, the processor 11 extracts a set of users from the entire user base of the second digital platform DP, for each cluster, in which the frequency of keywords listed in the keyword list is above a predetermined level. At this time, the processor 11 can refer to a set of online behavioral data provided from the second digital platform DP.

[0181] The processor 11 further refers to the user database of the second digital platform DP and calculates the mean vector of the user feature vectors of the extracted user groups for each cluster, thereby calculating the centroid of the user feature vectors for each cluster with respect to all users of the second digital platform DP. The centroid of each cluster calculated here is referred to as the second centroid.

[0182] In the subsequent S750, the processor 11 calculates the difference vector from the first centroid of the second double center as a correction vector for each cluster. After that, the processor 11 terminates the correction vector calculation process.

[0183] In this way, after calculating the correction vector for each cluster in S640, in the following step 650, the processor 11 corrects the cluster centroid for each cluster related to the first extracted user group calculated in S630 using the correction vector for the corresponding cluster. That is, for each cluster, the corrected cluster centroid is calculated by adding the correction vector for the corresponding cluster to the corresponding cluster centroid. The corrected cluster centroid corresponds to the value obtained by adding the correction vector to the cluster centroid before correction.

[0184] In S650, the processor 11 further generates a table describing the corrected cluster centroid for each cluster, which is then used as a second cluster centroid table describing the cluster centroid for each cluster of the second digital platform DP.

[0185] In the subsequent S660, the processor 11 refers to the second cluster centroid table and the user database of the second digital platform DP to extract, for each cluster, a group of users from the user group of the second digital platform DP that have user feature vectors belonging to the corresponding cluster.

[0186] As a result, the processor 11 identifies the extracted user group as a group in the second digital platform DP corresponding to a specific type of group in the first digital platform DP, and outputs the identification result. This identification result is useful, for example, for identifying the user group in the second digital platform DP that corresponds to a user group narrowed down based on data provided from the first digital platform DP. This identification is useful, for example, for setting advertising delivery targets.

[0187] As described above, according to this embodiment, the relationship between the user group in the first digital platform DP and the user group in the second digital platform DP is determined by using a cluster centroid corrected for differences in the keyword usage environment between digital platform DPs through a corrected search process.

[0188] Therefore, according to this embodiment, it is possible to more appropriately perform user behavior analysis and set advertising delivery targets across multiple digital platform DPs.

[0189] [Other embodiments] While exemplary embodiments of the present disclosure have been described above, the present disclosure is not limited to the embodiments described above, and various other forms can be adopted.

[0190] For example, the user feature vectors relating to the predicted online behavior of each user obtained from the time series prediction models M100, M200, M300, and M400 of the third, fourth, and fifth embodiments may be corrected to take into account the differences in environments between digital platforms DP, similar to the sixth embodiment.

[0191] For example, the operator can prepare a set of texts from a specific group of users over a certain period of time, collected in the first digital platform DP. The set of texts may be a group of behavioral description texts or a group of descriptive texts.

[0192] Information processing system 1 can generate multiple first vectors, which are multiple user feature vectors obtained by inputting each of the multiple input data based on this text set into the time series forecasting models M100, M200, M300, and M400 of the first digital platform DP, and multiple second vectors, which are multiple user feature vectors obtained by inputting each of the same multiple input data into the time series forecasting models M100, M200, M300, and M400 of the second digital platform DP. Information processing system 1 can calculate the difference vector of the second vector with respect to the first vector for each input data, and generate a weighted average of the difference vectors as a correction vector.

[0193] The information processing system 1 can correct the user feature vector in the first digital platform DP to a user feature vector that is suitable for the environment of the second digital platform DP by adding a correction vector to the user feature vector in the first digital platform DP.

[0194] The information processing system 1 described above handles online behavior data observed on multiple digital platforms DP. However, the online behavior data that explains online behavior observed on multiple digital platforms DP may be understood as online behavior data that explains online behavior observed in different environments.

[0195] In other words, the above-described embodiment can be used for analyzing online behavior observed in multiple different environments.

[0196] The function of one component in the above embodiment may be distributed among multiple components. The functions of multiple components may be integrated into one component. Some parts of the configuration of the above embodiment may be omitted. At least some parts of the configuration of the above embodiment may be added to or replaced by the configuration of other above embodiments. Any aspect of the technical concept specified by the wording of the claims constitutes an embodiment of the present disclosure.

[0197] [Technical Concept Disclosed in This Specified Specification] This specification can be understood to disclose the following technical concepts: [Item 1] An acquisition unit is configured to acquire descriptive data for each user that describes one or more online actions of the corresponding user, for multiple users. A vector generation unit is configured to generate one or more feature vectors corresponding to one or more online actions by generating feature vectors for each of the above users based on one or more linguistic pieces of information describing the one or more online actions of the corresponding user included in the description data, A vector processing unit is configured to generate user feature data that statistically explains the online behavior of the corresponding user by performing mathematical processing on the one or more feature vectors generated by the vector generation unit for each user. An information processing system equipped with the following features. [Item 2] The information processing system described in item 1, The descriptive data includes, for each of the one or more online actions, an action description text which is a string of characters that describes the content of the corresponding online action, and type data which describes the type of the corresponding online action. The vector generation unit generates feature vectors for each online action based on the action description text and the type data, thereby generating one or more feature vectors corresponding to one or more online actions. Information processing system. [Item 3] The information processing system described in item 2, The feature vector, based on the behavior description text and the type data, is a feature vector obtained by vectorizing the behavior description text and then increasing its dimensionality based on the type data. This increase in dimensionality is achieved by adding vector elements representing the corresponding online behavior type to the text vector. Information processing system. [Item 4] The information processing system described in item 1, The descriptive data includes, for each of the one or more online actions, descriptive text which is a string of characters that describes the content of the corresponding online action and the type of the corresponding online action. The vector generation unit generates one or more feature vectors corresponding to one or more online actions by vectorizing the descriptive text for each online action. Information processing system. [Item 5] An information processing system as described in item 2 or item 3, The aforementioned action description text is an information processing system that describes the content of the corresponding online action, specifically the input from the corresponding user in the corresponding online action, using the aforementioned string of characters. [Item 6] The information processing system described in item 4, The descriptive text is an information processing system that describes the content of the corresponding online action, specifically the input from the corresponding user in the corresponding online action, using the string of characters. [Item 7] An information processing system described in any one of items 1 to 6, The vector processing unit is configured to generate a single user feature vector representing the one or more feature vectors as user feature data by performing the mathematical processing on the one or more feature vectors. [Item 8] The information processing system described in item 7, Based on a comparison between the user feature vector for each user and the reference vector, an extraction unit extracts users from among the multiple users who have online behavior characteristics similar to the online behavior characteristics described by the reference vector. An information processing system further equipped with [the following features]. [Item 9] The information processing system described in item 7, A model building unit constructs an estimation model that uses the user feature vectors for each user as training data, and has at least a portion of the vector elements of the user feature vectors as explanatory variables, for estimating a predetermined feature possessed by the user that is explained by the explanatory variables. An information processing system further equipped with [the following features]. [Item 10] An information processing system described in any one of items 1 to 6, The vector processing unit is configured to generate, by performing the mathematical processing on one or more feature vectors, user feature data that probabilistically describes the online behavior of the corresponding user, or user feature data that describes the predicted online behavior of the corresponding user, as user feature data. Information processing system. [Item 11] The information processing system described in item 1, The aforementioned one or more online actions are multiple online actions, The one or more feature vectors mentioned above are multiple feature vectors corresponding to the multiple online actions, and include a feature vector for each online action. The explanatory data describes the corresponding user's multiple online actions in chronological order, and for each online action, it includes linguistic information describing the corresponding online action. The vector generation unit generates a feature vector based on the language information for each online action, thereby generating time-series data of the multiple feature vectors corresponding to the multiple online actions. The vector processing unit inputs the time series data into a time series prediction model to generate user characteristic data that probabilistically describes the online behavior of the corresponding user, or user characteristic data that describes the predicted online behavior of the corresponding user. Information processing system. [Item 12] The information processing system described in item 1, The aforementioned one or more online actions are multiple online actions, The one or more feature vectors mentioned above are multiple feature vectors corresponding to the multiple online actions, and include a feature vector for each online action. The description data includes, for each online action, action description text which is a string of characters that describes the content of the corresponding online action, and type data which describes the type of the corresponding online action. The vector generation unit generates, for each online action, a first feature vector obtained by vectorizing the action description text, and a second feature vector obtained by vectorizing the type data, as feature vectors for each online action. The vector processing unit inputs time-series data, in which the first and second feature vectors are arranged alternately, into a time-series prediction model to generate user feature data that probabilistically describes the online behavior of the corresponding user, or user feature data that describes the predicted online behavior of the corresponding user. Information processing system. [Item 13] The information processing system described in item 7, The aforementioned group of users includes a first user group and a second user group, Each of the one or more linguistic pieces of information includes a keyword that describes the corresponding online behavior among the one or more online behaviors, The aforementioned information processing system further, A clustering unit configured to cluster a specific type of group within the first user group into multiple clusters based on the frequency of occurrence of multiple keywords in the online behavior of each user belonging to the specific type of group, A centroid calculation unit is configured to calculate the centroid of the user feature vectors of one or more users belonging to the corresponding cluster as the cluster centroid for each cluster. A correction unit configured to correct the cluster centroid for each cluster based on the user feature vector of each user in the first user group and the user feature vector of each user in the second user group, A discrimination unit configured to identify a specific type of group within the second user group based on the corrected cluster centroid, An information processing system equipped with the following features. [Item 14] The information processing system described in item 13, The correction unit, For each of the aforementioned network, Identify the set of keywords that appear in the corresponding cluster, Among the first user group, the centroid of the user feature vector of the set of users that use the keyword group is calculated as the first centroid. Among the second user group, the centroid of the user feature vector of the set of users that use the keyword group is calculated as the second double center. Based on the difference between the first centroid and the second centroid, the cluster centroid of the corresponding cluster is corrected. Information processing system. [Item 15] A computer program for causing a computer to function as the acquisition unit, the vector generation unit, and the vector processing unit, which are included in the information processing system described in any one of items 1 to 14. [Item 16] A method of information processing performed by a computer, For multiple users, obtain descriptive data for each user that explains one or more of the corresponding user's online behaviors, For each user, a feature vector is generated based on each of the one or more linguistic pieces of information that describe the one or more online actions of the corresponding user included in the descriptive data, thereby generating one or more feature vectors corresponding to the one or more online actions. For each user, user feature data is generated that statistically explains the online behavior of the corresponding user by performing mathematical processing on the one or more feature vectors that have been generated. Information processing methods including [Explanation of Symbols]

[0198] 1... Information processing system, 11... Processor, 13... Memory, 15... Storage, 17... User interface, 19... Communication interface, DP... Digital platform, M100, M200, M300, M400... Time series forecasting model, M110, M210, M310, M410... Embedding layer, M120, M420... LSTM network layer, M130... Output layer.

Claims

1. An acquisition unit is configured to acquire descriptive data for each user that describes one or more online actions of the corresponding user, for multiple users. A vector generation unit is configured to generate one or more feature vectors corresponding to one or more online actions by generating feature vectors for each of the above users based on one or more linguistic pieces of information describing the one or more online actions of the corresponding user included in the description data, A vector processing unit is configured to generate user feature data describing the online behavior of the corresponding user for each user by performing statistical processing on the one or more feature vectors generated by the vector generation unit. An information processing system equipped with the following features.

2. An acquisition unit is configured to acquire descriptive data for each user that describes one or more online actions of the corresponding user, for multiple users. A vector generation unit is configured to generate one or more feature vectors corresponding to one or more online actions by generating feature vectors for each of the above users based on one or more linguistic pieces of information describing the one or more online actions of the corresponding user included in the description data, A vector processing unit is configured to calculate a representative vector of the one or more feature vectors generated by the vector generation unit, for each user, as a user feature vector that statistically explains the online behavior of the corresponding user. An information processing system equipped with the following features.

3. The information processing system according to claim 2, The descriptive data includes, for each of the one or more online actions, an action description text which is a string of characters that describes the content of the corresponding online action, and type data which describes the type of the corresponding online action. The vector generation unit generates one or more feature vectors corresponding to one or more online actions by generating a feature vector based on the action description text and the type data for each online action. Information processing system.

4. The information processing system according to claim 3, The feature vector, based on the behavior description text and the type data, is a feature vector obtained by vectorizing the behavior description text and then increasing its dimensionality based on the type data. This increase in dimensionality is achieved by adding vector elements representing the corresponding online behavior type to the text vector. Information processing system.

5. The information processing system according to claim 2, The descriptive data includes, for each of the one or more online actions, descriptive text which is a string of characters that describes the content of the corresponding online action and the type of the corresponding online action. The vector generation unit generates one or more feature vectors corresponding to one or more online actions by vectorizing the descriptive text for each online action. Information processing system.

6. The information processing system according to claim 3, The aforementioned action description text is an information processing system that describes the content of the corresponding online action, specifically the input from the corresponding user in the corresponding online action, using the aforementioned string of characters.

7. The information processing system according to claim 5, The descriptive text is an information processing system that describes the content of the corresponding online action, specifically the input from the corresponding user in the corresponding online action, using the string of characters.

8. An information processing system according to any one of claims 2 to 7, Based on a comparison between the user feature vector for each user and the reference vector, an extraction unit extracts users from among the multiple users who have online behavior characteristics similar to the online behavior characteristics described by the reference vector. An information processing system further equipped with [the following features].

9. An information processing system according to any one of claims 2 to 7, A model building unit constructs an estimation model that uses the user feature vectors for each user as training data, and has at least a portion of the vector elements of the user feature vectors as explanatory variables, for estimating a predetermined feature possessed by the user that is explained by the explanatory variables. An information processing system further equipped with [the following features].

10. An acquisition unit configured to acquire descriptive data for each user that describes one or more online actions of a corresponding user, A vector generation unit is configured to generate one or more feature vectors corresponding to one or more online actions by generating feature vectors for each of the above users based on one or more linguistic pieces of information describing the one or more online actions of the corresponding user included in the description data, A vector processing unit is configured to generate user feature data that probabilistically explains the online behavior of the corresponding user, or user feature data that explains the predicted online behavior of the corresponding user, by inputting one or more feature vectors generated by the vector generation unit into a prediction model for each user. Equipped with Information processing system.

11. An acquisition unit configured to acquire descriptive data for each of multiple users, describing a series of online actions of a corresponding user, wherein for each online action, the acquisition unit includes descriptive data containing linguistic information describing the corresponding online action. A vector generation unit is configured to generate time-series data of the multiple feature vectors corresponding to the multiple online actions by generating a plurality of feature vectors for each user based on a plurality of linguistic information that describes the plurality of online actions of the corresponding user included in the description data, A vector processing unit is configured to input the time series data generated by the vector generation unit into a time series prediction model for each user, thereby generating user characteristic data that probabilistically explains the online behavior of the corresponding user, or user characteristic data that explains the predicted online behavior of the corresponding user. Equipped with Information processing system.

12. An acquisition unit configured to acquire descriptive data for each of multiple users, for each user, which describes multiple online actions of a corresponding user, and for each online action, includes action description text, which is a string describing the content of the corresponding online action, and type data, which describes the type of the corresponding online action. For each user, a vector generation unit is configured to generate, for each online action, a first feature vector obtained by vectorizing the action description text and a second feature vector obtained by vectorizing the type data, with respect to the corresponding user's multiple online actions described in the description data. A vector processing unit is configured to generate user feature data that probabilistically explains the online behavior of the corresponding user, or user feature data that explains the predicted online behavior of the corresponding user, by inputting time series data obtained by alternately arranging the first feature vector and the second feature vector generated by the vector generation unit into a time series prediction model for each user. Equipped with Information processing system.

13. The information processing system according to claim 2, The aforementioned group of users includes a first user group and a second user group, Each of the one or more linguistic pieces of information includes a keyword that describes the corresponding online behavior among the one or more online behaviors, The aforementioned information processing system further, A clustering unit is configured to cluster a specific type of group within the first user group into multiple clusters based on the frequency of occurrence of multiple keywords in the online behavior of each user belonging to the specific type of group. A centroid calculation unit is configured to calculate the centroid of the user feature vectors of one or more users belonging to the corresponding cluster as the cluster centroid for each cluster. A correction unit configured to correct the cluster centroid for each cluster based on the user feature vector of each user in the first user group and the user feature vector of each user in the second user group, A discrimination unit configured to identify a specific type of group within the second user group based on the corrected cluster centroid, An information processing system equipped with the following features.

14. The information processing system according to claim 13, The correction unit, For each of the aforementioned network, Identify the set of keywords that appear in the corresponding cluster, Among the first user group, the centroid of the user feature vector of the set of users that use the keyword group is calculated as the first centroid. Among the second user group, the centroid of the user feature vector of the set of users that use the keyword group is calculated as the second double center. Based on the difference between the first centroid and the second centroid, the cluster centroid of the corresponding cluster is corrected. Information processing system.

15. A computer program for causing a computer to function as the acquisition unit, the vector generation unit, and the vector processing unit, which are included in the information processing system according to any one of claims 1 to 7 or any one of claims 10 to 12.

16. A method of information processing performed by a computer, For multiple users, obtain descriptive data for each user that explains one or more of the corresponding user's online behaviors, For each user, a feature vector is generated based on each of the one or more linguistic pieces of information that describe the one or more online actions of the corresponding user included in the descriptive data, thereby generating one or more feature vectors corresponding to the one or more online actions. For each user, user feature data describing the corresponding user's online behavior is generated by performing statistical processing on the one or more feature vectors that have been generated. Information processing methods including

Citation Information

Patent Citations

  • Low-Entropy Browsing History for Pseudo-Personalization of Content

    JP2022542624A

  • Privacy-preserving machine learning labeling

    JP2023514039A

  • Information processing system, information processing device, program, application software, terminal device, and information processing method

    JP2025024822A

  • Program, method, information processing device, and system

    JP7669075B1

  • Information processing device, display method, and program

    JP2020113218A