Information processing method, program, and information processing device
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2024-10-07
- Publication Date
- 2026-04-14
AI Technical Summary
Existing systems for analyzing user behavior on web pages generate complex data structures with multiple events and parameters mixed together, making it difficult for analysts to efficiently extract and analyze relevant information.
An information processing method that involves obtaining analysis data, generating event data with attributes and parameters, aggregating parameters based on predetermined rules, and outputting aggregated data in a format that is easier for analysts to use.
The method improves the work efficiency of analysts by simplifying the data structure and making it easier to analyze user behavior on web pages, thereby enhancing the ability to derive insights from the data.
Abstract
Description
Information processing method, program, and information processing device
[0001] The disclosed technology relates to an information processing method, a program, and an information processing device.
[0002] Conventionally, systems for analyzing user behavior on web pages are known. For example, a content provider system interacts with a network of websites to provide users with behavior-based content, where website operators add widgets to selected web pages of the sites, and the widgets report user-generated events to the content provider system, which analyzes the reported events and detects behavioral relevance of specific websites, web pages, products, and other types of items (see Patent Literature 1).
[0003] Special Publication No. 2011-508925
[0004] There are services that provide analytical data on the analysis of user behavior on web pages. For example, Google Analytics (registered trademark) 4 (GA4) analyzes data on an event-by-event basis, using events related to user behavior on web pages as the axis of measurement, and provides analytical data. However, while the provided data associates events with parameters, multiple events are mixed in the same file, and even the same event may contain multiple parameters, making the data structure very difficult to use.
[0005] Therefore, the disclosed technology aims to improve the work efficiency of analysts by generating data that is easy for analysts to use based on analytical data of user behavior on web pages.
[0006] In one aspect of the disclosure, an information processing method includes an information processing device acquiring analytical data having items including each event related to user behavior on a web page and each parameter related to each event, generating event data for each event based on the analytical data, including event attributes and parameters, aggregating parameters for the event data based on predetermined rules associated with the corresponding event, generating aggregated data based on the event data including the aggregated parameters, and outputting the aggregated data.
[0007] According to the disclosed technology, data that is easy for analysts to use can be generated based on analysis data of user behavior on web pages, thereby improving the work efficiency of analysts.
[0008] FIG. 1 is a diagram illustrating an example of a configuration of an information processing system according to an embodiment. FIG. 2 is a block diagram illustrating an example of a server according to an embodiment. FIG. 3 is a block diagram illustrating an example of a user device according to an embodiment. FIG. 4 is a diagram illustrating an example of analysis data according to an embodiment. FIG. 5 is a diagram illustrating an example of event data generated based on simplified analysis data according to an embodiment. FIG. 6 is a diagram illustrating an example of an event list according to an embodiment. FIG. 7 is a diagram illustrating an example of each event data according to an embodiment. FIG. 8 is a diagram illustrating an example of comprehensive aggregate data according to an embodiment. FIG. 9 is a diagram illustrating an example of each event aggregate data according to an embodiment. FIG. 10 is a flowchart illustrating an example of processing related to a server according to an embodiment.
[0009] Preferred embodiments of the present disclosure will be described with reference to the accompanying drawings, in which the same reference numerals denote the same or similar configurations.
[0010] [Embodiment] <System Configuration> Fig. 1 is a diagram illustrating an example of the configuration of an information processing system 1 according to an embodiment of the present disclosure. As illustrated in Fig. 1, the information processing system 1 includes a server 10 and one or more user devices 20A, 20B, and 20C. The server 10 and the one or more user devices 20 can transmit and receive data to and from each other via a network N. The server 10 may be configured with multiple processing devices (including databases). The number of user devices 20A, 20B, and 20C may be any number, and when they are not individually distinguished, they will be referred to as user devices 20.
[0011] The server 10 acquires analytical data from the user device 20 by uploading or the like, generates data that is easy for an analyst to analyze based on the analytical data, and outputs the generated data to the user device 20. The analytical data includes, for example, a report generated by GA4. The server 10 may be configured as a server or database on the cloud.
[0012] The user device 20 is, for example, an information processing device (or a processing terminal) used to analyze user behavior on a web page provided by the user. Examples of the user device 20 include a personal computer, a tablet terminal, and a mobile terminal such as a smartphone.
[0013] The user device 20 acquires analytical data from a service that analyzes user behavior on a web page, such as GA4. The user device 20 accesses a web page provided by the server 10 using, for example, a web browser and uploads analytical data using this web page. The user device 20 downloads and acquires aggregated data from the server 10 that has been processed to facilitate analysis of user behavior.
[0014] According to the information processing system 1 described above, the analyst's work efficiency can be improved by generating data that is easy for the analyst to use based on the analysis data of user behavior via web pages. Below, each component of the information processing system 1 that enables the above-mentioned processing to be executed will be described in detail.
[0015] 2 is a block diagram showing an example of a server 10 according to an embodiment of the present disclosure. For example, the server 10 includes one or more processors (CPUs: Central Processing Units) 110, one or more network communication interfaces 120, a memory 130, and one or more communication buses 170 for interconnecting these components.
[0016] The server 10 may optionally include a user interface 150. The user interface 150 includes a display and an input device (such as a keyboard and / or a mouse or some other pointing device).
[0017] Memory 130 may be, for example, a high-speed random access memory such as a DRAM, SRAM, or other random access solid-state memory device, or may be a non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices, or may be a non-transitory computer-readable recording medium.
[0018] The memory 130 stores data used by the information processing system 1. For example, the memory 130 stores various data such as analysis data, an event list, event data, and aggregated data. Details of each data will be described later using FIGS. 4 to 10.
[0019] Another example of memory 130 may be one or more storage devices (e.g., databases) separate from processor 110. In some embodiments, memory 130 stores programs, modules, and data structures, or a subset thereof, that are executed by processor 110.
[0020] The processor 110 executes a program stored in the memory 130 to constitute a control unit 112, and the control unit 112 constitutes an acquisition unit 113, a first generation unit 114, a counting unit 115, a second generation unit 116, and an output unit 117.
[0021] The control unit 112 controls the process of generating aggregated data that makes it easy to analyze user behavior based on the analysis data. For example, the control unit 112 logs in to the service that generates the aggregated data and processes billing for system usage.
[0022] The acquisition unit 113 acquires analytical data having items including events related to user behavior on a web page and parameters related to each event. For example, the acquisition unit 113 acquires analytical data using GA4. The analytical data contains a mixture of multiple events, and each event contains multiple attributes and parameters corresponding to these attributes in an unorganized manner. Furthermore, the analytical data may have a structure in which tables are nested within tables, or the items may be insufficiently defined (including strings, percentages, numerical values, etc. without distinction). The analytical data will be described later with reference to FIG. 4.
[0023] Events include, for example, scrolls, file downloads, clicks, page views, session starts, user-defined attributes (dimensions), and the like.
[0024] The first generating unit 114 generates event data including event attributes and parameters for each event based on the analysis data acquired by the acquiring unit 113. For example, the first generating unit 114 classifies the events by using the event name and generates event data for each event. The event data will be described later with reference to FIG. 5 or FIG. 7.
[0025] The aggregation unit 115 aggregates parameters based on a predetermined rule associated with the corresponding event for each piece of event data generated by the first generation unit 114. For example, the aggregation unit 115 acquires predetermined event data, analyzes data included in the event data, converts the data format to make it easier for a user (analyst) to analyze, and extracts attributes and parameters to reorganize the data structure.
[0026] A predetermined rule is associated with each event and defines how to process the parameters. For example, if the parameter is a numeric value, a numerical calculation process is defined, and if the parameter is a character string, a process related to the character string is defined.
[0027] The second generating unit 116 generates aggregated data based on the event data including the parameters aggregated by the aggregating unit 115. For example, the second generating unit 116 may generate the aggregated data in a data format (one-line record) in which each attribute of the event is represented as an item in a column. The aggregated data will be described later with reference to, for example, FIGS. 8 and 9 .
[0028] The output unit 117 outputs the aggregated data. For example, when a download request is made by the user device 20, the output unit 117 outputs the aggregated data to the user device 20. The aggregated data has one of the file formats that allows data editing or analysis, such as CSV or Excel.
[0029] According to the above process, data that is easy for analysts to use is generated based on the analysis data of user behavior on web pages, thereby improving the work efficiency of analysts.
[0030] The first generating unit 114 may classify the events according to the registered events based on an event list in which each event is registered. The event list is stored in the memory 130 and may include events specified by a user (analyst) or events extracted from the analysis data.
[0031] By using the event list through the above processing, the analyst can efficiently extract events that he or she wishes to analyze.
[0032] Furthermore, the first generating unit 114 may include a list generating unit that generates an event list based on predetermined events. The first generating unit 114 may determine whether each event included in the analysis data acquired by the acquiring unit 113 is included in the event list. When it is determined that there is a predetermined event that is not included in the event list, the first generating unit 114 may register the predetermined event in the event list.
[0033] The pre-defined events may include events that are included in the analysis data by default. Furthermore, the user can freely define the events to be analyzed in relation to the acquisition of analysis data. Events defined by the user are not included in the default event list. Even in this case, the first generation unit 114 can identify the user-defined events by comparing them with the events included in the event list, and can register these unique events in the event list.
[0034] The first generating unit 114 may generate an event list based on analysis data that includes events preset by a user behavior analysis service. In this case, no events are included in the initial event list, but the first generating unit 114 may extract events included in the analysis data one by one and register these events in the event list.
[0035] The above process automatically generates an event list, making it possible to register new events contained in the analysis data in the event list without placing a burden on the user.
[0036] The first generation unit 114 may also set events that are not to be registered in the event list. For example, the first generation unit 114 determines whether an event extracted from the analysis data corresponds to an event that cannot be registered, and if the event is an event that cannot be registered, the first generation unit 114 does not register the event in the event list. This allows events that the user does not need to analyze to be excluded from the analysis target, enabling efficient event analysis.
[0037] When the plurality of events indicate the same scroll event, the predetermined rule may include obtaining the maximum value from among a plurality of numerical values included in a plurality of parameters corresponding to the plurality of events.
[0038] For example, if there are identical scroll events (e.g., the same user, the same page, the same session, etc.) in the events included in the scroll event data, the aggregation unit 115 obtains the maximum value of the scroll rate (the ratio of scrolls to the entire page) included in the parameters and integrates the identical scroll events. Note that there are cases where the parameters simply contain numbers, and in this case, the aggregation unit 115 may determine the unit of the parameter's numerical value based on the meaning of the event included in the event data (e.g., column name).
[0039] By performing the above process, when there are multiple identical scroll events, only the maximum scroll rate is retained, thereby reducing the amount of data and improving the efficiency of analysis.
[0040] When multiple events indicate the same download event, the predetermined rule may include obtaining the sum of multiple numerical values included in multiple parameters corresponding to the multiple events.
[0041] For example, if there are identical download events among the events included in the download event data, the counting unit 115 obtains the total number of downloads included in the parameters and integrates the identical download events. For example, the counting unit 115 calculates the total number of downloads of each file that can be downloaded from the same web page.
[0042] By performing the above process, when there are multiple download events for each file from the same web page, the total number of downloads for each file is tallied and only the total value is retained in the aggregated data, thereby reducing the amount of data and improving the efficiency of analysis. The aggregation unit 115 can also aggregate the number of downloads for the same file. For example, the aggregation unit 115 may identify the downloaded file by using the file name and a URL indicating the file storage location, and count the number of downloads for each file. The aggregation unit 115 may also aggregate user behavior on an application screen instead of a web page. For example, event data on the application screen may be used as a primary key, and parameters for each event may be aggregated.
[0043] When multiple events each containing a string as a parameter indicate the same event, the predetermined rule may include concatenating (or merging) the strings contained in the parameters corresponding to the multiple events according to a predetermined criterion, such as concatenating the strings using a comma.
[0044] For example, if there are identical events (e.g., the same user, the same page, the same session, etc.) in the events included in the event data of the dimension set by the user, the aggregation unit 115 concatenates the strings included in each parameter (e.g., values entered by the user, etc.) and integrates the events of the same dimension.
[0045] By performing the above process, when there are multiple events of the same dimension, the parameter strings are concatenated, thereby reducing the amount of data and improving the efficiency of analysis.
[0046] The aggregated data may include individual event aggregated data for each event, or overall aggregated data that compiles the individual event aggregated data. That is, the second generation unit 116 generates individual event aggregated data and / or overall aggregated data as aggregated data. For example, the second generation unit 116 may first generate individual event aggregated data and then generate overall aggregated data as needed, or may generate overall aggregated data as a requirement. The overall aggregated data may have a data format in which, for example, one record (one row) per page view has each event as a column item.
[0047] The above process makes it possible to set the generation of aggregated data for each user, thereby increasing the variety of output to users.
[0048] The first generation unit 114 may also identify individual events based on identification information including a user ID (e.g., user_pseudo_id shown in FIG. 4) and a session ID included in the parameters (e.g., ga_session_id shown in FIG. 4). By using this identification information as a primary key, the first generation unit 114 can extract events with the same user ID but different session IDs as different events. This also makes it possible to analyze user behavior on a session-by-session basis.
[0049] 3 is a block diagram illustrating an example of a user device 20 according to an embodiment of the disclosure. For example, the user device 20 includes one or more processors (CPUs) 210, one or more network communication interfaces 220, a memory 230, and one or more communication buses 270 for interconnecting these components.
[0050] The user interface 250 includes a display 251 and an input device 252 (such as a keyboard and / or a mouse or some other pointing device).
[0051] Memory 230 may be, for example, a high-speed random access memory such as a DRAM, SRAM, or other random access solid-state storage device, or may be a non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices, or may be a non-transitory computer-readable recording medium.
[0052] The memory 230 stores data used by the information processing system 1. For example, the memory 230 stores analytical data for analyzing user behavior on a web page.
[0053] Another example of memory 230 may be one or more storage devices located remotely from processor 210. In some embodiments, memory 230 stores programs, modules, and data structures, or a subset thereof, that are executed by processor 210.
[0054] The processor 210 executes a program stored in the memory 230 to configure a client application 212. The client application 212 includes, for example, a web browser and an email application. The web browser can access web pages provided by the server 10.
[0055] Furthermore, the processor 210 can receive services provided by the server 10 by executing the programs stored in the memory 230. The processor 210 also has an output unit 213, an acquisition unit 214, and a display control unit 215 to execute the processing of the disclosed technology.
[0056] The output unit 213 outputs the analysis data stored in the memory 230 to the server 10 via the network communication interface 220. For example, the output unit 213 uploads the analysis data using a web page provided by the server 10.
[0057] The acquisition unit 214 acquires, from the server 10 via the network communication interface 220, the aggregated data generated by the server 10 based on the output analysis data.
[0058] The display control unit 215 controls the screen displayed on the display 251. For example, it controls the display of the aggregated data on the display screen.
[0059] As described above, the user device 20 displays the aggregated data generated by the server 10 and analyzes the data, thereby enabling the user to efficiently analyze user behavior on the web page.
[0060] <Examples of Data> Next, some of the data used in the disclosed technology will be described with reference to FIGS. 4 to 10. FIG.
[0061] FIG. 4 is a diagram illustrating an example of analysis data according to an embodiment. The example illustrated in FIG. 4 is analysis data in JSON format that can be acquired by GA4. The analysis data illustrated in FIG. 4 includes "event date," "event timestamp," "event_name," "event_params," "event_bundle_sequence_id," "user_pseudo_id," and the like. Furthermore, the analysis data may be directly connected to a cloud-based data storage service using an API or the like, and the data may be analyzed. Furthermore, when uploading and downloading analysis data, a single file containing multiple compressed files may be processed, rather than a single file.
[0062] "Event date" indicates the date and time when the event occurred, "event timestamp" indicates the time when the event occurred, "event_name" indicates the name of the event, "event_params" contains the event parameters, "event_bundle_sequence_id" indicates an ID for grouping together events that occurred at the same time, and "user_pseudo_id" indicates the user ID.
[0063] "event_params" also contains a "key" that indicates the attributes of the event and a parameter "value" that corresponds to that "key." "value" can be one of two types: a string "string_value" or a numeric "int_value."
[0064] As shown in Figure 4, "event_name" contains multiple parameters in "scroll", including, for example, parameter values (numeric values or character strings) corresponding to each key. The data structure is disorganized and complex, making it difficult to analyze.
[0065] 5 is a diagram showing an example of event data generated based on simple analysis data according to an embodiment. In the simple analysis data on the left shown in FIG. 5, "event_name" includes "pagescroll" and "filedownload," and each event has multiple keys (attributes) and parameter values corresponding to each key.
[0066] The first generation unit 114 generates event data for the pagescroll (pagaview) event (table in the upper right) and event data for the filedownload event (table in the lower right) shown on the right in Fig. 5 based on the simple analysis data shown on the left in Fig. 5. For example, the first generation unit 114 dynamically creates and executes an SQL statement using the column names for each "key" stored in the columns of the simple analysis data. The first generation unit 114 references values that contain values from int and string values as data to be inserted into the event data.
[0067] The first generating unit 114 can generate a part of the event table using, for example, the following SQL statements: CREATE TABLE pagescroll((common column), key1, key2) CREATE TABLE filedownload((common column), key1, key2, key3)
[0068] By the above processing, even if an event contains multiple attributes and parameters (numeric values or character strings) for the attributes, the data can be restructured into a format that is easy to analyze, with each attribute as an item in a column.
[0069] 6 is a diagram illustrating an example of an event list according to an embodiment. Fig. 6 includes some of the events prepared on the GA4 side. For example, "click," "file_download," and "scroll" are included in the event list.
[0070] When the analysis data includes a new event (for example, a custom dimension) set by the user, the first generating unit 114 may include this new event in the event list.
[0071] An example of generating aggregated data will be described using Figures 7 to 9. Figure 7 is a diagram showing an example of each event data according to an embodiment. The content of the "value" of the simple event data shown in Figure 7 can have various patterns, such as only numeric values or a mixture of alphanumeric characters, depending on the event.
[0072] 7A shows an example of simplified event data for page scroll. Two events are extracted for page A, which has the same PV (pageview) as "page A." Since the predetermined rule for the same scroll event is to obtain the maximum value, the aggregation unit 115 obtains 90% of the 50% and 90% values and aggregates the two scroll events for page A as events with a scroll rate of 90% of the maximum value.
[0073] 7B shows an example of simple event data for filedownload. Two events are extracted for pageA, which has the same PV (pageview) as "pageA." Since the predetermined rule for the same file download event is to calculate a total value, the counting unit 115 calculates "2" and "1" downloads, for a total of "3," and counts the two download events for pageA as an event with a total value of "3."
[0074] 7C is an example of simple event data for custom dimension A. Two events are extracted for page A, which has the same PV (pageview) as "page A." In this case, the counting unit 115 combines "ABC" and "nbv" as a predetermined rule for the same event for custom dimension A, and therefore counts the two custom dimension A events for page A as an event having the string "ABC, nbv."
[0075] 8 is a diagram illustrating an example of comprehensive aggregate data according to an embodiment. For example, the comprehensive aggregate data illustrated in FIG. 8 is data that combines the aggregate data of each event. The comprehensive aggregate data includes aggregate values for page scrolls, file downloads, and custom dimension A based on page views.
[0076] For example, "pageA" includes a page scroll of "0.9," file downloads of "3 times," and custom dimension A of "ABC, nbv." "pageB" includes a page scroll of "1." Note that the parameter values included in the analysis data are assumed to be numeric only; in this case, the numerical unit is specified as "%" based on the event data "scroll" (column name), and "0.9" and "1" are considered 90% and 100%, respectively, to generate the event aggregation data.
[0077] 9A and 9B are diagrams illustrating an example of event aggregation data according to an embodiment. Fig. 9A shows aggregation data for scroll events, in which the event data shown in Fig. 7A is aggregated for each page view. For example, for "pageA," the scroll parameter value is aggregated to the maximum scroll rate of "90%."
[0078] 9B shows the aggregated data of file download events, where the event data shown in FIG. 7B is aggregated for each page view. For example, for "pageA," the parameter value for the number of downloads is aggregated to a total value of "3 (=2 + 1)."
[0079] Figure 9(C) shows aggregated data for custom dimension events, where the event data shown in Figure 7(C) is aggregated by page view. For example, for "pageA," the character strings included in the parameter values of each event are concatenated with commas to form "ABC, nbv."
[0080] 8 or 9, the output unit 117 may output a file designated by the user. In this way, the information processing system 1 may determine the billing amount based on the aggregated data output to the user device 20.
[0081] <Explanation of Operation> Next, a description will be given of each operation of the information processing system 1. Fig. 10 is a flowchart showing an example of processing related to the server 10 according to an embodiment. The processing shown in Fig. 10 indicates, for example, processing that is performed after the user device 20 logs in to a service provided by the server 10.
[0082] In step S102, the acquisition unit 113 of the server 10 acquires analysis data having items including events related to user behavior on the web page and parameters related to each event.
[0083] In step S104, the first generation unit 114 of the server 10 determines whether or not there is a new event not included in the event list among the events (e.g., event_name) included in the acquired analysis data. If there is a new event (step S104—YES), the process proceeds to step S106. If there is no new event (step S104—NO), the process proceeds to step S108.
[0084] In step S106, the first generation unit 114 of the server 10 registers the new event in the event list and updates the event list.
[0085] In step S108, the first generation unit 114 of the server 10 generates an event table for each event identified based on the event list (see, for example, FIGS. 5 and 7). The event table reorganizes each key and its parameter value included in the "event_params" of the analysis data into a single row of records to facilitate analysis. For example, the first generation unit 114 uses the above-mentioned user ID+session ID as a primary key, enabling more detailed analysis of events with different primary keys even for the same page view.
[0086] In step S110, the tallying unit 115 of the server 10 tally parameters of the same event (for example, the same page view) for each event table based on a predetermined rule associated with the corresponding event.
[0087] In step S112, the second generating unit 116 of the server 10 generates aggregated data based on the aggregated event data (for example, FIG. 8 or FIG. 9).
[0088] In step S114 , the output unit 117 of the server 10 outputs the aggregated data generated by the second generation unit 116 to the user device 20 .
[0089] Through the above processing, data that is easy for analysts to use is generated based on the analysis data of user behavior via web pages, thereby improving the work efficiency of analysts.
[0090] Although the embodiments have been described in detail above, the present invention is not limited to the above embodiments, and various modifications and variations other than the above embodiments are possible within the scope of the claims, as described below.
[0091] [Modifications] For example, the disclosed technology may appropriately integrate the processes on the server side and the user device side, or transfer the processes to the other device, without departing from the spirit of the technology. For example, when analytical data is input into an AI (artificial intelligence) machine learning model, unorganized data cannot be used for proper learning. Therefore, the second generation unit 116 may generate the aggregated data described in the above embodiment as training data for training the machine learning model. This reconstructs the analytical data in a form that is easy to analyze, thereby improving the learning efficiency of the learning model. Furthermore, the server 10 may generate training data for each analytical data holder, or may generate training data by aggregating analytical data for each industry of each analytical data holder, or may generate training data by aggregating all analytical data.
[0092] Specifically, when the server 10 receives a learning data acquisition request from a user device, the server 10 sets the user's past aggregated data as learning data based on the user ID included in the acquisition request, sets the aggregated data linked to the industry ID as learning data based on the industry ID, and generates all aggregated data as learning data when all aggregated data is indicated. The server 10 may determine the amount to be charged to the user depending on the amount of aggregated data included in the learning data.
[0093] The second generation unit 116 may input the aggregated data generated by the above process into a large-scale language model (LLM) or a machine learning model along with instructions to perform further aggregation or analysis. For example, the large-scale language model or the machine learning model may be pre-loaded with data of a schema definition document that defines the structure of the analysis data in a database. This allows the large-scale language model or the machine learning model to understand the structure of the aggregated data, thereby enabling analysis and aggregation based on the schema definition document.
[0094] Furthermore, definitions of conversion and / or churn rate are read into the large-scale language model or machine learning model. Since there are various definitions of conversion and churn rate, in order to use a uniform definition, data containing general definitions of conversion and / or churn rate is read into the large-scale language model or machine learning model.
[0095] The second generation unit 116 may also have a function of inputting prompts to the large-scale language model or machine learning model to identify pages or paths that contribute to sales and analyzing the impact of those pages on purchases. The second generation unit 116 may also have a function of automatically generating graphs and tables based on the analysis results of the aggregated data, and may be provided with format and design rules for visually presenting information.
[0096] The second generation unit 116 may also have a function to integrate multiple JSON files and delete unnecessary data strings to generate a lightweight file. The second generation unit 116 may also have a function to extract only necessary data from a large-scale dataset and perform efficient analysis.
[0097] For example, the second generation unit 116 may input a prompt including instructions to aggregate the period, number of page views, number of sessions, and number of users, along with the aggregated data, to the large-scale language model, and obtain the aggregation results. Furthermore, after the aggregation, the second generation unit 116 may input the following instructions as an additional prompt to the large-scale language model, and obtain the analysis results. List the content and pages that visitors most frequently view before making sales or conversions. Analyze the impact of these pages on purchases. Identify blogs, product description pages, or FAQ pages that visitors frequently view before making payments or conversions, and analyze how these pages affect sales. Analyze pages with high viewing times and click rates, and analyze which pages visitors pass through before converting or making payments. The above are merely examples, and the present invention is not limited to these examples.
[0098] 1...information processing system, 10...server, 20...user device, 110...processor, 130...memory, 112...control unit, 113...acquisition unit, 114...first generation unit, 115...tallying unit, 116...second generation unit, 117...output unit, 210...processor, 212...client application, 213...output unit, 214...acquisition unit, 215...display control unit
Claims
1. Information processing device, To obtain analytical data provided by a service that analyzes user behavior on a web page or application, which includes items containing each event related to user behavior on a web page or application, and each parameter related to each event. Based on the aforementioned analysis data, event data including event attributes and parameters is generated for each event. For the aforementioned event data, parameters are aggregated based on predetermined rules associated with the corresponding event. To generate aggregated data based on the event data including aggregated parameters, Outputting the aforementioned aggregated data, An information processing method that performs the following.
2. Generating the aforementioned event data means The information processing method according to claim 1, comprising classifying each event according to the registered event list.
3. The aforementioned information processing device To generate the event list based on pre-configured events, To determine whether each event included in the aforementioned analysis data is included in the aforementioned event list, The information processing method according to claim 2, further comprising: if it is determined that there is a predetermined event not included in the event list, registering the predetermined event in the event list.
4. The information processing method according to any one of claims 1 to 3, wherein, when multiple events indicate the same scroll event, the predetermined rule includes obtaining the maximum value from among multiple numerical values included in multiple parameters corresponding to the multiple events.
5. The information processing method according to any one of claims 1 to 3, wherein, when multiple events indicate the same download event, the predetermined rule includes obtaining the sum of multiple numerical values included in multiple parameters corresponding to the multiple events.
6. The information processing method according to any one of claims 1 to 3, wherein, when multiple events containing strings as parameters indicate the same event, the predetermined rule includes concatenating the multiple strings contained in the multiple parameters corresponding to the multiple events according to a predetermined criterion.
7. The information processing method according to claim 1, wherein the aggregated data includes event aggregated data for each event, or comprehensive aggregated data which is a compilation of the event aggregated data.
8. Generating the aforementioned event data means The information processing method according to claim 1, comprising identifying individual events based on identification information including a user ID and a session ID included in the parameters.
9. The information processing method according to claim 1, further comprising inputting a prompt to a large-scale language model that has read schema definition data of the analysis data, which includes instructions to aggregate the aggregated data and at least one of the period, number of page views, number of sessions, and number of users, and obtaining the aggregation results.
10. In an information processing device, To obtain analytical data provided by a service that analyzes user behavior on a web page or application, which includes items containing each event related to user behavior on a web page or application, and each parameter related to each event. Based on the aforementioned analysis data, event data including event attributes and parameters is generated for each event. For the aforementioned event data, parameters are aggregated based on predetermined rules associated with the corresponding event. To generate aggregated data based on the event data including aggregated parameters, Outputting the aforementioned aggregated data, A program that executes the command.
11. To obtain analytical data provided by a service that analyzes user behavior on a web page or application, the data having items that include each event related to user behavior on a web page or application, and each parameter related to each event. Based on the aforementioned analysis data, a first generation unit generates event data for each event, including event attributes and parameters. The aforementioned event data includes an aggregation unit that aggregates parameters based on predetermined rules associated with the corresponding event, A second generation unit generates aggregated data based on the event data including aggregated parameters, An output unit that outputs the aforementioned aggregated data, An information processing device equipped with the following features.