Information processing device, information processing method, and information processing program

The information processing device efficiently generates tokens from trajectory data by encoding position data into character strings, classifying them, and tokenizing them, addressing the challenge of handling large volumes of user location data and enabling effective machine learning model training.

JP2025085911AActive Publication Date: 2025-06-06RAKUTEN GROUP INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023199618
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-06-06
Estimated Expiration
2043-11-27

AI Technical Summary

Technical Problem

Existing machine learning techniques struggle to efficiently generate tokens as learning data from trajectory data, which includes location information associated with user geographical movement, due to the enormous amount of data collected from a large number of users.

Method used

An information processing device and method that encodes consecutive position data into character strings assigned to geographical areas, classifies these strings into sets, and tokenizes them to efficiently generate tokens representing trajectory data.

Benefits of technology

This approach allows for the efficient conversion of trajectory data into tokens, reducing noise and data volume while extracting trajectory characteristics, thereby enabling effective training of machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085911000001_ABST
    Figure 2025085911000001_ABST
Patent Text Reader

Abstract

To efficiently generate a token as learning data from trajectory data.SOLUTION: An information processing device generates multiple character strings by encoding each of multiple pieces of consecutive positional data into multiple character strings. Each of the multiple character strings is a character string allocated to an area including a location identified by each of the multiple pieces of location data. The information processing device sorts the multiple character strings into multiple sets of character strings and divides each of the multiple sets of character strings into multiple tokens.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a machine learning technique using trajectory data. [Background technology]

[0002] In the technology of Natural Language Processing (NLP), a machine learning model is known that performs pre-training using a large amount of data and then re-trains (fine-tuning) using supervised data so that it can be adapted to downstream tasks (e.g., Patent Document 1). A machine learning model that undergoes such a two-stage learning process is also called a foundation model.

[0003] In such a base model, pre-learning is performed using general-purpose data that is not biased toward a specific task. Furthermore, pre-learning is generally performed using a large amount of unsupervised data. One example of a pre-learning technique is MLM (Masked Language Modeling). MLM is a learning technique that masks some of the multiple tokens (phrases) included in an input sentence (input data) and predicts the masked tokens. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2023-097204 A Summary of the Invention [Problem to be solved by the invention]

[0005] By pre-learning and fine-tuning corresponding to the desired downstream task, a learning model applicable to the downstream task can be realized, and thus far, learning models applicable to all language tasks have been developed. On the other hand, although there has been growing interest in learning models that use trajectory data including location information associated with geographical movement of a user, concrete realization has been awaited. On the other hand, location information acquired by a user device held by a user can be used as user trajectory data, but the amount of data collected from a large number of users is enormous, and it is important to train a learning model using the trajectory data efficiently. To achieve this, it is necessary to efficiently generate tokens as learning data from the trajectory data.

[0006] The present invention has been made in consideration of the above-mentioned problems, and has an object to provide a technique for efficiently generating tokens as learning data from trajectory data. [Means for solving the problem]

[0007] In order to solve the above problem, one aspect of an information processing device according to the present invention includes an encoding unit that generates multiple character strings by encoding each of multiple consecutive position data into a character string, wherein each of the multiple character strings is a character string assigned to an area that includes a position identified by each of the multiple position data, a classification unit that classifies the multiple character strings into multiple sets of character strings, and a tokenization unit that divides each of the multiple sets of character strings into multiple tokens.

[0008] In order to solve the above problem, one aspect of an information processing method according to the present invention includes an encoding step of generating a plurality of character strings by encoding each of a plurality of consecutive position data into a character string, wherein each of the plurality of character strings is a character string assigned to an area including a position identified by each of the plurality of position data; a classification step of classifying the plurality of character strings into a plurality of sets of character strings; and a tokenization step of dividing each of the plurality of sets of character strings into a plurality of tokens.

[0009] In order to solve the above problem, one aspect of the information processing program according to the present invention is an information processing program for causing a computer to execute information processing, the program causing the computer to execute processes including: an encoding process for generating multiple character strings by encoding each of multiple consecutive position data into a character string, wherein each of the multiple character strings is a character string assigned to an area including a position identified by each of the multiple position data; a classification process for classifying the multiple character strings into multiple sets of character strings; and a tokenization process for dividing each of the multiple sets of character strings into multiple tokens. Effect of the Invention

[0010] According to the present invention, it is possible to efficiently generate tokens as learning data from trajectory data. The above-mentioned objects, aspects, and advantages of the present invention, as well as objects, aspects, and advantages of the present invention not described above, will be understood by those skilled in the art from the following detailed description of the invention by referring to the accompanying drawings and the claims. [Brief description of the drawings]

[0011] [Figure 1] FIG. 1 shows an example of the configuration of an information processing system according to an embodiment. [Diagram 2] FIG. 2 illustrates an example of a functional configuration of an information processing device according to an embodiment. [Diagram 3]FIG. 3 shows a flow chart of the tokenization process. [Figure 4A] FIG. 4A shows an example of a map on which multiple regions are located. [Figure 4B] FIG. 4B is a diagram illustrating the transformation of data from trajectory data to tokens. [Figure 4C] FIG. 4C shows an example of tokens generated from hash values ​​assigned to 49 hexagons. [Diagram 5] FIG. 5 shows a flowchart of the learning process. [Figure 6] FIG. 6 illustrates an example of a hardware configuration of an information processing device according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] Hereinafter, with reference to the attached drawings, an embodiment for carrying out the present invention will be described in detail. Among the components disclosed below, those having the same functions are given the same reference numerals, and their description will be omitted. Note that the embodiment disclosed below is an example of a means for realizing the present invention, and should be appropriately modified or changed depending on the configuration of the device to which the present invention is applied and various conditions, and the present invention is not limited to the following embodiment. Furthermore, not all of the combinations of features described in the present embodiment are necessarily essential to the solution of the present invention.

[0013] [Information processing system configuration] FIG. 1 shows a configuration example of an information processing system 1 according to the present embodiment. The information processing system 1 includes an information processing device 10 and a user device 11. The information processing device 10 and the user device 11 are configured to be able to communicate with each other via a network 12. The network 13 may include the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a mobile communication network, and the like. Although FIG. 1 shows one user device 11, the information processing system 1 is configured to have a plurality of user devices, and in the present disclosure, the plurality of user devices may be collectively referred to as the user device 11. The user device 11 is operated by a user 13. In the present disclosure, the terms user device and user may be understood to be synonymous.

[0014] The user device 11 is a mobile terminal that can be carried by the user 13. The user device 11 is, for example, a device such as a smartphone or a tablet, and is configured to be able to communicate with the information processing device 10 via the network 12. The user device 11 includes a positioning unit that can acquire position data (position information) of the user device 11. The positioning unit is, for example, a GPS (Global Positioning System) sensor. The user device 11 acquires multiple pieces of continuous position data along the movement of the user 13 and transmits the data to the information processing device 10. The user device 11 may also attach time information (timestamp) at which the position data was acquired to the position data and transmit the data to the information processing device 10. The user device 11 may acquire the position data at regular intervals or at a predetermined time. The time at which the position data is acquired may be instructed by another device. The user device 11 may also transmit the position data with the time information attached to it to an external device other than the information processing device 10 via the network 12.

[0015] The information processing device 10 acquires a plurality of consecutive position data received from the user device 11 as trajectory data. Then, the information processing device 10 performs processing for training a learning model, which will be described later, using the trajectory data. The information processing device 10 can acquire a large amount of different trajectory data from a large number of user devices including the user device 11. In this case as well, the information processing device 10 performs processing, which will be described later, for each piece of trajectory data.

[0016] [Functional configuration of information processing device] The information processing device 10 according to the present embodiment is configured to acquire trajectory data including a plurality of position data, and generate a plurality of character strings by encoding each of the plurality of position data included in the trajectory data into a character string. Furthermore, the information processing device 10 is configured to classify (clusterize) the plurality of character strings into a plurality of sets (clusters) of character strings, and divide each of the plurality of sets of character strings into a plurality of tokens. Furthermore, the information processing device 10 is configured to perform pre-learning and fine-tuning of a language model as a machine learning model using the plurality of tokens.

[0017] FIG. 2 shows an example of the functional configuration of the information processing device 10 according to the present embodiment. The information processing device 10 has, as an example of its functional configuration, a trajectory data acquisition unit 201, an encoding unit 202, a clustering unit 203, a tokenization unit 204, a learning data generation unit 205, a pre-learning unit 206, a fine tuning unit 207, and a storage unit 210. The storage unit 210 is configured to be able to store a language model (natural language processing model) 211, a first token set 212 which is an unsupervised learning data set, and a second token set 213 which is a supervised learning data set. Note that the entire information processing device 10 may not be provided in one device, but may be provided in a plurality of devices. For example, a part of the information processing device 10 may be provided in an external server device. In this case, the following functions are realized by cooperation between the information processing device 10 and the external server device.

[0018] The trajectory data acquisition unit 201 acquires trajectory data including a plurality of consecutive position data from each of a plurality of user devices including the user device 12. The encoding unit 202 encodes (converts) each of a plurality of position data included in the trajectory data acquired by the trajectory data acquisition unit 101 into a character string. The clustering unit 203 classifies (clusters) the plurality of character strings encoded by the encoding unit 102 into a plurality of sets (clusters) of character strings. The tokenization unit 204 divides each of the plurality of sets of character strings grouped by the clustering unit 203 into a plurality of tokens. That is, the tokenization unit 204 generates a plurality of tokens. The training data generation unit 205 generates a first token set 212, which is an unsupervised training data set, and a second token set 213, which is a supervised training data set, from the plurality of tokens generated by the tokenization unit 204. The training data generation unit 205 stores the first token set 212 and the second token set 213 in the storage unit 210. The pre-training unit 206 pre-trains the language model 211 using the first token set 212. The fine tuning unit 207 performs fine tuning corresponding to a predetermined task on the pre-trained language model 211 using the second token set 213. Hereinafter, the process executed by the information processing device 10 will be described, divided into a tokenization process and a learning process.

[0019] [Tokenization process] First, the tokenization process according to this embodiment will be described. Fig. 3 shows a flowchart of the tokenization process executed by the information processing device 10. The tokenization process is started in a state in which the information processing device 10 can acquire trajectory data including multiple consecutive position data of the user device 11 (i.e., the user 13) from the user device 11 via the network 12. Alternatively, the tokenization process is started in a state in which the information processing device 10 can acquire trajectory data of the user device 11 from an external device.

[0020] In S31, the trajectory data acquisition unit 201 acquires trajectory data including a plurality of consecutive position data. In this embodiment, the position data is assumed to be data consisting of latitude and longitude. Alternatively, the position data may be data indicating a position at any coordinate on a map. In this embodiment, the position data is provided with time information (time stamp) at which the position data was acquired by the user device 11.

[0021] In S32, the encoding unit 202 encodes (converts) each of the multiple position data included in the trajectory data acquired by the trajectory data acquisition unit 101 into a character string. Here, the character string to be encoded may be text information in any format. In this embodiment, the encoding unit 102 uses multiple areas arranged in advance on a map. A character string is assigned to each of the multiple areas based on a geographical position. The encoding unit 102 encodes each of the position data into a character string assigned to the area including the position specified by each of the position data. In other words, the encoding unit 102 converts continuous position data into discrete data (character strings) by mapping each of the position data to one of the multiple areas.

[0022] FIG. 4A shows an example of a map 400 on which a plurality of regions are arranged. The map 400 shows an example of a plurality of regions and character strings assigned to each of the regions. In FIG. 4A, each region has a hexagonal shape as an example of a polygon. In the case of a hexagon, it is possible to further reduce a quantization error (error when converting position data into a character string) that occurs when a user moves. In addition, a hexagon also makes it possible to easily approximate a radius. Note that the shape of the region is not limited to a hexagon as long as it is a region that does not overlap with adjacent regions and is in contact with them. For example, the shape of the region may be another polygon, such as a pentagon.

[0023] In Fig. 4A, as an example of a character string assigned to each region (hexagon), a hash value (hash expression) corresponding to the geographical location of each region is shown. In Fig. 4A, the hash value corresponding to each hexagon can be obtained by inputting the latitude and longitude of a predetermined position (e.g., a center position) in each hexagon as arguments (input values) to a hash function. The encoding unit 102 encodes (converts) the latitude and longitude indicated by the position data into a hash value corresponding to the region in which the latitude and longitude are located.

[0024] FIG. 4B is a diagram showing the transformation of data from trajectory data to a token. In FIG. 4B, trajectory data 410 is trajectory data (position data sequence) including four pieces of position data each consisting of latitude and longitude with time information added. The encoding unit 202 generates a hash value sequence 411 by converting each piece of position data of the trajectory data 410 into a hash value. The hash value in the hash value sequence 411 corresponds to a hash value assigned to an area including a position specified by the position data among the multiple areas of the map 400 shown in FIG. 4A. In the hash value sequence 411, the position data (35.524, 139.759, 08:00) and (35.527, 139.757, 08:01) in the trajectory data 410 are converted to the same hash value (82f5a52bfff, 08:01) while maintaining the time information. This corresponds to the fact that the position at latitude=35.524 and longitude=139.759 and the position at latitude=35.527 and longitude=139.757 are located in the same area (the same hexagon in the example of FIG. 4A) arranged on the map. On the other hand, the position data (35.538, 139.769, 11:00) and (35.559, 139.755, 16:00) in the trajectory data 410 are converted into different hash values. This corresponds to the fact that the position at latitude=35.538 and longitude=139.769 and the position at latitude=35.559 and longitude=139.755 are located in different areas (different hexagons in the example of FIG. 4A) arranged on the map.

[0025] In this embodiment, each character string encoded by the encoding unit 102 is composed of multiple blocks corresponding to multiple hierarchical geographical areas (large division areas to small division areas). Here, one or more common (i.e., the same) characters are used for common geographical areas. For example, character strings corresponding to the same numerical parts of latitude and longitude are expressed by one or more common characters, and as a result, are composed of multiple blocks.

[0026] In the example of FIG. 4A, all 49 hexagons shown in the figure form the same large division (upper) area, and the large division area is applied with the character "82f5a". Therefore, the character "82f5a" is applied to the beginning of all 49 hexagons. In addition, the seven hexagons surrounded by a thick line form the same medium division (middle) area for each of the seven hexagons, and a common character is applied to each of the seven hexagons. That is, the characters "e8", "ee", "e1", "52", "53", "ed", or "ec" are applied following "82f5a". In addition, each of the seven hexagons forms a small division (lower) area, and a different character is applied to each of them. For example, the characters "7fff", "5fff", "dfff", "9fff", "bfff", "3fff", or "1fff" are applied following "e8" to each of the seven hexagons in the medium division area to which "e8" is assigned. Thus, in the example of FIG. 4A, the hash value (character string) assigned to each area is made up of at least three blocks.

[0027] To explain using addresses, assume that the addresses corresponding to a location specified by two latitudes and longitudes are "A City, B Town, 1-chome" and "A City, B Town, 2-chome." In this case, the character strings corresponding to the two addresses are configured so that one or more characters representing "A City" and "B Town" are common, and one or more characters representing "1-chome" and "2-chome" are different. Therefore, the character string in this example is an expression having multiple blocks corresponding to multiple geographical areas, namely "cities," "towns," and "chome."

[0028] After encoding into character strings, in S33, the clustering unit 203 classifies (clusters) the multiple character strings encoded by the encoding unit 102 into multiple sets of character strings. By clustering, redundant (unnecessary) data is deleted. That is, noise is reduced. In this embodiment, the clustering is performed based on spatiotemporal proximity. Specifically, the clustering unit 203 performs clustering based on the similarity of the multiple character strings encoded by the encoding unit 102 (i.e., the proximity of the positions specified by the position data) and the similarity of the time information attached to the character strings (i.e., the proximity of the times when the position data are acquired). For example, the clustering unit 203 clusters, i.e., merges, into one set, one or more character strings that have the same predetermined number of characters and have time information within a predetermined time. In this embodiment, the information on the predetermined number and the predetermined time (i.e., the range of proximity of position and time) may be set in advance in the information processing device 10, or may be set by any program stored in the storage unit (ROM 602 or RAM 603 in FIG. 6). Alternatively, the clustering unit 203 may set or adjust the information on the predetermined number and the predetermined time according to a predetermined instruction (e.g., an instruction by an operator).

[0029] An example of clustering will be described with reference to Fig. 4B. The clustering unit 203 judges the similarity of multiple hash values ​​included in the hash value sequence 411 having time information generated by the encoding unit 202. For example, the clustering unit 203 classifies multiple hash values ​​in which a predetermined number of characters match among multiple character strings constituting a hash value and which have time information within a predetermined time into one set. This corresponds to classifying multiple character strings encoded from multiple position data within a predetermined range acquired within a predetermined time into one set.

[0030] Assume that the clustering unit 203 determines that the time information attached to the first two hash values ​​of (82f5a52bfff,08:00), (82f5a52bfff,08:00), (82f5a525fff,11:00), and (82f5ae15fff,16:00) included in the hash value sequence 411 is within a predetermined time. The two hash values ​​match. In this case, the clustering unit 203 groups the first two hash values ​​into one set. That is, the first two hash values ​​(82f5a52bfff,08:00) and (82f5a52bfff,08:00) are grouped into one set (82f5a52bfff,08:01). Here, the latest time information (=08:01) among the time information attached to the hash value and the position data is attached as the time information, but is not limited to this. For example, the earliest time information or average time information may be adopted. Then, the clustering unit 203 generates a clustered hash value sequence 412. If no clustered hash value exists, the hash value sequences generated by the encoding unit 202 and the clustering unit 203 will be the same.

[0031] By performing clustering, multiple character strings (multiple hash values ​​in this embodiment) encoded from multiple location data are organized into one or more sets (clusters) based on spatiotemporal proximity. This reduces noise in the data and controls the density of trajectory points. The density of the trajectory points can be controlled by adjusting the conditions related to spatiotemporal proximity (range of location and time proximity).

[0032] In S34, the tokenization unit 204 generates multiple tokens by tokenizing (dividing) the sets of character strings clustered by the clustering unit 203. As described above, each character string is composed of multiple hierarchical blocks, and the tokenization unit 204 generates multiple tokens by using the multiple blocks. For example, the tokenization unit 204 divides each of the sets of character strings clustered by the clustering unit 203 into multiple hierarchical blocks to generate multiple tokens. Each block corresponds to one token. In this embodiment, the character strings are hash values, and the tokenization unit 204 divides each hash value into multiple sub-hash values ​​and generates each sub-hash value as a token.

[0033] An example of tokenization will be described with reference to Fig. 4B. The tokenizer 204 divides each of the multiple hash values ​​included in the hash value sequence 412 into multiple tokens (sub-hash values) and generates a token sequence 413 consisting of the multiple tokens. The token sequence 413 includes multiple identical tokens (e.g., "52"), but each token is generated from a different clustered hash value, and therefore is treated as a different token (data).

[0034] In this manner, multiple tokens are generated from the trajectory data. By performing encoding and clustering processes on the trajectory data, multiple tokens can be generated from the trajectory data so as to represent the characteristics of the trajectory (i.e., so that the characteristics of the trajectory are extracted). In addition, multiple tokens can be generated by reducing the amount of trajectory data. In the example of FIG. 4B, a token sequence 413 representing the characteristics of the trajectory is generated from trajectory data 410 having time information. As is clear from the figure, the amount of data of the token sequence 413 is significantly reduced from the amount of data of the trajectory data 410.

[0035] The reduction of data amount by tokenization will be further described with reference to FIG. 4C. FIG. 4C shows an example of tokens generated from hash values ​​assigned to the 49 hexagons shown in FIG. 4A. In FIG. 4C, the hash value group 420 includes 49 hash values ​​assigned to each of the 49 hexagons shown in FIG. 4A. The token group 421 is a collection of tokens generated from the hash value group 420 according to the above-mentioned procedure. As can be seen from FIG. 4C, the data amount of the token group 421 is significantly reduced from the data amount of the hash value group 420. Therefore, for example, when trajectory data is obtained from a large number of users via user devices, a large-scale data set including a large number of tokens representing trajectory characteristics can be generated from the trajectory data.

[0036] In this way, the information processing device 10 converts a plurality of position data into a plurality of character strings based on the positions specified by the position data, and classifies the plurality of character strings into a plurality of clusters based on the proximity (similarity) between the character strings and the time information at which the position data was acquired. Then, the information processing device 10 divides the character strings included in the plurality of clusters into a plurality of tokens. Through such processing, the trajectory data including the plurality of position data is converted into tokens from which the characteristics of the trajectory are extracted after noise is reduced. As a result, when trajectory data is obtained from a large number of users, a large-scale data set including a large number of tokens indicating the characteristics of the trajectory can be generated from the trajectory data. Such a large-scale data set is used to train the language model 211.

[0037] [Learning process] Next, the learning process according to this embodiment will be described. Fig. 5 shows a flowchart of the learning process executed by the information processing device 10. The tokens generated by the tokenization unit 204 are used in the learning process. First, in S51, the learning data generation unit 205 generates a first token set and a second token set from the multiple tokens generated by the tokenization unit 204.

[0038] The training data generation unit 205 generates an unsupervised training data set as the first token set 212 from the multiple tokens generated by the tokenization unit 204, and stores the first token set 212 in the storage unit 210. That is, the first token set 212 is a training data set that does not depend on any task. In this embodiment, in order to pre-train the language model 211 by the MLM (Masked Language Modeling) method, the training data generation unit 205 generates multiple sets of tokens as the first token set 212, in which some of the multiple tokens generated from one trajectory data are masked. The first token set 212 is used for pre-training by the pre-training unit 206.

[0039] Furthermore, the training data generation unit 205 generates a supervised training data set as a second token set 213 from the multiple tokens generated by the tokenization unit 204, and stores the second token set in the storage unit 210. The second token set includes multiple sets in which labels (correct answer data) for a target task are attached to multiple tokens generated from one trajectory data. The second token set 213 is used for fine tuning by the fine tuning unit 206. The target task may be a task for a location-based service related to urban planning or traffic.

[0040] In S52, the pre-learning unit 206 pre-learns the language model 211 using the first token set (unsupervised learning data set). The language model 211 is, for example, a machine learning model incorporating an architecture called a transformer. That is, the language model 211 is a language model based on a transformer. As such a transformer, BERT (Bidirectional Encoder Representation from Transformers) is known. Some of the tokens included in the first token set 211 are masked, and the pre-learning unit 206 trains the language model 211 by estimating the masked tokens. The language model 211 is trained in a task-independent manner, making it possible to obtain a broad understanding of the trajectory data. In this embodiment, since it is possible to pre-learn the language model 211 from a data set of a large amount of tokens generated from trajectory data, the language model 211 can be referred to as an LTM (Large trajectory model) based on an LLM (Large Language Model).

[0041] In S53, the fine tuning unit 206 fine-tunes the pre-trained language model 211 using a second token set (supervised learning data set) generated for a target task. The parameters (weights) of the pre-trained language model 211 are adjusted so that the pre-trained language model 211 can understand a plurality of tokens generated from trajectory data. Fine tuning is learning that gives a clear task to the pre-trained language model 211. Through fine tuning, the parameters of the language model 211 are adjusted to optimize its performance on the target task, making it possible to configure the language model 211 adapted to the task.

[0042] In this manner, in this embodiment, trajectory data composed of user position data obtained by a user device is used as user trajectory data. Such trajectory data can be obtained from many users, is factual information obtained in the real world, and can be useful learning data. In this embodiment, a large amount of trajectory data obtained from many users is converted into a plurality of tokens representing trajectory features after reducing noise and reducing the amount of data by the above-mentioned procedure. Since the amount of data of the plurality of tokens is significantly reduced from the trajectory data, even when a large amount of trajectory data is used, the language model 211 can be efficiently trained using the plurality of tokens.

[0043] In the above embodiment, an example has been described in which the encoded character string is composed of a plurality of blocks corresponding to hierarchical geographical areas and is tokenized using the plurality of blocks, but the nature of the plurality of blocks is not limited to this. For example, the encoded character string may be composed of a plurality of blocks based on a predetermined rule set on a map and may be tokenized using the plurality of blocks.

[0044] [Hardware configuration of information processing device] Next, a description will be given of an example of the hardware configuration of the information processing device 10. Fig. 6 is a block diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. The information processing apparatus 10 according to the present embodiment can be implemented on any single or multiple computers, mobile devices, or any other processing platform. 6, the information processing device 10 is illustrated as being implemented in a single computer, but the information processing device 10 according to the present embodiment may be implemented in a computer system including multiple computers. The multiple computers may be connected to each other via a wired or wireless network so as to be able to communicate with each other.

[0045] 6, the information processing device 10 may include a CPU (Central Processing Unit) 601, a ROM (Read Only Memory) 602, a RAM (Random Access Memory) 603, a HDD (Hard Disk Drive) 604, an input unit 605, a display unit 606, a communication I / F (communication unit) (interface) 607, and a system bus 608. The information processing device 10 may also include an external memory. The CPU 601 generally controls the operation of the information processing device 10, and controls each of the components (602 to 607) via a system bus 608, which is a data transmission path.

[0046] The ROM 602 is a non-volatile memory that stores a control program and the like necessary for the CPU 601 to execute processing. The program includes instructions (codes) for executing the processing according to the above-described embodiment. The program may be stored in a non-volatile memory such as the HDD 604 or an SSD (Solid State Drive) or an external memory such as a removable storage medium (not shown). The RAM 603 is a volatile memory and functions as a main memory, a work area, etc. of the CPU 601. That is, the CPU 601 loads necessary programs, etc. from the ROM 602 to the RAM 603 when executing processing, and realizes various functional operations by executing the programs, etc. The RAM 603 may include the storage unit 210 shown in FIG.

[0047] The HDD 604 stores, for example, various data and various information required when the CPU 601 performs processing using a program. The HDD 604 also stores, for example, various data and various information obtained when the CPU 601 performs processing using a program. The input unit 605 is composed of a keyboard and a pointing device such as a mouse. The display unit 606 is configured with a monitor such as a liquid crystal display (LCD), etc. The display unit 606 may be configured in combination with the input unit 605 to function as a GUI (Graphical User Interface).

[0048] The communication I / F 607 is an interface that controls communication between the information processing device 10 and an external device. The communication I / F 607 provides an interface with a network and executes communication with the external device via the network. Various data, various parameters, and the like are transmitted and received between the information processing device 10 and the external device via the communication I / F 607. In this embodiment, the communication I / F 607 may execute communication via a wired LAN (Local Area Network) or a dedicated line that complies with a communication standard such as Ethernet (registered trademark). However, the network that can be used in this embodiment is not limited to this, and may be configured as a wireless network. This wireless network includes wireless PANs (Personal Area Networks) such as Bluetooth (registered trademark), ZigBee (registered trademark), and UWB (Ultra Wide Band). It also includes wireless LANs (Local Area Networks) such as Wi-Fi (Wireless Fidelity) (registered trademark), and wireless MANs (Metropolitan Area Networks) such as WiMAX (registered trademark). It also includes wireless WANs (Wide Area Networks) such as 4G and 5G. It should be noted that the network is sufficient as long as it can connect the devices to each other so that they can communicate with each other, and the communication standard, scale, and configuration are not limited to those described above.

[0049] At least some of the functions of each element of the information processing device 10 shown in Fig. 2 can be realized by the CPU 601 executing a program. However, at least some of the functions of each element of the information processing device 10 shown in Fig. 2 may be operated as dedicated hardware. In this case, the dedicated hardware operates under the control of the CPU 601.

[0050] The disclosure of this embodiment includes the following configuration. [1] An encoding unit that generates a plurality of character strings by encoding each of a plurality of consecutive position data into a character string, wherein each of the plurality of character strings is a character string assigned to an area including a position specified by each of the plurality of position data; a grouping unit that groups the plurality of character strings into a plurality of sets of character strings; and a tokenization unit that divides each of the plurality of sets of character strings into a plurality of tokens; An information processing device having the above configuration.

[0051] [2] The information processing device according to [1], wherein each of the plurality of position data is composed of latitude and longitude.

[0052] [3] The information processing device according to [1] or [2], wherein the area is one of a plurality of areas pre-arranged on a map.

[0053] [4] The information processing device according to any one of [1] to [3], wherein the encoding unit generates the plurality of character strings by adding time information at which each of the plurality of position data was acquired, and the grouping unit group the plurality of character strings into the plurality of sets of character strings based on similarity between the character strings and similarity between the time information.

[0054] [5] An information processing device according to any one of [1] to [4], wherein each of the multiple character strings generated by the encoding unit is composed of multiple blocks, and the tokenization unit divides each of the multiple sets of character strings into multiple tokens using the multiple blocks.

[0055] [6] The information processing device according to any one of [1] to [5], wherein each of the multiple character strings is a hash representation based on each of the multiple location data.

[0056] [7] An information processing device according to any one of [1] to [6], further comprising: a first token set generation unit that generates a first token set from the plurality of tokens by masking some of the tokens; and a pre-training unit that pre-trains a language model based on a transformer using the first token set.

[0057] [8] The information processing device described in [7], further comprising: a second token set generation unit that generates a second token set from the plurality of tokens by labeling the tokens for a specified task; and a fine tuning unit that fine-tunes the pre-trained language model using the second token set. [Explanation of symbols]

[0058] 1: information processing system, 10: information processing device, 11: user device, 12: network, 13: user, 201: trajectory data acquisition unit, 202: encoding unit, 203: clustering unit, 204: tokenization unit, 205: learning data generation unit, 206: pre-learning unit, 207: fine tuning unit, 210: memory unit, 211: language model, 212: first token set, 213: second token set

Claims

1. an encoding unit that generates a plurality of character strings by encoding each of a plurality of consecutive position data into a character string, each of the plurality of character strings being a character string assigned to an area including a position specified by each of the plurality of position data; a grouping unit that groups the plurality of character strings into a plurality of sets of character strings; a tokenizer for dividing each of the plurality of sets of character strings into a plurality of tokens; An information processing device having the above configuration.

2. The information processing apparatus according to claim 1 , wherein each of the plurality of position data is composed of a latitude and a longitude.

3. The information processing device according to claim 1 , wherein the area is any one of a plurality of areas arranged in advance on a map.

4. the encoding unit generates the plurality of character strings by adding time information at which each of the plurality of position data was acquired; the grouping unit grouping the plurality of character strings into the plurality of sets of character strings based on similarity of character strings and similarity of time information; The information processing device according to claim 1 .

5. Each of the plurality of character strings generated by the encoding unit is composed of a plurality of blocks, The tokenization unit divides the plurality of sets of character strings into a plurality of tokens by using a plurality of blocks of each of the plurality of sets of character strings. The information processing device according to claim 1 .

6. The information processing apparatus according to claim 1 , wherein each of the plurality of character strings is a hash expression based on each of the plurality of position data.

7. a first token set generation unit that generates a first token set by masking some of the tokens from the plurality of tokens; a pre-training unit that pre-trains a transformer-based language model using the first token set; The information processing device according to claim 1 , further comprising:

8. a second token set generation unit that generates a second token set by attaching a label for a predetermined task to the tokens from the plurality of tokens; The information processing apparatus according to claim 7 , further comprising a fine tuning unit that fine-tunes the pre-trained language model using the second set of tokens.

9. An information processing method executed by an information processing device, an encoding step of generating a plurality of character strings by encoding each of a plurality of consecutive position data into a character string, each of the plurality of character strings being a character string assigned to an area including a position specified by each of the plurality of position data; a grouping step of grouping the plurality of character strings into a plurality of sets of character strings; a tokenization step of dividing each of the plurality of sets of character strings into a plurality of tokens; An information processing method comprising:

10. An information processing program for causing a computer to execute information processing, the program comprising: an encoding process for generating a plurality of character strings by encoding each of a plurality of consecutive position data into a character string, each of the plurality of character strings being assigned to an area including a position specified by each of the plurality of position data; a grouping process for grouping the plurality of character strings into a plurality of sets of character strings; and a tokenization process for dividing each of the plurality of sets of character strings into a plurality of tokens. Information processing program.

Citation Information

Patent Citations

  • Information processing device, method for operating information processing device, and operation program of information processing device

    JP2023097204A