Data management device, data management method, and data management program

The data management technology constructs a latent space that reflects user intentions, addressing the challenge of recommending highly related data by allowing users to interactively position and update data relationships, thereby improving relevance.

JP7850986B2Active Publication Date: 2026-04-24UNIV OF TSUKUBA +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
UNIV OF TSUKUBA
Filing Date
2022-07-19
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Conventional data management systems struggle to recommend highly related data based on user preferences, as they primarily manage data using bibliographic information, making it difficult to extract papers that are subjectively relevant.

Method used

A data management technology that constructs a latent space embedding data positions corresponding to relationships between data, allowing users to move and update these positions to reflect their intentions, and recommends highly related data based on user interactions.

Benefits of technology

Enables the construction of a latent space that reflects user intentions, effectively recommending data that is highly related to specific data points, enhancing the relevance of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850986000009
    Figure 0007850986000009
  • Figure 0007850986000010
    Figure 0007850986000010
  • Figure 0007850986000011
    Figure 0007850986000011
Patent Text Reader

Abstract

To construct a latent space that embeds data in positions according to data relationships and reflects a user's intention.SOLUTION: A processor (10) of a data management device (1) performs learning processing (S11) including: embedding processing for embedding each data included in a dataset (DS0) into a latent space (LS) on the basis of a relationship between data (S111); display processing for displaying the latent space (LS) embedded with each data including in the dataset DS on a display (S112); moving processing for moving a position in the latent space (LS) of at least one data including in the dataset (LS) according to a user's operation (S113); and update processing (S114) for updating the relationship between data on the basis of a position in the latent space (LS) of the data moved in the movement processing (S113).SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a data management device, a data management method, and a data management program for managing data. [Background technology]

[0002] The number of academic papers that researchers need to manage is increasing year by year, becoming enormous. These academic papers are often managed using reference management software such as EndNote® or BibTex (Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] ELSEVIER. Mendeley. https: / / www.mendeley.com / . (Accessed on 12 / 30 / 2021). [Overview of the project] [Problems that the invention aims to solve]

[0004] However, conventional literature management software manages papers according to their bibliographic information. Therefore, while it's easy to search for papers by a specific author or papers containing specific terms, it's difficult to extract papers that are highly related to a particular paper. In particular, what constitutes a highly related paper is subjective to the user. Consequently, while the extraction of highly related papers should reflect the user's preferences, conventional literature management software has been unable to do so. It should be noted that this problem is not limited to the management of document data such as academic papers and patent documents, but can occur in the management of any data.

[0005] One aspect of the present invention has been made in view of the above-mentioned problems, and its objective is to realize a data management technology that can construct a latent space in which data is embedded at positions corresponding to the relationships between data, and which reflects the user's intentions. Furthermore, it is also to realize a data management technology that can recommend data highly related to specific data to the user using such a latent space. [Means for solving the problem]

[0006] A data management device according to embodiment 1 of the present invention comprises at least one processor, the processor performing a learning process which includes embedding each data included in a dataset into a latent space based on the relationships between the data; displaying the latent space in which each data included in the dataset is embedded on a display; moving the position of at least one data included in the dataset in the latent space in response to user operation; and updating the relationships between the data based on the position of the data moved in the latent space during the moving process.

[0007] According to the above configuration, a latent space is created in which data is embedded at positions corresponding to the relationships between data, and a latent space that reflects the user's intentions can be constructed.

[0008] In the data management device according to aspect 2 of the present invention, in addition to the configuration of aspect 1, the data is a node of a heterogeneous mixed network, and in the embedding process, the processor embeds each data included in the dataset into the latent space based on the adjacency matrix corresponding to each metapath and the weight corresponding to each metapath, and in the update process, the processor updates the adjacency matrix corresponding to each metapath based on the position in the latent space of the data moved in the move process.

[0009] According to the above configuration, it is possible to construct a latent space in which data is embedded at positions corresponding to the adjacency matrix for each meta-path and the weights for each meta-path, and which reflects the user's intention.

[0010] In the data management device according to Aspect 3 of the present invention, in addition to the configuration of Aspect 2, in the update process, the processor selects a predetermined number of data in order from the ones closer to the position of the data before movement in the latent space for the adjacency matrix corresponding to each meta-path, and if the meta-path is stretched between the moved data and the selected data, sets the elements of the adjacency matrix corresponding to these data to 0, selects a predetermined number of data in order from the ones closer to the position of the data after movement in the latent space, and if the meta-path is stretched between the moved data and the selected data, sets the elements of the adjacency matrix corresponding to these data to 1. Such a configuration is adopted.

[0011] According to the above configuration, it is possible to construct a latent space in which data is embedded at positions corresponding to the adjacency matrix for each meta-path and the weights for each meta-path, and which reflects the user's intention.

[0012] In the data management device according to Aspect 4 of the present invention, in addition to any of the configurations of Aspects 1 to 3, the processor includes an embedding process of embedding each data included in the data set in which new data has been added into the latent space based on the relationship between the data updated by the update process, a selection process of selecting at least one data included in the data set based on the position of the new data in the latent space, and a presentation process of presenting the data selected by the selection process to the user, and executes a recommendation process. Such a configuration is adopted.

[0013] According to the above configuration, when new data is added to the dataset, in the latent space where the data is embedded at a position corresponding to the relationship between the data, the data selected based on the position in the latent space reflecting the user's intention can be presented (for example, recommended) to the user.

[0014] In the data management device according to Aspect 5 of the present invention, in addition to the configuration of Aspect 4, in the selection process, the processor selects a predetermined number of data in order from the ones closer to the new data in the latent space. This configuration is adopted.

[0015] According to the above configuration, when new data is added to the dataset, in the latent space where the data is embedded at a position corresponding to the relationship between the data, in the latent space reflecting the user's intention, the data close to the new data can be presented (for example, recommended) to the user.

[0016] The data management method according to Aspect 6 of the present invention includes: an embedding process in which at least one processor embeds each data included in the dataset into a latent space based on the relationship between the data; a display process in which the latent space in which each data included in the dataset is embedded is displayed on a display; a movement process in which the position of at least one data included in the dataset in the latent space is moved according to a user operation; and an update process in which the relationship between the data is updated based on the position of the data moved in the movement process in the latent space. The learning process is executed.

[0017] According to the above configuration, it is a latent space in which data is embedded at a position corresponding to the relationship between the data, and a latent space reflecting the user's intention can be constructed.

[0018] The data management program according to Aspect 7 of the present invention is a program for operating a computer as the data management device according to any one of Aspects 1 to 5, and is characterized in that the computer is caused to execute each of the above processes.

[0019] According to the above configuration, a latent space can be constructed using a computer, in which data is embedded in positions corresponding to the relationships between data, and which reflects the user's intentions. [Effects of the Invention]

[0020] According to one aspect of the present invention, a latent space in which data is embedded at positions corresponding to the relationships between data can be constructed, and which reflects the user's intentions. [Brief explanation of the drawing]

[0021] [Figure 1] This is a block diagram showing the configuration of a data management device related to one embodiment of the present invention. [Figure 2] This is a flowchart showing the flow of a learning method included in a data management method according to one aspect of the present invention. [Figure 3] This is a flowchart showing the flow of a recommendation method included in a data management method according to one aspect of the present invention. [Figure 4] This figure shows a screen displayed on a display to allow the user to move data in a latent space, relating to one embodiment of the present invention. [Figure 5] This graph, relating to one embodiment of the present invention, shows the dependence of an evaluation index indicating the effectiveness of the learning process on the number of feedback iterations. [Modes for carrying out the invention]

[0022] (Configuration of the data management device) The configuration of the data management device 1 according to one embodiment of the present invention will be described with reference to Figure 1. Figure 1 is a block diagram showing the configuration of the data management device 1.

[0023] The data management device 1 is implemented using a general-purpose computer and, as shown in Figure 1, comprises a processor 11, a primary memory 12, a secondary memory 13, an input / output interface 14, and a bus 15. The processor 11, primary memory 12, secondary memory 13, and input / output interface 14 are interconnected via the bus 15.

[0024] The secondary memory 13 stores the data management program P1 and the dataset DS. The data management program P1 is a program that causes the computer to execute the data management method S1, which will be described later. The dataset DS is the collection of data to be managed.

[0025] The processor 11 loads the data management program P1 stored in the secondary memory 13 onto the primary memory 12. Then, the processor 11 executes each step of the data management method S1 described later, according to the instructions contained in the data management program P1 loaded onto the primary memory 12. During this process, the processor 11 refers to the dataset DS stored in the secondary memory 13.

[0026] Examples of devices that can be used as processor 11 include CPUs (Central Processing Units) and GPUs (Graphics Processing Units). Examples of devices that can be used as primary memory 12 include semiconductor RAM (Random Access Memory). Examples of devices that can be used as secondary memory 13 include HDDs (Hard Disk Drives).

[0027] Input and / or output devices are connected to the input / output interface 14. An example of an output device connected to the input / output interface 14 is a display. This display is used, for example, to display the latent space LS, which will be described later. An example of an input device connected to the input / output interface 14 is a mouse. This mouse is used, for example, to move the position of data embedded in the latent space LS, which will be described later.

[0028] Examples of interfaces that can be used as input / output interface 14 include PCI (Peripheral Component Interconnect) interfaces and USB (Universal Serial Bus) interfaces.

[0029] The data management program P1 may be recorded on a computer-readable recording medium, such as a tangible, non-temporary recording medium. Examples of such recording media include, in addition to the secondary memory 13 mentioned above, tapes, disks, cards, semiconductor memory, and programmable logic circuits.

[0030] (Heterogeneous mixed network) The data constituting the dataset DS is assumed to be given as nodes in a heterogeneous network. Here, a heterogeneous network refers to a network represented by a graph in which the sum of the number of node types and the number of edge types is greater than 2.

[0031] For example, if the data constituting the dataset DS consists of document data representing academic papers, we can consider a heterogeneous network with nodes such as author, term, and publication year in addition to the document data. For example, if the author of both document data u and v is A, a path is established between document data u and document data v via author A in the heterogeneous network. Also, if the term T is included in both document data u and v, a path is established between document data u and document data v via term T in the heterogeneous network. Furthermore, if the publication year of both document data u and v is Y, a path is established between document data u and document data v via publication year Y in the heterogeneous network. In addition, presentation conferences and other information may be added as nodes.

[0032] Alternatively, if the data constituting the dataset DS is document data representing patent documents, a heterogeneous network can be considered with nodes such as applicants (companies), terms, and filing years, in addition to the document data. For example, if the applicant for both document data u and v is A, a path is established between document data u and document data v via applicant A in the heterogeneous network. Also, if both document data u and v contain term T, a path is established between document data u and document data v via term T in the heterogeneous network. Furthermore, if the filing years for both document data u and v are Y, a path is established between document data u and document data v via filing year Y in the heterogeneous network. In addition, patent classifications (e.g., IPC or F-terms) and inventors may be added as nodes.

[0033] Furthermore, the dataset DS can also be constructed using data representing books. In this case, in addition to nodes representing books, we can consider a heterogeneous network that includes, for example, nodes representing authors, nodes representing publishers, nodes representing book classification codes (e.g., C codes), and nodes representing terms. In this case, if two books have the same author, a path is established between the nodes representing those books via the node representing the author. Similarly, if two books have the same publisher, a path is established between the nodes representing those books via the node representing the publisher. Also, if two books have the same book classification code, a path is established between the nodes representing those books via the node representing the book classification code. Furthermore, if the same term is used in two books (full text, summary, or title), a path is established between the nodes representing those books via that term.

[0034] Furthermore, the dataset DS can also be constructed using data representing users in a Social Networking Service (SNS). In this case, in addition to nodes representing users, we can consider a heterogeneous network that includes, for example, nodes representing posted articles, nodes representing terms, nodes representing tags, and nodes representing URLs (e.g., linked URLs). In this case, if two users post the same article, a path is established between the nodes representing those users via the node representing that article. Also, if two users post articles containing the same term, a path is established between the nodes representing those users via the node representing that term. Also, if two users post articles containing the same tag, a path is established between those users via the node representing that tag. Also, if two users post articles containing the same URL, a path is established between those users via the node representing that URL.

[0035] Furthermore, the dataset DS can also be constructed using data representing advertisements. In this case, in addition to nodes representing advertisements, we can consider a heterogeneous network that includes, for example, nodes representing advertisers (e.g., companies), nodes representing viewers, and nodes representing content. In this case, if two advertisements are advertised by the same advertiser, a path is established between the nodes representing those advertisements via the node representing the advertiser. Also, if two advertisements are viewed by the same viewer, a path is established between the nodes representing those advertisements via the node representing the viewer. Furthermore, if two advertisements contain the same content, a path is established between the nodes representing those advertisements via the node representing that content.

[0036] Furthermore, the dataset DS can also be constructed using data representing videos on video streaming sites. In this case, in addition to nodes representing videos, we can consider a heterogeneous network that includes, for example, nodes representing creators, viewers, titles, and tags. In this case, if two videos have the same creator, a path is established between the nodes representing those videos via the node representing the creator. Also, if two videos are viewed by the same viewer, a path is established between the nodes representing those videos via the node representing the viewer. Furthermore, if two videos have the same title, a path is established between the nodes representing those videos via the node representing the title. Also, if two videos have the same tag, a path is established between the nodes representing those videos via the node representing the tag.

[0037] Furthermore, the dataset DS can also be constructed using data representing products on an e-commerce (EC) site. In this case, in addition to nodes representing products, we can consider a heterogeneous network that includes, for example, nodes representing product classification codes, nodes representing terms, and nodes representing buyers. In this case, if two products have the same product classification code, a path is established between the nodes representing those products via the node representing that product classification code. Also, if the product descriptions of two products contain the same term, a path is established between the nodes representing those products via the node representing that term. Furthermore, if two products were purchased by the same buyer, a path is established between the nodes representing those products via the node representing that buyer.

[0038] Furthermore, the dataset DS can also be constructed using data representing articles in news. In this case, in addition to nodes representing articles, we can consider a heterogeneous network that includes nodes representing authors (reporters, contributors, authors, etc.) and nodes representing terms. In this case, if two articles have the same author, a path is established between the nodes representing those articles via the node representing the author. Also, if two articles contain the same term, a bus is established between the nodes representing those articles via the node representing that term.

[0039] Furthermore, the dataset DS can also be composed of data representing websites on the WWW (World Wide Web). In this case, in addition to nodes representing websites, we can consider a heterogeneous network that includes, for example, nodes representing terms and nodes representing domain names. In this case, if the content of two websites contains the same term, a path is established between the nodes representing those websites via the node representing that term. Also, if the URLs of two websites contain the same domain name, a path is established between the nodes representing those websites via the node representing that domain name.

[0040] Furthermore, the dataset DS can also be constructed using data representing applications. In this case, in addition to nodes representing applications, we can consider a heterogeneous network that includes, for example, nodes representing platforms, nodes representing genres, and nodes representing billing methods. In this case, if two applications share the same platform, a path is established between the nodes representing those applications via the node representing that platform. Similarly, if two applications share the same genre, a node representing that genre is established between the nodes representing those applications. And if two applications share the same billing method, a path is established between the nodes representing those applications via the node representing that platform.

[0041] In a heterogeneous network viewed at the schema level, the type of path is called a metapath. In the heterogeneous network of academic papers mentioned above, possible metapaths between document data include, for example, document data-author-document data and document data-term-document data. It is also possible to consider metapaths involving two or more nodes, such as document data-term-document data-term-text data. Similarly, in the heterogeneous network of patent documents mentioned above, possible metapaths between document data include document data-applicant-document data and document data-filing year-document data. It is also possible to consider metapaths involving two or more nodes, such as document data-applicant-document data-term-document data.

[0042] If the data constituting the dataset DS is given as nodes in a heterogeneous mixed network, the relationships between the data are, for example, represented by the adjacency matrix A corresponding to each metapath p∈P between the data. p , and the weights w corresponding to each metapath p∈P between the data. p It is expressed by [this method].

[0043] Furthermore, the dataset DS managed by the data management device 1 is not limited to sets of documents (or data representing documents) such as academic papers and patent documents. Any set of data that can become a node in a heterogeneous mixed network can be managed by the data management device 1. Sets of data representing books, sets of data representing users on social networking services, sets of data representing videos on video streaming sites, sets of data representing products on e-commerce sites, sets of data representing articles in news articles, sets of data representing websites on the World Wide Web, and sets of data representing applications can also be managed by the data management device 1.

[0044] (Data management process flow) A data management method S1 according to one embodiment of the present invention will be described with reference to Figures 2 and 3. Figure 2 is a flowchart showing the flow of the learning process S11 included in the data management method S1. Figure 3 is a flowchart showing the flow of the recommendation process S12 included in the data management method S1.

[0045] The learning process S11 included in the data management method S1 includes, as shown in Figure 2, an embedding process S111, a display process S112, a move process S113, and an update process S114. In the learning process S11, the processor 10 of the data management device 1 first executes the embedding process S111, then executes the display process S112 and the move process S113 in parallel, and then executes the update process S114.

[0046] The embedding process S111 is a process that embeds each data point included in the dataset DS into the latent space LS based on the relationships between the data points. In this embodiment, the latent space LS is an L-dimensional Euclidean space R L This is used. In addition, the adjacency matrix A corresponding to each metapath p∈P represents the relationship between the data. p , and the weight w corresponding to each metapath p∈P p Let's consider this. That is, in this embodiment, in the embedding process S111, each data included in the dataset DS is used to create the adjacency matrix A corresponding to each metapath p∈P.p and the weights w corresponding to each meta-path p ∈ P p are embedded into the latent space based on. A specific example of the embedding method used in the embedding process S111 will be described later.

[0047] The display process S112 is a process of displaying the latent space LS in which each data included in the dataset DS is embedded on the display. When the dimension L of the latent space LS is 2 or more, for example, a projection image or a projective image of the latent space LS into a two-dimensional space is displayed on the display.

[0048] The movement process S113 is a process of moving the position of at least one data in the latent space LS included in the dataset DS according to a user operation. In the present embodiment, as the user operation, a drag-and-drop operation of the data in the latent space LS displayed on the display is used. In the movement process S113, the user moves the data in the latent space LS so that data with high relevance approaches each other and data with low relevance moves away from each other.

[0049] The update process S114 is a process of updating the relationship between the data referred to in the embedding process S111 based on the position of the data moved by the movement process S113 in the latent space LS. In the present embodiment, the adjacency matrix A p corresponding to each meta-path p ∈ P is updated based on the position of the data moved by the movement process S113 in the latent space LS.

[0050] Specific examples of the update method used in the update process S114 are as follows. That is, (1) k data (k is an arbitrary natural number) are selected in order from the data closer to the position of the data before movement in the latent space LS. (2) When a meta-path p is spanned between the moved data z and the selected data z, that is, when A p zx = 1, A p zxChange the value of from 1 to 0. (3) In the latent space LS, select k data points in order from those closest to the position of the moved data. (4) If a metapath p was established between the moved data z and the selected data y, that is, A p zy If = 0, A p zy The value of is changed from 0 to 1. By performing the above process for each data moved by the user, the adjacency matrix A corresponding to the metapath p is obtained. p The update is complete.

[0051] In this case, the updated adjacency matrix A p This is given by the following equation (1). Here, Z is the set of data moved by the user. Also, Ne(Z) is the set of k data selected in order from closest to furthest to furthest from furthest from furthest furthest from furthest furthest from furthest furthest from furthest furthest from furthest furthest from furthest furthest from fur

number

[0052] According to the learning process S11 described above, the relationships between data referenced when embedding data into the latent space LS are updated by user operations such as moving data in the latent space LS. Therefore, if a user moves data in the latent space LS so that highly related data moves closer to each other and less related data moves further away from each other, the user can reflect the relationships between data referenced when embedding data into the latent space LS in accordance with the user's expectations.

[0053] The recommendation process S12 included in the data management method S1, as shown in Figure 3, includes an embedding process S121, a selection process S122, and a presentation process S123, and is executed when new data is added to the dataset DS. In the recommendation process S12, the processor 10 of the data management device 1 first executes the embedding process S121, then the selection process S122, and then the presentation process S123. Hereinafter, the dataset DS∪{x} to which new data x has been added will be referred to as dataset DS'.

[0054] The embedding process S121 is a process in which each data item included in the dataset DS' to which the new data has been added is embedded into the latent space LS based on the relationships between the data updated by the update process S114. The embedding method used in embedding process S121 is the same as the embedding method used in embedding process S111.

[0055] The selection process S122 is the process of selecting at least one data item included in the dataset DS based on the position of the new data item in the latent space LS. In this embodiment, in the selection process S122, m data items (where m is any natural number) are selected in order from those closest to the position of the new data item in the latent space LS.

[0056] Presentation process S123 is the process of presenting the data selected by selection process S122 to the user. In this embodiment, presentation process S123 presents the m data selected by selection process S122 to the user. The presentation method used in presentation process S123 is not particularly limited. For example, one possible method is to display a list of titles of the m data selected by selection process S122 on the display.

[0057] According to the recommendation process S12 described above, when new data is added to the dataset DS, data that is in the vicinity of the new data in the latent space LS can be presented to the user. As mentioned above, the relationships between data referenced when embedding data in the latent space LS reflect the relationships assumed by the user. Considering this point, according to the recommendation process S12 described above, when new data is added to the dataset DS, data that the user is likely to judge to be highly related to the new data can be presented to the user.

[0058] When the data management method S1 according to this embodiment is applied to a dataset DS consisting of academic papers, it is possible to recommend academic papers that are highly relevant to newly added academic papers to the dataset DS. In this case, if the elements constituting the heterogeneous mixed network include "terms," ​​academic papers that are highly relevant in content can be recommended, and if the elements constituting the heterogeneous mixed network include "publication year," academic papers that are highly relevant in time can be recommended. In particular, in the data management method S1 according to this embodiment, the latent space is optimized by the learning process S11 so that a higher weight is given to the relationships that the user considers important among the various relationships expressed as a heterogeneous mixed network. This makes it possible to recommend academic papers that meet the user's needs.

[0059] Furthermore, when the data management method S1 according to this embodiment is applied to a dataset DS consisting of patent documents, it is possible to recommend patent documents that are highly related to new patent documents added to the dataset DS. In this case, if the elements constituting the heterogeneous mixed network include "terms," ​​patent documents that are highly relevant in content can be recommended, and if the elements constituting the heterogeneous mixed network include "applicant," patent documents of the same applicant or patent documents of co-applicants can be recommended. In particular, in the data management method S1 according to this embodiment, the latent space is optimized by the learning process S11 so that a higher weight is given to the relationships that the user considers important among the various relationships expressed as a heterogeneous mixed network. This makes it possible to recommend patent documents that are suitable for the user's needs.

[0060] (Specific examples of embedding methods in embedding processes) In the embedded process S111, the processor 11 processes each data contained in the dataset DS into an adjacency matrix A corresponding to each metapath p∈P. p , and the weight w corresponding to each metapath p∈P p Based on this, it is embedded in the latent space LS. The latent space LS is, as mentioned above, an L-dimensional Euclidean space R L The following describes a specific example of the embedding process S111 using a two-layer GCN.

[0061] The dataset DS is denoted as D (the decorative letter D in mathematical formulas), and the data contained in dataset D are denoted as d1, d2, ..., dn. Furthermore, the embedding vectors of data d1, d2, ..., dn (the position vectors of the points on the latent space LS to which they are embedded) are denoted as z. d1 ,z d2 ,…,z dn (In mathematical formulas, z di An arrow is written above it, and the embedded vector z d1 ,z d2 ,…,z dn Let Z be the matrix formed by arranging these elements. Also, X represents the feature quantities of each node. D This is how it is written. In this specific example, the feature quantity X of each node DOnly the data number is provided. If each data di has a title (text data), the Bag-of-Words of that title is used as the feature X for each node. D It may be given as such.

[0062] Embedded vector z d1 ,z d2 ,…,z dn The matrix Z formed by arranging these elements is the output GCN, as shown in equation (2) below. φ (X, A~) is given. GCN output GCN φ (X,A~) is given by equation (3) below.

[0063]

number

number

[0064] Here, φ={W0,W1} is the parameter set. W0 represents the weights of the first layer of GCN, and W1 represents the weights of the second layer of GNC. ReLU represents the ReLU (Reflected Linear Unit) function. A~ (with a ~ above A in the formula) is a matrix (weighted adjacency matrix) defined by equation (4) below, and A· (with a· above A in the formula) is a matrix defined by equation (5) below.

[0065]

number

number

[0066] Furthermore, the weight wp in equation (4) above is defined by equation (7) below, using the contribution rate γp defined by equation (6) below, for example. Also, D in equation (5) above represents the order matrix. Also, p in equation (6) below θThis represents the decoder defined by equation (8) below. In equation (8) below, σ(·) is the sigmoid function and θ={a,b} is the parameter set.

[0067]

number

number

number

[0068] (Examples) The effectiveness of the learning process S11 was verified using a dataset of 50 research papers held by a researcher. For the verification, metapaths considered included paper-author-paper, paper-term-term, paper-publication year-paper, and paper-conference-publication society-paper. In addition to these metapaths, an adjacency matrix was used that incorporated relationships that could not be extracted using these metapaths.

[0069] First, embedding process S111 was performed on these 50 paper data to construct the latent space LS0. Next, display process S112 and movement process S113 were performed on 10 of these 50 paper data (test data) to record their positions in the latent space LS0 after movement. Figure 4 shows the screen displayed on the display at this time. On this screen, when the mouse cursor is moved over a point corresponding to a paper data, the bibliographic information of that paper data is displayed. The user, while checking this bibliographic information, moved the positions of the paper data in the latent space LS0 so that paper data with strong relationships move closer to each other, and paper data with weak relationships move further away from each other.

[0070] Next, for each of the remaining 40 research paper data (training data), the display process S112, the move process S113, and the update process 114 were performed in order. At this time, the effect of the training process S11 was evaluated for each of the latent spaces LS10 obtained by moving the 1st to 10th research paper data, LS20 obtained by moving the 11th to 20th research paper data, LS30 obtained by moving the 21st to 30th research paper data, and LS40 obtained by moving the 31st to 40th research paper data. Specifically, an evaluation index (recall@k) was calculated based on the position of the test data in each of the latent spaces LS10, LS20, LS30, and LS40, and the position of the test data after the move in the initial latent space LS0. The closer these two positions are, the higher the value of this evaluation index, and the further apart they are, the lower the value.

[0071] Figure 5 is a graph of the scores obtained in this way. From this graph, it can be seen that the value of the evaluation metric increases as the number of moving paper data increases, that is, as the number of user feedbacks increases. This confirms that the learning process S11 can create a latent space that reflects the relationships between papers as assumed by the user.

[0072] (Additional notes) The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in the embodiments described above are also included in the technical scope of the present invention. [Explanation of Symbols]

[0073] 1. Data Management Device 11 processors 12 Primary Memory 13 Secondary memory 14 Input / Output Interfaces 15 bus S1 Data Management Method S11 Learning Process S111 Recessed Processing S112 Display Processing S113 Movement Processing S114 Update process S12 Recommendation Processing S121 Recessed Processing S122 Selection Process S123 Presentation process

Claims

1. Equipped with at least one processor, The aforementioned processor, The embedding process involves embedding each data point in the dataset into a latent space based on the relationships between the data points, A display process that displays the latent space in which each data included in the dataset is embedded on a display, A movement process that moves the position of at least one data in the latent space included in the dataset in response to user operation, The learning process includes an update process that updates the relationships between the aforementioned data based on the position of the data moved in the move process in the latent space. A data management device characterized by the following features.

2. The aforementioned data represents nodes in a heterogeneous mixed network. In the aforementioned embedding process, the processor embeds each data contained in the dataset into the latent space based on the adjacency matrix corresponding to each metapath and the weights corresponding to each metapath. In the update process, the processor updates the adjacency matrix corresponding to each metapath based on the position in the latent space of the data moved in the move process. The data management device according to feature 1.

3. In the update process, the processor performs the following steps in the adjacency matrix corresponding to each metapath: (1) select a predetermined number of data in the latent space in order of proximity to the data before movement; (2) if a metapath exists between the moved data and the selected data, set the elements of the adjacency matrix corresponding to these data to 0; (3) select a predetermined number of data in the latent space in order of proximity to the data after movement; and (4) if a metapath exists between the moved data and the selected data, set the elements of the adjacency matrix corresponding to these data to 1. The data management device according to feature 2.

4. The aforementioned processor, An embedding process in which each data point included in the dataset to which new data has been added is embedded into the latent space based on the relationships between the data points updated by the update process, A selection process that selects at least one data item included in the dataset based on the position of the new data in the latent space, The system performs a recommendation process which includes a presentation process that presents the data selected by the selection process to the user. The data management device according to feature 1.

5. In the selection process, the processor selects a predetermined number of data in the latent space in order from the one closest to the new data. The data management device according to feature 4.

6. At least one processor, The embedding process involves embedding each data point in the dataset into a latent space based on the relationships between the data points, A display process that displays the latent space in which each data included in the dataset is embedded on a display, A movement process that moves the position of at least one data in the latent space included in the dataset in response to user operation, The learning process includes an update process that updates the relationships between the aforementioned data based on the position of the data moved in the move process in the latent space. A data management method characterized by the following features.

7. A data management program for operating a computer as a data management device according to any one of claims 1 to 5, wherein the computer is made to perform each of the above processes. A data management program characterized by the following features.

Citation Information

Patent Citations

  • Information processing method, device, and program

    JP2014191757A

  • Model generation method and device for heterogeneous graph node expression

    JP2022002079A