Methods, apparatus, devices, and storage media for data querying
The method and device for data querying through one-hot encoding and secure fragment exchange in SMPC technologies address the complexity issues, enabling efficient and secure multi-party data analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LEMON CO LTD
- Filing Date
- 2024-03-01
- Publication Date
- 2026-04-10
AI Technical Summary
The increasing computational and communication complexities in secure multi-party computing (SMPC) technologies hinder efficient data circulation and analysis across different data owners, necessitating a more efficient multi-party data query protocol.
A method and device for data querying that involves generating data fragments based on one-hot encoding and secure fragment exchange between devices, allowing classification and aggregation of data attributes while ensuring security, using structured query language (SQL) for multi-party database queries.
This approach simplifies and enhances the computational and communication efficiency of multi-party data queries, ensuring secure and reliable data analysis across different data owners.
Smart Images

Figure 2026510727000001_ABST
Abstract
Description
Technical Field
[0001] [Cross-reference to Related Applications] This application claims priority based on a Chinese patent application filed on March 1, 2023, with the invention title "Data Query Method, Apparatus, Device, and Storage Medium" and application number 202310189039.1, and the entire content disclosed in the above Chinese patent application is incorporated herein by reference.
[0002] [Technical Field] Exemplary embodiments of the present invention generally relate to the field of computers, and more particularly, to data query methods, apparatuses, devices, and storage media.
Background Art
[0003] Data, as a production factor, is playing an increasingly important role in social life. For data to exert greater value, collaboration in the data flow is a prerequisite. Currently, different data is often held by different owners, and the phenomenon of data silos is spreading. Therefore, it is necessary to complete the common analysis of data while ensuring the security of the data of all stakeholders and giving full play to the value of the data. Secure multi-party computing (SMPC) technology is a common technology for solving the data circulation problem. However, with the increase in the amount of data, the computational complexity and communication complexity of this technology increase significantly.
Summary of the Invention
[0004] In a first aspect of the present invention, a method for data query is provided. The method includes A step of obtaining first data via a first device based on a data query request, wherein the request specifies first data attributes to be classified and aggregated, and at least second and third data attributes used for classification, and the first data includes at least data associated with the second data attributes. The steps include generating a third data fragment of a data query result based on a first data fragment of first data and a second data fragment of second data received from a second device, wherein the second data includes at least data associated with a third data attribute, and the third data fragment includes elements corresponding to the first data attribute that are classified and aggregated based on at least the second and third data attributes.
[0005] In a second aspect of the present invention, a device for data querying is provided. The device is An acquisition module configured to acquire first data based on a data query request via a first device, wherein the request specifies first data attributes to be classified and aggregated, and at least second and third data attributes used for classification, and the first data includes at least data associated with the second data attributes, A first generation module is configured to generate a third data fragment of a data query result based on a first data fragment of first data and a second data fragment of second data received from a second device, wherein the second data includes at least data associated with a third data attribute, and the third data fragment includes elements corresponding to a first data attribute that is classified and aggregated based at least on the second and third data attributes.
[0006] In a third aspect of the present invention, an electronic device is provided. The electronic device includes at least one processing unit and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions to be executed by the at least one processing unit. When an instruction is executed by the at least one processing unit, the electronic device is made to execute the method of the first aspect.
[0007] In a fourth aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, which is executed by a processor to realize the method of the first aspect.
[0008] In a fifth aspect of the present invention, a computer program product is provided. The computer program product includes computer executable instructions, and the method of the first aspect is realized when the computer executable instructions are executed by a processor.
[0009] It should be understood that the contents described in the summary of the present invention are not intended to limit the main or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will be readily apparent from the following description. [Brief explanation of the drawing]
[0010] The above-mentioned features and other features, advantages, and aspects of each embodiment of the present invention will become clearer with reference to the following detailed description in conjunction with the attached drawings. In the attached drawings, the same or similar reference numerals indicate the same or similar elements.
[0011] [Figure 1] A schematic diagram of an exemplary environment in which embodiments of the present invention can be realized is shown. [Figure 2] A schematic diagram of a multi-party data query process according to several embodiments of the present invention is shown. [Figure 3]A flowchart of a data query method according to several embodiments of the present invention is shown. [Figure 4] A schematic block diagram of a data query device according to several embodiments of the present invention is shown. [Figure 5] A block diagram of an electronic device that can realize several embodiments of the present invention is shown. [Modes for carrying out the invention]
[0012] The embodiments of the present invention will be described in more detail below with reference to the drawings. Although specific embodiments of the present invention are shown in the drawings, the present invention may be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, it should be understood that these embodiments are provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and should not be used to limit the scope of protection of the present invention.
[0013] In the description of embodiments of the present invention, the term “including” and similar terms should be understood as open inclusion, i.e., “including but not limited to.” The term “based on” should be understood as “based at least partially.” The term “one embodiment” or “the embodiment” should be understood as “at least one embodiment.” The term “several embodiments” should be understood as “at least several embodiments.” Other explicit and implicit definitions may include the following:
[0014] The phrase "in response to" indicates the occurrence of a corresponding event or the fulfillment of a condition. It should be understood that the timing of execution of a subsequent action performed in response to such an event or condition is not necessarily strongly correlated with the timing of the event's occurrence or the condition's fulfillment. Depending on the circumstances, the subsequent action may be executed simultaneously with the occurrence of the event or the fulfillment of the condition, or it may be executed some time after the event or condition has been fulfilled.
[0015] It is understood that data related to this technical solution (including, but not limited to, the data itself, the acquisition of the data, or the use of the data) must comply with applicable laws and related designated requirements.
[0016] It is understood that, before using the technical solutions disclosed in each embodiment of the present invention, the user should be notified in an appropriate manner in accordance with relevant laws and regulations regarding the type of personal information related to the present invention, the scope of use, the usage scenario, etc., and their consent should be obtained.
[0017] For example, in response to receiving an active request from a user, a prompt is sent to the user, and in response to receiving an unsolicited request from a user, for example, information is sent to the user to explicitly inform them that the requested operation requires the acquisition and use of the user's personal information. This allows the user to independently choose whether or not to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present invention, based on the information provided.
[0018] As a selective but non-limiting implementation form, a method of transmitting presentation information to a user in response to receiving an uncommitted request from the user may, for example, be a method that utilizes a pop-up window, and the presentation information can be displayed in the form of text within the pop-up window. Further, the pop-up window may further include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0019] It should be understood that the above-described notification and user authorization acquisition process are merely schematic and do not limit the implementation forms of the present invention, and other methods that comply with relevant laws and regulations can also be applied to the implementation forms of the present invention.
[0020] As described above, since different data are often held by different owners, it is necessary to complete the common analysis of the data while ensuring the data security of all related parties. Multi-party data common analysis based on SMPC technology is a common method for solving the data circulation problem and exerting the data value. However, with the increase in the amount of data, the computational complexity and communication complexity of such a method increase significantly. In order to meet the actual business requirements, it is necessary to design an efficient common query protocol.
[0021] Multi-party database queries based on structured query language (SQL) have application value in many application scenarios. In the scenario of traditional database queries, data in the form of relational tables and other tables are held by the same data owner, and query analysis can be performed using a traditional SQL query engine. However, in the multi-party database query scenario, relational tables (or data) are held by different data owners, and it is necessary to design a multi-party security query protocol to complete query analysis while ensuring data security.
[0022] Embodiments of the present invention propose a data query solution for realizing aggregate queries of multi-party data. According to this solution, participants in a multi-party security query can obtain query results by performing operations according to the aggregate query protocol proposed herein and inferring information about other participants based on the data generated during the operations. This solution implements an aggregate query process based on multi-party security computing and effectively solves common query problems in multi-party scenarios.
[0023] Hereinafter, several exemplary embodiments of the present invention will be described in conjunction with Figures 1 and 2.
[0024] Figure 1 shows a schematic diagram of an exemplary environment 100 in which an embodiment of the present invention can be realized.
[0025] Environment 100 may include multiple devices such as a first device 105 and a second device 110. These devices may be of any type, including terminal devices and servers. Terminal devices may include, but are not limited to, mobile devices, fixed devices, portable devices, and include mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio receivers, e-book devices, virtual reality (VR) all-in-one devices, game consoles, gamebooks, or any combination thereof, and include accessories and peripherals for these devices or any combination thereof. In some embodiments, terminal devices may also support any type of user interface (e.g., a "wearable" circuit).
[0026] Servers may include, but are not limited to, mainframes, edge computing nodes, rack-mount servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, and desktop computers. In some embodiments, servers may be implemented as virtual machines, containers, or bare-metal servers.
[0027] Environment 100 may further include multiple databases, such as a first database 115 and a second database 120. The first device 105 and the second device 110 can retrieve data from the first database 115 and the second database 120, respectively, and store the data in the first database 115 and the second database 120. The data stored in the first database 115 and the second database 120 may belong to different data owners. Accordingly, the first device 105 and the second device 110 may be associated with different data owners.
[0028] The first device 105 and the second device 110 can communicate via wired and / or wireless means, such as transmitting data and information. Communication in environment 100 can follow any suitable communication protocol, and the scope of the present invention is not limited in this respect.
[0029] It should be understood that Figure 1 shows two devices 105 and 110 and their corresponding databases 115 and 120 for illustrative purposes only. In implementation, the environment 100 may include multiple devices that can communicate with each other. Each device can retrieve data from one or more databases and store that data in one or more different databases.
[0030] Furthermore, it should be understood that the configuration and function of environment 100 are described for illustrative purposes only and do not imply any limitation to the scope of the invention. For example, depending on the specific implementation, environment 100 may include more or fewer devices, units, modules, assemblies, and / or components.
[0031] In environment 100, the first device 105 and the second device 110 can execute aggregate queries for multi-party data in response to data query requests. An example process for aggregate queries for multi-party data will be described below in conjunction with Figure 2.
[0032] Figure 2 shows a process 200 for multi-party data querying performed by the first apparatus 105 and the second apparatus 110 according to several embodiments of the present invention.
[0033] As shown in Figure 2, in 205, the first device 105 can represent a query participant or computing party and retrieve first data based on a data query request. A data query request may come from any data query party, such as a data owner or other third party associated with the first device 105 or the second device 110. Such a request may be implemented in any suitable way, for example, via a Structured Query Language (SQL) query statement.
[0034] The data to be queried may be any appropriate data, such as user data from data owners (e.g., related data owners of the first device 105 and the second device 110 and other data owners). The data to be queried may have multiple attributes. Each attribute may have one or more values. A data query request indicates a first data attribute to be classified and aggregated (also called the “data attribute to be aggregated” or “attributes to aggregate”), and at least a second and a third data attribute used for that classification.
[0035] The first data may be generated by the first device 105 based on data stored in a local or associated device (such as the first database 115). In some embodiments, the first data may be generated by encrypting the stored data, thereby further improving data security. For example, the first device 105 can retrieve data associated with a second data attribute from the local or associated first database 115 based on a second data attribute indicated by a data query request. Subsequently, the first data can be generated by performing one-hot encoding, such as one-hot coding, on the stored data based on the value of the second data attribute.
[0036] For example, the stored data to be retrieved may be in the form of a table (such as a relational table), where each column corresponds to one attribute of the data being queried, as shown in Table 1 below. [Table 1]
[0037] In Table 1, Attribute 1 represents the second data attribute, and its value includes XY and XX. For example, Table 1 also includes an Identifier (ID) column.
[0038] Based on the XY and XX values contained in attribute 1, one-hot encoding can be performed on the data in Table 1, as shown in Table 2 below. [Table 2]
[0039] In Table 1, the value of attribute 1 for ID1 is XY. Therefore, after one-hot encoding, for the row where ID1 is located, the value of the column element corresponding to the value XY is 1, and the value of the column element corresponding to the value XX is 0, as shown in Table 2. As a result, for example, the first data containing the associated data for attribute 1 is obtained.
[0040] In some embodiments, if a local or associated device of the first device 105 holds a first data attribute that is classified and aggregated, the first data may further include data associated with the first data attribute.
[0041] In step 210, the first device 105 secretly fragments the first data and obtains multiple data fragments of the first data, as shown in Table 3 below. [Table 3]
[0042] Here, <...> represents a fragment representation of a variable. Any suitable data fragmentation method or algorithm currently known or to be developed in the future may be used herein, and the scope of the present invention is not limited in this respect.
[0043] The second device 110 can perform the same processing as the first device 105. As shown in Figure 2, in 215, the second device 110 retrieves second data based on the request of a data query. The second data includes at least data associated with a third data attribute indicated by the request of the data query. In some embodiments, the second data may further include associated data of the first data attribute that is classified and aggregated.
[0044] For example, the second device 110 can retrieve related data of the third data attribute stored in a local or associated device (such as the second database 120), as shown in Table 4 below. [Table 4]
[0045] In Table 4, attribute 2 represents the third data attribute, and its values include AA, BB, and CC. For example, the second data further includes related data for the attribute to be aggregated.
[0046] The second data can be generated by performing one-hot encoding (e.g., one-hot encoding) on the stored data based on the value of the third data attribute, as shown in Table 5 below. [Table 5]
[0047] In Table 4, since the value of attribute 2 for ID1 is AA, after one-hot encoding, for the row where ID1 is located, the value of the column element corresponding to value AA is 1, the value of the column element corresponding to value BB is 0, and the value of the column element corresponding to value CC is 0, as shown in Table 5. As a result, for example, a second set of data including attribute 2 and related data of the attribute to be aggregated can be obtained.
[0048] In step 220, the second device 110 secretly fragments the second data and obtains multiple data fragments of the second data, as shown in Table 6 below. [Table 6]
[0049] In block 225, the first device 105 and the second device 110 exchange data fragments. For example, the first device 105 can locally hold one data fragment of the first data (referred to as the "first data fragment") and send the other data fragment of the first data (referred to as the "fourth data fragment") to the second device 110. Similarly, the second device 110 can locally hold one data fragment of the second data (referred to as the "sixth data fragment") and send the other data fragment of the second data (referred to as the "second data fragment") to the first device 105.
[0050] In block 230, a data fragment (referred to as the "third data fragment") of the data query result can be generated based on the first data fragment of the first data held by the first device 105 and the second data fragment of the second data received from the second device 110. The third data fragment includes elements corresponding to the first data attributes that are classified and aggregated based on at least the second and third data attributes.
[0051] For example, in the first device 105, the first data fragment (for example, in the format of Table 3) and the second data fragment (for example, in the format of Table 6) may be concatenated in the format shown in Table 7 below. [Table 7]
[0052] As shown in Table 7, the data can be sorted based on IDs during the concatenation process. Any suitable data sorting method can be employed, and the scope of the present invention is not limited in this respect.
[0053] In some embodiments, to generate a third data fragment of the results of a data query, the first device 105 can generate a data fragment (referred to as the "seventh data fragment") that includes elements corresponding to multiple combinations of values of the second and third data attributes, based on the first data fragment of the first data and the second data fragment of the second data. Subsequently, the third data fragment can be generated by classifying and aggregating the first data attributes to be aggregated, with one combination of those multiple values as one classification.
[0054] For example, by multiplying the columns of values (e.g., XY, XX, etc.) of the first data attribute of the first data fragment (shown in Table 3, for example) and the columns of values (e.g., AA, BB, CC, etc.) of the second attribute of the second data fragment (shown in Table 6, for example) element by element, the data fragment shown in Table 8 below can be obtained. [Table 8]
[0055] Here, (XY,AA), (XY,BB), (XY,CC), (XX,AA), (XX,BB), and (XX,CC) represent various classifications.
[0056] Next, each classification is multiplied by the attribute to be aggregated according to the corresponding element, and the multiplication results are summed to obtain a third data fragment. In some embodiments, the third data fragment may further include an element indicating whether the value of the first data attribute to be aggregated is empty for each classification. For example, the count result for each classification can be calculated to determine whether that classification is empty or not. The resulting third data fragment is shown in Table 9 below. [Table 9]
[0057] Here, <0> This indicates that the classification is empty. <1> This indicates that the classification is not empty.
[0058] In step 235, the second device 110 generates a data fragment (referred to as the "fifth data fragment") of the data query result based on the fourth data fragment of the first data received from the first device 105 and the sixth data fragment of the second data that is pending. The fifth data fragment contains elements corresponding to the first data attributes that are classified and aggregated based on the fourth data fragment of the first data and the sixth data fragment of the second data. The fifth data fragment can be generated by the second device 110 using the same operations as the first device 105, so specific details will not be repeated.
[0059] At point 240, the first device 105 exchanges data fragments of the data query results with the second device 110. For example, the first device 105 sends a third data fragment of the data query results to the second device 110 and receives a fifth data fragment of the data query results from the second device 110.
[0060] In step 245, the first device 105 generates the results of a data query based on the third data fragment and the fifth data fragment. For example, the first device 105 can recover the final result based on the third data fragment and the fifth data fragment, and obtain the final query result by deleting rows with empty classifications, as shown in Table 10. [Table 10]
[0061] Accordingly, the second device 110 can recover the final query result (not shown) based on the locally generated query result data fragment and the data fragment received from the first device 105, and the specific operation is the same as that of the first device 105, so it will not be repeated.
[0062] The aggregated query solution according to the embodiment of the present invention is simpler and more effective, significantly reducing the computational and communication complexity for all parties involved, and enabling secure, reliable, and efficient data queries.
[0063] The following describes an exemplary algorithm flow. In this example, the `group by` keyword in an SQL query statement is used to perform security computation in a multi-party security computation scenario. Without loss of generality, we assume that there are two participants, P0 (e.g., associated with the first device 105) and P1 (e.g., associated with the second device 110), each owning a corresponding relational table. Both P0 and P1 are data owners and computing parties, completing a common query while ensuring privacy. `group by` can have one or more data attributes, each attribute belonging to either participant P0 or P1. The attributes in `group by` are not necessarily the result of a common computation between participants P0 and P1, and may not be, for example, dense intermediate data in multi-party computing.
[0064] In the setup phase, we assume that both P0 and P1 have a single `group by` column (attribute), and that after sorting based on the Identity Identifier (ID) column, they will contain N tuples (rows). We assume that P0 has a relational table L, with `group by` column (attribute) k0, and the column (attribute) that needs to be aggregated after `group by` is v0. We assume that P1 has a relational table R, with `group by` column (attribute) k1, and the column (attribute) that needs to be aggregated after `group by` is v1.
[0065] In the local computation phase, P0 performs one-hot coding on column L[k0] of the relational table L=(k0,v0). For example, it reconstructs the relational table,
number
Number
[0066] After one-hot encoding, the value of the attribute s corresponding to the j-th tuple in the relational table i 0 is
Number
[0067] Similarly, P1 performs one-hot encoding on the column L[k1] of the relational table R = (k1, v1). For example, the relational table is reconstructed to
Number
Number
Number
[0068] In the multi-party computing phase, P0 and P1 secretly perform fragment and data fragment exchange, and P0
number
number
[0069] The following classification and aggregation algorithm can be used to generate query results.
number
[0070] Relational Table <s>=( <c0> , <c1>,Agg(v0),Agg(v1), <f>The output can be a relational table in which each tuple of S, c0 and c1, represents the current classification (or "category"), Agg(v0) and Agg(v1) represent the aggregated value in the current classification, and f indicates whether the classification is empty or not.
[0071] For illustrative purposes only and without implying any limitations, the group by computation process in a two-person computing party has been described, but it should be understood that this computation process may be applicable to scenarios with any number of computing parties. Furthermore, for illustrative purposes only and without implying any limitations, a computation process where each participant has one group by attribute has been described, but it should be understood that this process may be applicable to scenarios with any multiple data attributes.
[0072] Figure 3 shows a flowchart of an exemplary data query method 300 according to several embodiments of the present invention. Method 300 may be implemented in a first apparatus 105 or a second apparatus 110. For the sake of discussion, Method 300 will be described from the perspective of the first apparatus 105.
[0073] As shown in Figure 3, in block 310, the first device 105 retrieves first data based on a data query request, the request specifies first data attributes to be classified and aggregated, and at least second and third data attributes used for classification, the first data including at least data associated with the second data attributes. In block 320, the first device 105 generates a third data fragment of the data query result based on a first data fragment of the first data and a second data fragment of the second data received from the second device 110, the second data including at least data associated with the third data attributes, and the third data fragment including elements corresponding to the first data attributes that are classified and aggregated based at least on the second and third data attributes.
[0074] In some embodiments, the first device 105 can transmit a fourth data fragment of the first data to the second device 110.
[0075] In some embodiments, the first device 105 can receive a fifth data fragment of the data query results from the second device 110, the fifth data fragment containing elements corresponding to first data attributes that are classified and aggregated based on the fourth data fragment of the first data and the sixth data fragment of the second data. The first device 105 can then generate the data query results based on the third data fragment and the fifth data fragment.
[0076] In some embodiments, the first device 105 can generate first data by retrieving data associated with a stored second data attribute based on a data query request and performing one-hot encoding on the data based on the value of the second data attribute.
[0077] In some embodiments, the first data may further include data associated with the first data attributes that are classified and aggregated.
[0078] In some embodiments, the first device 105 can generate a seventh data fragment based on a first data fragment of first data and a second data fragment of second data, the seventh data fragment containing elements corresponding to multiple combinations of values of the second data attribute and the third data attribute. The first device 105 can further generate a third data fragment by classifying and aggregating the first data attribute, with one combination of those multiple values as a classification.
[0079] In some embodiments, the third data fragment further includes an element indicating whether the value of the aggregated first data attribute is empty for each classification.
[0080] In some embodiments, the first device 105 can transmit a third data fragment to the second device 110, causing the second device 110 to generate the results of a data query.
[0081] In relation to the operation of the first apparatus 105 and the second apparatus 110, it should be understood that the features and corresponding effects described above are also applicable to method 300, with reference to Figures 1 and 2, and will not be repeated here.
[0082] Figure 4 shows a schematic block diagram of the configuration of a device 400 for data querying according to several embodiments of the present invention. The device 400 may be implemented by the first device 105 or the second device 110 in Figure 1. For the sake of discussion, the device 400 will be described in terms of the first device 105.
[0083] As shown in Figure 4, the apparatus 400 includes an acquisition module 410 and a first generation module 415. The acquisition module 410 is configured to acquire first data via the first apparatus based on a data query request, the request specifying first data attributes to be classified and aggregated, and at least second and third data attributes used for classification, the first data including at least data associated with the second data attributes. The first generation module 415 is configured to generate a third data fragment of the data query result based on a first data fragment of the first data and a second data fragment of second data received from the second apparatus, the second data including at least data associated with the third data attributes, and the third data fragment including elements corresponding to the first data attributes that are classified and aggregated based at least on the second and third data attributes.
[0084] In some embodiments, the apparatus 400 may further include a transmission module configured to transmit a fourth data fragment of the first data to the device.
[0085] In some embodiments, the apparatus 400 may further include a receiving module configured to receive a fifth data fragment of the results of a data query from a device, wherein the fifth data fragment includes elements corresponding to a first data attribute that is classified and aggregated based on a fourth data fragment of a first data and a sixth data fragment of a second data, and a second generating module configured to generate the results of a data query based on the third data fragment and the fifth data fragment.
[0086] In some embodiments, the acquisition module 410 may further be configured to acquire data associated with a stored second data attribute based on a data query request, and to generate first data by performing one-hot coding on the data based on the value of the second data attribute.
[0087] In some embodiments, the first data may further include data associated with the first data attributes that are classified and aggregated.
[0088] In some embodiments, the first generation module 410 may further generate a seventh data fragment based on a first data fragment of the first data and a second data fragment of the second data, and the seventh data fragment may be configured to generate a third data fragment by classifying and aggregating the first data attributes, with the seventh data fragment containing elements corresponding to multiple combinations of values of the second and third data attributes, and one combination of those multiple values as one classification.
[0089] In some embodiments, the third data fragment may further include an element indicating whether the value of the first data attribute being aggregated is empty for each classification.
[0090] In some embodiments, the transmission module may further be configured to transmit a third data fragment to a second device, thereby causing the second device to generate the results of a data query.
[0091] In relation to the operation of the first device 105 and the second device 110, it should be understood that the above features and corresponding effects are also applicable to device 400, as described above, with reference to Figures 1 and 2, and will not be repeated here.
[0092] Figure 5 shows a block diagram of an electronic device 500 that can carry out one or more embodiments of the present invention. It should be understood that the electronic device 500 shown in Figure 5 is illustrative only and should not limit the function and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 may be used to realize the electronic device 110 of Figure 1.
[0093] As shown in Figure 5, the electronic device 500 is in the form of a general-purpose electronic device. The electronic device 500 may include, but is not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be an actual or virtual processor and can perform various processes based on a program stored in the memory 520. In a multiprocessor system, the parallel processing capability of the electronic device 500 is improved by having multiple processing units execute computer executable instructions in parallel.
[0094] The electronic device 500 typically includes multiple computer storage media. Such media may include, but are not limited to, volatile and non-volatile media, removable and non-removable media, and may be any obtainable media accessible by the electronic device 500. Memory 520 may be volatile memory (e.g., registers, fast cache, random access memory (RAM), etc.), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.), or a specific combination thereof. Storage device 530 may be removable or non-removable media, may include machine-readable media such as flash memory drives, magnetic disks, or any other media, may be used to store information and / or data (e.g., training data for training), and may be accessible within the electronic device 500.
[0095] The electronic device 500 may further include other removable / non-removable, volatile / non-volatile storage media. Not shown in Figure 5, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a “floppy disk”) and a removable optical disk drive for reading from or writing to a non-volatile optical disk may be provided. In these cases, each drive may be connected to a path (not shown) by one or more data medium interfaces. The memory 520 may also include a computer program product 525 having one or more program modules, which are configured to perform various methods or operations of various embodiments of the present invention.
[0096] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 500 may be implemented as a single computing cluster or multiple computing machines, which can communicate via communication connections. Therefore, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0097] The input device 550 may be one or more input devices such as a mouse, keyboard, or trackball. The output device 560 may be one or more output devices such as a display, speaker, or printer. The electronic device 500 may further communicate with one or more external devices (not shown), such as a storage device or display device, via the communication unit 540 as needed, or with one or more devices that enable a user to interact with the electronic device 500, or with any device (e.g., a netbook card, modem) that enables the electronic device 500 to communicate with one or more other electronic devices. Such communication may be performed via an input / output (I / O) interface (not shown).
[0098] An exemplary embodiment of the present invention provides a computer-readable storage medium in which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to realize the above method. An exemplary embodiment of the present invention further provides a computer program product, the computer program product including computer-executable instructions, which is tangibly stored in a non-temporary computer-readable medium, and the computer-executable instructions are executed by a processor to realize the above method.
[0099] Herein, each aspect of the present invention has been described with reference to flowcharts and / or block diagrams of methods, apparatus, devices, and computer program products realized by the present invention. It should be understood that each box in the flowcharts and / or block diagrams, and each combination of boxes in the flowcharts and / or block diagrams, may be realized by computer-readable program instructions.
[0100] These computer-readable program instructions may be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device to generate a machine that, when executed by the processing unit of the computer or other programmable data processing device, generates a device for performing one or more functions / operations specified in one or more boxes in a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which may operate the computer, programmable data processing device, and / or other device in a particular manner so that the computer-readable medium containing the instructions constitutes a product containing instructions for performing each aspect of the functions / operations specified in one or more boxes in a flowchart and / or block diagram.
[0101] By loading computer-readable program instructions into a computer, other programmable data processing device, or other device, a series of operational steps are performed on the computer, other programmable data processing device, or other device to generate a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing device, or other device to perform the functions / operations specified in one or more boxes in a flowchart and / or block diagram.
[0102] The flowcharts and block diagrams in the attached drawings illustrate the implementable architectures, functions, and operations of several implementable systems, methods, and computer program products according to the present invention. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or part of an instruction, and a module, program segment, or part of an instruction contains one or more executable instructions for implementing a specified logical function. In some implementations as replacements, the functions represented in the boxes may occur in a different order than those shown in the drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or in reverse order depending on the functions involved. It should be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, may be implemented by a special-purpose hardware-based system that performs a specified function or operation, or by a combination of special-purpose hardware and computer instructions.
[0103] While the various realizations of the present invention have been described above, the above descriptions are illustrative, not exhaustive, and not limited to the realizations disclosed. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the realizations described. The choice of terms used herein is intended to best interpret the principle, practical application, or improvement to the art in the market of each realization, or to enable those skilled in the art to understand each realization disclosed herein.< / f> < / c0> < / s>
Claims
1. A method for data querying, said method A step of obtaining first data via a first device based on a data query request, wherein the request specifies first data attributes to be classified and aggregated, and at least second and third data attributes used for the classification, and the first data includes at least data associated with the second data attributes. A step of generating a third data fragment of the result of the data query based on a first data fragment of the first data and a second data fragment of the second data received from a second device, wherein the second data includes at least data associated with the third data attribute, and the third data fragment includes elements corresponding to the first data attribute that are classified and aggregated based at least on the second and third data attributes, method.
2. The step of transmitting a fourth data fragment of the first data to the second device further includes: The method according to claim 1.
3. A step of receiving a fifth data fragment of the result of the data query from the second device, wherein the fifth data fragment includes elements corresponding to the first data attribute that is classified and aggregated based on the fourth data fragment of the first data and the sixth data fragment of the second data. The further step includes generating the result of the data query based on the third data fragment and the fifth data fragment, The method according to claim 2.
4. The step of obtaining the aforementioned first data is: The steps include: obtaining data associated with the stored second data attribute based on the request of the data query; The step of generating the first data by performing one-hot coding on the data based on the value of the second data attribute, The method according to claim 1.
5. The first data further includes data associated with the first data attributes that are classified and aggregated. The method according to claim 1.
6. The step of generating the third data fragment is: A step of generating a seventh data fragment based on the first data fragment of the first data and the second data fragment of the second data, wherein the seventh data fragment includes elements corresponding to a combination of multiple values of the second data attribute and the third data attribute, The steps include generating the third data fragment by classifying and aggregating the first data attributes, with one combination of the aforementioned multiple values as one classification, The method according to claim 1.
7. The third data fragment further includes an element indicating whether the value of the aggregated first data attribute is empty for each classification. The method according to claim 1.
8. The further step includes transmitting the third data fragment to the second device so that the second device generates the results of the data query, The method according to claim 1.
9. A device for data querying, said device, An acquisition module configured to acquire first data based on a data query request via a first device, wherein the request specifies first data attributes to be classified and aggregated, and at least second and third data attributes used for the classification, and the first data includes at least data associated with the second data attributes, A first generation module configured to generate a third data fragment of the results of a data query based on a first data fragment of the first data and a second data fragment of the second data received from a second device, wherein the second data includes at least data associated with the third data attribute, and the third data fragment includes elements corresponding to the first data attribute that are classified and aggregated based at least on the second and third data attributes, Device.
10. The second device further includes a transmission module configured to transmit a fourth data fragment of the first data to the second device. The apparatus according to claim 9.
11. A receiving module configured to receive a fifth data fragment of the results of the data query from the second device, wherein the fifth data fragment includes elements corresponding to the first data attributes that are classified and aggregated based on the fourth data fragment of the first data and the sixth data fragment of the second data, The system further includes a second generation module configured to generate the results of the data query based on the third data fragment and the fifth data fragment, The apparatus according to claim 10.
12. The aforementioned acquisition module further, Based on the request of the data query, retrieve the data associated with the stored second data attribute. The system is configured to generate the first data by performing one-hot coding on the data based on the value of the second data attribute. The apparatus according to claim 9.
13. The first data further includes data associated with the first data attributes that are classified and aggregated. The apparatus according to claim 9.
14. The preceding 1 generation module further, Based on the first data fragment of the first data and the second data fragment of the second data, a seventh data fragment is generated, and the seventh data fragment includes elements corresponding to combinations of multiple values of the second data attribute and the third data attribute. The system is configured to generate the third data fragment by classifying and aggregating the first data attributes, with one combination of the aforementioned multiple values being treated as a single classification. The apparatus according to claim 9.
15. The third data fragment further includes an element indicating whether the value of the aggregated first data attribute is empty for each classification. The apparatus according to claim 9.
16. The aforementioned transmission module further, The system is configured to transmit the third data fragment to the second device so that the second device generates the results of the data query. The apparatus according to claim 10.
17. An electronic device, At least one processing unit, Includes at least one memory, The at least one memory is coupled to the at least one processing unit to store instructions to be executed by the at least one processing unit, and when the instructions are executed by the at least one processing unit, causes the electronic device to perform the method according to any one of claims 1 to 8. electronic equipment.
18. A computer program is stored therein, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is realized. Computer-readable storage medium.