INTER-ORGANIZATIONAL LEARNER USER MATCHING METHOD AND SYSTEM WITHOUT DATA SHARING
Patent Information
- Authority / Receiving Office
- TR · TR
- Patent Type
- Applications
- Current Assignee / Owner
- ATAKUM HALK EĞİTİM MERKEZİ MÜDÜRLÜĞÜ
- Filing Date
- 2026-04-23
- Publication Date
- 2026-06-22
Abstract
Description
INTER-ORGANIZATIONAL LEARNING WITHOUT DATA SHARING USER MATCHING METHOD AND SYSTEM 1. TECHNICAL FIELD This invention encompasses distributed data processing, privacy-protected record matching, and data engineering. Similarity calculation, user behavior analysis, and computer-implemented methods. It relates to learner-user matching methods. The invention specifically involves the organization of user records belonging to different institutions or different data systems. Raw data belonging to the user is not shared between systems; each institution has its own data. processing in the environment, conversion into a common comparison structure, converted data by comparing the representations and utilizing verified matching results Inter-organizational user matching by updating the matching criteria It relates to a method and system that enables its realization. 2. STATE OF KNOWLEDGE OF THE ART Record matching is the process of identifying records in different datasets that belong to the same natural person, user, or... It is a technical field aimed at determining whether something belongs to an entity or not. The methods in this field... In general, this involves selecting the areas to be compared, identifying candidate registration pairs, This involves performing a similarity calculation and applying a decision criterion. In the privacy-protected record matching area, the parties' directly identifying data... Various techniques have been developed to enable matching without opening them to each other. Statistical and computational aspects of record matching in field studies; privacy. The issue has been addressed along with its dimensions, and fundamental, open technical problems have been identified. In Bloom's filter-based approaches, fault-tolerant encoded data is used instead of explicit ID data. Approximate matching can be done using the representations. This approach, raw While it is important that identifiers are not explicitly shared, Matching success largely depends on the selected fields, coding rules, and similarity. It adheres to the criteria. Comparing all records in large datasets is computationally efficient. Since 35 was not suitable, alternative double reduction techniques were developed. In these techniques, detailed The record pairs to be included in the similarity calculation are determined using pre-defined reduction rules. It is limited. In learnable candidate narrowing approaches, however, the fixed number is determined manually. 1.5 Instead of rules, data-learned reduction structures can yield more suitable results. It has been shown. Matching and federated learning are considered together on vertically divided data. In the studies, the effect of matching errors on learning was examined and the parties' It has been shown that collaborative learning can occur without explicitly sharing data. Secure multilateral computational approaches require a reliable third party. fault-tolerant, privacy-protected record matching can be performed without hearing about it. It is stated. Application examples include corporate privacy-protected record matching. This demonstrates its applicability in real-world use cases on a large scale. However, the following technical problems persist in the current state of the technology: • A common framework for data fields, naming conventions, and data formats across different institutions. the absence of a standard and clear process chain in converting it into a structure, • Behavioral, temporal, statistical, and pattern-based user characteristics combined fixed-length, comparable, and identical in structure used across institutions the fact that a data representation that can be produced is not clearly defined, • Raw data or aggregated intermediate data should only be stored without being taken outside the organization. a detailed guide on matching with transformed transfer representations The communication structure is not sufficiently clear. • Candidate pair reduction, multi-component similarity calculation, threshold application, and indeterminate case the fact that its management is not defined together within a single integrated method, • Weighing using verified match and verified non-match samples coefficients, tolerance ranges, decision thresholds, and candidate narrowing criteria an updated, learning matching structure in a clear and detailed form It is undefined. The invention in question aims to remedy the technical deficiencies mentioned above by analyzing the data each User-owned multi-component attribute sets that transform data into a common format within the organization. extracting, converting into fixed-length comparison representations, and recovering raw data Candidate narrowing and similarity calculation based on unsuitable transfer representations. a distributed system that performs and updates its matching criteria based on verified results. It offers methods and systems. 35 2 5 3. PURPOSE OF THE INVENTION The purpose of this invention is to organize user records in data systems belonging to different institutions. A computer that enables the matching of raw data belonging to the user without transferring it outside the organization. The aim is to present a method and system implemented by [the organization / institution]. Another objective of the invention is to consolidate user data from different data structures into a common data structure. transformative, behavioral, temporal, statistical, and pattern-based user characteristics together a processing system that uses and produces comparable technical data representations from them It is to define. Another objective of the invention is to narrow down the candidate record pair on the generated data representations. Component-based similarity calculation, weighted combined match score generation, threshold assessment. and to provide a distributed matching mechanism that includes uncertainty management. Another purpose of the invention is to derive from verified matching and verified non-matching samples. by utilizing feature weights, tolerance ranges, decision thresholds, and candidate narrowing. The goal is to establish a learning, adaptive structure that updates its criteria. Another objective of the invention is to create a system that does not require a central raw data pool, but rather allows external access to the data. a communication system that produces transformed data representation, candidate selection information and communication control information. The aim is to improve data security and reduce data transmission load. 4. EXPLANATION OF THE FIGURES 4.1 No form is included in the application. 5. EXPLANATION OF REFERENCES IN THE FIGURES 5.1 Since no form is included in the application, a reference list is not provided. DETAILED DESCRIPTION OF THE 6TH INVENTION 6.1 General system structure 6.1.1 Invention of user mapping between multiple institutions or data systems It relates to a distributed data processing structure that enables this. 6.1.2 Each institution shall only store user data located within its own local domain. It works in that environment. 35 6.1.3 Raw data shall not be transferred outside the institution's boundaries. 6.1.4 Transformed data created for comparison purposes only and sent outside the institution. Representatives, candidate selection information, and communication control information are communicated. 3 5 6.1.5 System; local data retrieval unit, common format conversion unit, attribute subtraction unit, comparison representation unit, transfer representation unit unit, candidate narrowing unit, similarity calculation unit, decision making unit, learner adaptation unit, communication layer and control log unit It consists of. 6.2 Local data retrieval and conversion to common format unit 6.2.1 The local data retrieval unit retrieves user-related data from corporate data sources. It is receiving. 6.2.2 The data in question includes transaction logs, session logs, event logs, and time. Stamp information, usage frequency information, processing times, success or failure. logs, interaction patterns, client-type information, and similar user behavior data. It contains the data. 6.2.3 Common Format Conversion Unit, data structures from different institutions In order to make them comparable, the following steps are taken: • mapping domain names to a common schema dictionary, • Consolidating the date and time fields into a single time display, • bringing numerical fields to a common scale, • Conversion of categorical fields to common code fields, • Standardization of missing data indicators, • merging duplicate records, • semantically equivalent data fields under the same attribute gathering. 6.2.4 At the end of this stage, independent of different corporate data structures, A shared intermediate data structure is obtained. 6.3 Feature Extraction Unit 6.3.1 Feature extraction unit, user-related from a common intermediate data structure. It generates multi-component feature sets. 6.3.2 The extracted features include at least three of the following groups: 35 • behavioral attributes, • temporal attributes, • Statistical summary attributes, • pattern attributes, 4 5 • Consistency and variability attributes. 6.3.3 Behavioral attributes; session frequency, transaction intensity, event type distribution, from completion rate, repetition rate, navigation transition and resource access patterns It is derived. 6.3.4 Temporal attributes; intraday usage patterns, weekday and weekend distribution, time difference between events, session duration distribution, seasonal intensity and sequential It is derived from usage ranges. 6.3.5 Statistical summary attributes; mean, median, percentile values, distribution It is derived from the amplitude, standard deviation, entropy, and coefficient of variation. 6.3.6 Pattern attributes; process transition sequences, repeating subsequences, event type It is derived from structures of being seen together and sequential behavioral patterns. 6.3.7 Consistency and variability attributes; across different time periods or different the consistency and variability of behavior that emerged in the sessions It shows. 6.4 Comparison representation creation unit 6.4.1 The comparison representation creation unit sets the extracted feature sets as constants. It transforms it into a lengthy comparative representation. 6.4.2 During this conversion; • Attribute values are displayed on a common scale, • It is divided into specific intervals when necessary. • Empty or missing values are marked with separate indicator fields. • Attribute elements are placed in predetermined fixed positions, • The representation structure is tagged with a version number. 6.4.3 Comparative representation is established according to the same rules across institutions. They are directly comparable in structure. 6.4.4 Comparative representation in an application; behavioral sub-representation, temporal sub-representation by combining the representation, statistical sub-representation and pattern sub-representation is obtained. 35 6.5 Transfer representation creation unit 6.5.1 Comparative representation, before being transmitted externally, to a transfer representation. It is being transformed. 5 6.5.2 Transfer representation, directly returning raw data or shared intermediate data. It is a transformed data structure that is unsuitable for production. 6.5.3 This transformation process includes quantification, component grouping, feature masking, and single-item processing. Directional transformation, fixed-length encoding, and partial feature compression. It includes at least one of the following operations. 6.5.4 Thus, data transmitted between institutions is suitable for comparison. However, it does not possess raw data characteristics. 6.6 Candidate Reduction Unit 6.6.1 Candidate narrowing section, to prevent all records from being compared with each other. For this purpose, it generates at least one candidate selection information for each transfer representation. 6.6.2 Candidate selection information; band code, bucket code, block key or rough code. This is in the form of a classification label. 6.6.3 Transfer representations from different systems may only be made from the same or compatible candidate systems. Detailed comparisons are made in terms of pairs of records that have election information. 6.6.4 This structure reduces the number of record pairs to be compared and simplifies the calculation. It keeps its load at a controllable level. 6.7 Similarity calculation unit 6.7.1 Similarity calculation unit, multi-component similarity sub-unit for candidate registration pairs. It generates scores. 6.7.2 The sub-scores must include at least three of the following: • behavioral similarity score, • temporal proximity score, • Statistical overlap score, • pattern fit score, • Consistency score. 6.7.3 Different comparison criteria are used for different attributes. 6.7.4 Absolute difference, common difference, or interval-based calculations for numerical components proximity; overlap coefficient in distribution components; order consistency in ordered patterns or 35 transition adjustments are being implemented. 6.7.5 The sub-scores obtained represent the discrimination level determined for each attribute. The coefficients are weighted and converted into a single combined match score. 6 5 6.8 Decision-making unit 6.8.1 The decision-making unit sets the combined match score at least one upper threshold and at least one lower threshold. It is evaluated with a threshold and at least a tolerance range. 6.8.2 If the combined match score is above the upper threshold value A matching decision is being generated. 6.8.3 If the combined match score is below the lower threshold value A non-matching decision is being made. 6.8.4 Registration if the combined match score falls between the lower and upper thresholds. The pair is marked as an uncertain status. 6.8.5 A record generating a score above the upper threshold with more than one candidate record In that case, the candidate with the highest combined match score will be registered as a single match. One is selected and the others are marked as having multiple match conflicts. 6.9 Learner adaptation unit 6.9.1 The learner adaptation unit, subsequently verified match and verified The system updates its criteria by utilizing examples of mismatches. 6.9.2 The updated criteria include at least one of the following: • weighting coefficients for attribute elements, • tolerance ranges, • Criteria for generating candidate selection information, • upper and lower threshold values, • Boundary values of the uncertain region. 6.9.3 The update process only applies to verified tags located on the local system, more using previously generated matching outputs and performance records is being carried out. 6.9.4 The learner adaptation unit works only on benchmark packets instead of raw data. making changes and updating the criteria in versioned packages It is hiding. 6.9.5 Thanks to this structure, the system adapts to the characteristics of data from different institutions and over time. its own matching behavior according to changing usage patterns It can be updated in 35 ways. 6.10 Communication layer 7 5 6.10.1 The communication layer enables inter-organizational data exchange only through defined communication channels. It is carried out through messages. 6.10.2 Every communication message must include at least the following fields: • system identifier, • Representative version number, • Schema version number, • candidate selection information, • representation of transfer, • timestamp, • integrity verification field. 6.10.3 Raw data, plain ID or direct identifier in communication messages The area is not being moved. 6.10.4 Communication, request-response, bulk packet transfer or periodic synchronization It can be carried out with at least one of the following methods. 6.11 Audit record and traceability section 6.11.1 The system provides a representation of each matching operation, based on the criteria used. audit package including candidate selection information, decision thresholds and decision outcome. It creates a record. 6.11.2 Audit log, retrospective traceability, performance review and learner It is used to verify the adaptation unit. 6.11.3 The audit log does not contain raw data. 6.12 Sequence of work 6.12.1 Local user data is being collected. 6.12.2 Different data structures are converted into a common format. 6.12.3 Behavioral, temporal, statistical, and pattern attributes are extracted. 6.12.4 A fixed-length comparison representation is being created. 6.12.5 The comparative representation is being transformed into an irreproducible transfer representation. 6.12.6 Candidate selection information is being generated. 35 6.12.7 Transfer of representation and communication information to the other institution or institutions It is being communicated. 6.12.8 Candidate registration pairs are being determined. 6.12.9 Multi-component similarity sub-scores are being calculated. 8 5 6.12.10 Combined match points are being generated. 6.12.11 A matching decision is made by evaluating the threshold and tolerance. 6.12.12 Criteria are updated using verified results. 7. HOW THE INVENTION IS APPLIED TO INDUSTRY The invention encompasses educational technologies, learning management systems, and examination and assessment infrastructures. multi-membership systems, public and private sector data integration structures, health, finance, It is applicable in insurance and similar data-intensive privacy fields. The invention involves software developed for cloud environments, on-premises data centers, or hybrid computing architectures. It can be implemented based on this principle. The implementation of the invention involves collecting raw data in a central data pool. It is not necessary. 9
Claims
1. A computer that performs user matching between different data systems. It is a data processing method implemented by; its characteristic is that in every data system User data of different structures are common in the local environment according to a shared schema dictionary. conversion to a standard format, user-related information from converted data. behavioral attributes, temporal attributes, statistical summary attributes, and Removing at least three of the pattern attributes, and the removed attributes having common features. converting it into a fixed-length comparison representation by scaling it, word the subject is not suitable for retrieving raw data from the comparative representation Generating a transfer representation, including at least one candidate selection information for the transfer representation. together, the representation of the transfer in question is in the form of a predefined communication message. Transmission of candidate selection information from different data systems Based on this, the reduction to candidate registration pairs, attributes for candidate registration pairs. Calculating similarity sub-scores separately for each family, A combined match score is obtained by weighting the scores according to their discrimination coefficients. conversion, combined match score must have at least one upper threshold, at least one lower threshold and at most Evaluation with a small tolerance range and matching as a result of this evaluation, This involves generating either a mismatch or an ambiguous situation decision.
2. According to claim 1, it is a data processing method and its characteristic is; the conversion to a common format. mapping domain names to a common schema dictionary, combining date and time fields into a single system. converting to time representation, bringing numerical fields to a common scale, converting categorical fields to common code fields, missing data indicators This involves standardizing the data and merging duplicate records.
3. Data processing method according to claim 1, characterized by its ability to process behavioral attributes in session. frequency, transaction intensity, event type distribution, completion rate, repetition rate, and source access patterns, temporal attributes, time difference between events, session duration distribution, intraday usage patterns, and seasonal intensity; statistical summary. mean, median, percentile values, range, standard of attributes 35 deviations and entropy, while pattern features consist of sequential process transitions and repetitions. It is derived from sub-sequences. 5.
4. This is a data processing method according to claim 1, and its characteristic is; fixed-length comparison. behavioral sub-representation, temporal sub-representation, statistical sub-representation of representation and consists of pattern sub-representations, the sub-representations in question being previously placement in designated fixed locations and a version number of the representation structure It is labeled with.
5. Data processing method according to claim 1, characterized by; quantization of transport representation, component grouping, attribute masking, one-way transformation, fixed length because it involves at least one of the encoding and partial feature compression processes Suitable for retrieving raw data or intermediate data converted to a common format. It is the absence of.
6. Data processing method according to Claim 1, characterized by; transmission of candidate selection information. Bandcode, bucket code, block key, or other representations are generated from specific components of the representation. It must include at least one of the coarse classification labels and a detailed similarity calculation. only on pairs of records where candidate selection information is the same or compatible It is the execution.
7. Data processing method according to Claim 1, characterized by its behavioral similarity sub-scores. similarity score, temporal proximity score, statistical overlap score, and pattern fit. It must include at least three of the sub-scores and each of the sub-scores must be related to the attribute. combined matching score obtained by multiplying by the family's discrimination coefficient It is a transformation.
8. Data processing method according to Claim 1, characterized by its upper threshold of combined match score. If the value is above the threshold, a matching decision is made; if it is below the lower threshold value... In that case, the decision is not to match, and if it falls between the lower and upper thresholds, it is uncertain. The situation involves making a decision.
9. Data processing method according to Claim 1, characterized by its ability to process a record into multiple candidate records. If it produces a score above 35, it will have the highest combined match score. the candidate record being selected as a unique match and the other candidate records being multiple It is marked as a match conflict. 11 5 10. Data processing method according to claim 1, characterized by; verified matching and Weights of feature families are calculated using verified mismatch examples. coefficients, tolerance ranges, candidate selection information production criteria, upper threshold and the lower threshold values and the boundaries of the uncertain state region are defined locally. It is an update.
11. Data processing method according to claim 10, characterized by its ability to update criteria only. Validated tags available in the local system, previous matching outputs, and performance. this is done using records and instead of raw data being sent outside the institution, only This involves communicating versioned benchmark package or representation version information.
12. Data processing method according to claim 1, characterized by; system identifier of the communication message, representation version number, scheme version number, candidate selection information, transfer representation, the communication in question must include a timestamp and integrity verification field. the message contains explicit identity information, raw data or directly identifying information It is not to be moved.
13. It is a data processing method according to claim 1, and its characteristic is that the method has a central raw data It is run on a distributed architecture that does not require a data pool, and each data system has its own... By processing user data within its own boundaries, the organization only transfers it outside of its own boundaries. It is the extraction of representative and communication control information.
14. A data processing system that implements the method according to claims 1 to 13, and its characteristic is that each User data with different structures in the data system is shared with a common schema in the local environment. According to its dictionary, it retrieves at least one local data point that converts to a common format and then converts it to a common format. The transformation unit extracts behavioral attributes from data converted into a common format. extracts temporal features, statistical summary features, and pattern features. At least one feature extraction unit extracts those features as fixed-length files. Transfer that is not suitable for comparison representation and raw data retrieval. At least one representation-generating unit that transforms into a representation is a candidate for transfer representations. At least one candidate narrowing unit that generates 35 election data points, on candidate registration pairs At least one similarity sub-score calculates a combined match score. The calculation unit calculates the combined match score with upper threshold, lower threshold and tolerance range. at least one decision-making unit that makes matching decisions by evaluating and only 12 5 A communication that enables the transmission of communication control information through a transfer representation. It includes a layer.
15. Data processing system according to claim 14, characterized by: verified matching and Weighting of feature families using verified mismatch examples. coefficients, tolerance ranges, candidate selection information production criteria, upper thresholds, and at least one that updates the lower threshold values and the boundaries of the uncertain state region It includes a learner adaptation unit.
16. The data processing system according to claims 14 and 15 is characterized by; for each matching operation representation version, criteria package version, candidate selection information, threshold set used, and a traceability system that stores the decision outcome as an audit log without raw data It includes the unit. 13