Method for searching individuals in big data set in distributed manner
By constructing an undirected graph structure in a large data set and mapping it to a two-dimensional integer grid graph, four types of state initialization and rank adjustments are performed, individual identification and conflict avoidance problems in large-scale node structure data sets are solved, and individual precise positioning and unique identification are realized, which is suitable for multi-field applications.
Patent Information
- Application Number
- CN202510454255.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
In the existing technology, when individual identification and conflict avoidance are found in large-scale node structure data centers, there are problems such as low efficiency, high conflict probability, and difficulty in adjusting dynamic nodes, especially in graph neural networks.
The undirected graph structure is used to map to the two-dimensional integer grid graph, and the four types of state initialization and adjacency detection are used to adjust the order to ensure that the states of adjacent nodes are different, and state expansion or compression is carried out when the state density is high, and a globally unique and locally adaptive coding method system is built.
It realizes accurate positioning and unique identification of individuals in a big data environment, has state unique traceability and compression, and is suitable for distributed deployment in multiple fields, and is suitable for matching maps, image modeling, graph neural networks and file systems.
Smart Images

Figure CN120373427A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for discriminating individual uniqueness in large-scale structured data. Background Art
[0002] When the prior art performs individual identification and conflict avoidance in large-scale structures composed of graph elements, such as neural networks, social networks, and task scheduling graphs, it generally adopts global scheduling, assigns unique identifiers, tags, or account IDs, and hash mapping methods. There are deficiencies in low efficiency, high conflict probability, and difficulty in dynamic node adjustment. Especially in artificial intelligence AI modeling and graph learning, such as graph neural network GNN, the defects of node label conflicts and non-unique state assignments have a significant impact on the training effect. Summary of the Invention
[0003] The object of the present invention is to provide a method for distributedly searching for individuals in a large data set, and the technical problem to be solved is conflict avoidance and unique state assignment in a large-scale node structure data set.
[0004] The present invention adopts the following technical solutions: A method for distributedly searching for individuals in a large data set, which is used for a set of two-dimensional data objects of a limited scale, includes the following steps:
[0005] I. Establish an undirected graph
[0006] In a set of objects with two-dimensional plane characteristics, each object is regarded as an original point, and an undirected graph structure is constructed accordingly. In this undirected graph, each node corresponds to an original point object, and the connection lines between nodes represent the existence of a direct adjacent relationship or semantic association relationship between two objects. The edges in the undirected graph satisfy the constraints of non-crossing and non-overlapping;
[0007] II. Structure mapping
[0008] Map the undirected graph structure into the first quadrant of the plane rectangular coordinate system to form a two-dimensional integer point lattice graph;
[0009] III. Initialization of four types of states
[0010] According to the parity of the X and Y values of each node in the plane rectangular coordinate system, the node states are divided into four types: State 1: X and Y values are odd and odd, State 2: X and Y values are odd and even, State 3: X and Y values are even and odd, State 4: X and Y values are even and even;
[0011] IV. Adjacency detection
[0012] For all pairs of points with connection relationships in the two-dimensional integer point lattice graph, perform state conflict detection where the states of adjacent nodes are different to meet the requirement that the states of adjacent nodes are different;
[0013] V. Sequence Adjustment
[0014] During the status conflict detection process, when a status conflict occurs between adjacent nodes, sequence adjustment is performed. The subsequent nodes are re-searched for mapping positions in the order of right shift, up shift, and diagonal shift until the statuses of all adjacent nodes are different from each other.
[0015] When a new node is added to the undirected graph structure of the method of the present invention, it enters the two-dimensional integer point lattice graph according to the mapping method in Steps 2 to 5, and the four types of statuses are initialized and adjacency detection is performed. If a status conflict occurs, local sequence adjustment is performed.
[0016] When a node in the undirected graph structure of the method of the present invention is deleted, the node position corresponding to it in the two-dimensional integer point lattice graph is released, the status identifier corresponding to the coordinate point is cleared, and the relevant connecting edges are removed together.
[0017] When the node density of the method of the present invention increases or the status space is crowded, status expansion is performed. The status expansion includes: two-layer parity subclassification, adding semantic and label dimensions, and introducing a modulo-N status rotation method.
[0018] When the node density of the method of the present invention increases or the status space is crowded, status compression is performed, and the four statuses are represented by a binary 2-bit coding method. 6. The method for distributed searching for individuals in a large dataset according to claim 1, wherein: the direct adjacent relationship in the first step means that there is a clear structural or spatial proximity between objects; the semantic association relationship means the logical connection established between objects based on meaning, structure, or context.
[0019] In the second step of the method of the present invention, an object with significant features is set as the starting point (A11) and mapped to the coordinate value (1, 1). Then, in the scanning manner from left to right and from bottom to top along the X and Y axes of the plane rectangular coordinate system, other nodes are embedded into the two-dimensional integer grid points one by one.
[0020] Before mapping the undirected graph structure to the first quadrant of the plane rectangular coordinate system in the second step of the method of the present invention, a plane embeddability determination is performed. The condition is: if each object in the object set has four representative basic attributes or description dimensions, and any two of them form a basic attribute or description dimension pair, the plane embeddability condition is satisfied.
[0021] In the fourth step of the method of the present invention, a status conflict detection for adjacent nodes with different statuses from each other is performed in sequence.
[0022] In the fourth step of the method of the present invention, a status conflict detection for adjacent nodes with different statuses from each other is performed, and at the same time, the status detection between adjacent nodes is performed in batches and in parallel.
[0023] Compared with the prior art, the present invention is based on graph structure mapping, parity state classification, and sequence adjustment, and is used for distributed search and precise positioning of individuals in a large dataset. Through four-color state division and local adjustment, a globally unique and locally adaptive coding method system is constructed, which has the characteristics of unique and traceable states, precise backtracking, flexible state compression and expansion, expandable into multiple types of states, strong distributed deployment ability, strong versatility, and is applicable to multiple fields such as planar maps, image modeling, graph neural networks, and file systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic diagram of the structure mapping of an embodiment of the present invention.
[0025] Figure 2 is the schematic diagram of the four-color state coding principle of the present invention.
[0026] Figure 3 is Figure 2 an enlarged view of points Amn, B(m+1)n, Cm(n+1), and D(m+1)(n+1) of
[0027] Figure 4 is a schematic diagram of the state sequence adjustment process of the present invention.
[0028] Figure 5 is the schematic diagram of the original graph structure of an embodiment of the present invention.
[0029] Figure 6 is the backtracking identification flowchart of an embodiment of the present invention..
[0030] Figure 7 is the flowchart of dynamic node addition of an embodiment of the present invention.
[0031] Figure 8 is the input path diagram of the present invention applied in the GNN model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0033] The method of the present invention for distributed searching for individuals in a large dataset (the method) is applicable to a set of two-dimensional data objects of limited scale.
[0034] The method of the present invention includes the following steps:
[0035] I. Establish an undirected graph
[0036] As Figure 1As shown, in a data with two-dimensional plane features, there is a finite-scale object set, such as all humans, all parking spaces in the city, and a set G2 of all data sample individuals in the same database. In this embodiment, each object in the finite-scale object set is regarded as an original point, and an undirected graph structure is constructed based on this to describe the structural relationship between objects. In the undirected graph, each node corresponds to an original point (object), and the edges (connections) between the nodes indicate that there is a direct adjacent relationship or semantic association relationship between the two objects. The construction of the edges in the undirected graph satisfies the basic embedding constraints of non-intersection and non-overlapping, ensuring that the overall graph structure has a certain two-dimensional mappability.
[0037] Direct proximity refers to objects with clear structural or spatial proximity. Common relationships include: proximity in spatial position, such as the border between two houses or parking spaces on a map; direct connection in network structure, such as friend relationships in social networks; distance neighbors with similar features, such as the pair of points with the shortest Euclidean distance in vector space; data sorting or indexing adjacency, such as consecutive records in a file system.
[0038] Semantic association refers to the logical connection between objects based on meaning, structure or context, such as: hyponymy, such as "animal" and "dog"; synonymy, such as "vehicle" and "car"; antonym, such as "high" and "low"; causal relationship, such as "infection" and "fever"; conditional relationship, such as "if...then..."; conditional relationship, such as "if...then..."; hypothetical result relationship, such as "if...then it may be..."; transitional relationship, such as "although...but..."; parallel and / or concomitant relationship, such as "occurs at the same time"; characteristic attribute relationship, such as "apple-red", "user-age". As long as there is a clear connection between two objects in any of the above relationships, the edge relationship between nodes in the undirected graph structure can be established.
[0039] 2. Structure Mapping
[0040] like Figure 5 As shown, the undirected graph structure is mapped to the first quadrant of the plane rectangular coordinate system to form a two-dimensional integer grid graph, or it is called a topological arrangement flattened on the integer points in the first quadrant. In this embodiment, an object (node) with significant features is set as the starting point A11, the node is mapped to the coordinate value (1,1), and then the other nodes are embedded one by one into the two-dimensional integer grid points along the X and Y axes of the plane rectangular coordinate system in a scanning manner from left to right and from bottom to top.
[0041] The method of the present invention constructs a structured grid graph with embeddability, uniqueness, and compressible state representation ability by mapping an undirected graph structure to a two-dimensional integer coordinate space. To ensure that the structure mapping method is applicable to real big data scenarios, it is first necessary to determine the planar embeddability of the significant features (data characteristics) of the object set.
[0042] (1) Determination of planar embeddability
[0043] The determination of planar embeddability is a method for judging whether a set of objects (data objects) is suitable for two-dimensional structure mapping based on the combined relationship of dominant factors, and is used to ensure the clarity and distinguishability of subsequent node coordinate embeddings.
[0044] The specific method for determining planar embeddability is as follows: If each object in a certain object set (i.e., the nodes in the two-dimensional integer point lattice graph structure) can extract four common dominant factors, and any two of these dominant factors can form a dominant characteristic factor pair for constructing a data plane structure with distinguishability and distribution interpretability, it is considered to meet the planar embeddability (two-dimensional embedding) condition. As shown in Table 1.
[0045] Table 1 Typical determination examples
[0046] Application area Dominant factor Planar structure Urban parking management Longitude, latitude, parking space type, usage frequency Longitude-latitude plane Medical atlas analysis Tongue image features, body temperature, time label, symptom words Tongue image × time plane Document system retrieval Keywords, catalog number, timestamp, access volume Keywords × structure plane User behavior model Access frequency, stay time, interest label, user group Time × frequency plane
[0047] "Distinguishability" means that this coordinate structure can map different nodes to non-overlapping or easily distinguishable positions, and "distribution interpretability" means that the nodes have an analyzable and interpretable distribution law in this structure.
[0048] In the method of the present invention, the dominant factor refers to the key feature attributes that can be extracted from each object (i.e., the nodes in the undirected graph), such as time, spatial coordinates, behavior frequency, and type classification. These factors can be used as the basic variables for coordinate axes or state division, and are closely related to the structural positions or semantic labels of the nodes.
[0049] If at least four representative and overall structurally distributed dominant factors can be extracted from each node in the object set, a dominant factor combination can be formed. Among them, two factors with strong distinguishing ability, expression ability, and spatial distribution correlation are selected as the dominant characteristic factors for this embedding, and are used to form the two main axes of the plane coordinate system, such as the X-axis and the Y-axis, to realize the mapping basis of the two-dimensional integer lattice points.
[0050] Specifically, for the extraction of the leading factors, it is necessary to satisfy the existence of common fields, measurability, and structural continuity among objects; the leading factor combination is four or more (at least four) main factors extracted from the candidate factor set; the leading characteristic factors are two key factors preferentially selected as the structural embedding plane axes in the leading factor combination, with distinctiveness and distribution interpretability.
[0051] Under the condition of meeting the above requirements, the method of the present invention can map the object set to a two-dimensional integer lattice to achieve stable structure encoding, state classification, and subsequent graph embedding modeling.
[0052] To more clearly illustrate the concepts of "leading factor", "leading factor combination", and "leading characteristic factor", two typical practical application examples are listed below.
[0053] Example 1: Two-dimensional coordinate integer point set model. In the modeling of the two-dimensional integer point lattice structure, assume that an object set corresponds to all integer points in the first quadrant of the plane rectangular coordinate system, and the coordinates are denoted as (x, y), where (positive integer set). In this model, each object node is a two-dimensional coordinate point, which itself has the following leading factors: horizontal position factor (i.e., x value), vertical position factor (i.e., y value); parity of the abscissa, parity of the ordinate.
[0054] The above four factors can form a set of leading factor combinations. Among them, "x value" and "y value" can be used as leading characteristic factors to form a two-dimensional plane lattice; the "parity combination" is used for subsequent state encoding (four-color state classification). This model meets the plane embeddability determination conditions in the method of the present invention. Example 2: Modeling of the identity status of all human beings. In a global population identity dataset, each individual can be abstracted as an object node, and the following leading factor combinations can be extracted from each person: gender (male / female), age stage (juvenile / elderly); country or region of belonging; time period (such as timestamp, year of birth, registration time). Among them, if "gender" and "age stage" are set as the two leading characteristic factors for this mapping, the following two-dimensional plane can be formed: horizontal axis (X axis): male and female (gender leading factor), vertical axis (Y axis): juvenile, elderly (age leading factor).
[0055] Furthermore, each human individual can correspond to a unique position coordinate (x, y) in this plane lattice, and a status identifier is assigned accordingly to complete the processes of structure mapping, state classification, and identity traceability.
[0056] Distinguishability means that the selected dominant feature factors should be able to map different individuals in the object set to positions or states with significant differences, so that any two object nodes do not overlap or are easily distinguishable under the coordinate projection of the dominant factors, thus achieving the structural separation of objects. For example: The age factor distinguishes teenagers from the elderly, and the coordinate positions are distinguishable; or the parity of the X coordinate separates states.
[0057] Distribution interpretability means that the selected factors present a describable distribution pattern with certain regularity or semantic meaning in the overall data structure, so that the arrangement of objects on the two-dimensional structure conforms to cognitive, statistical or engineering logic, facilitating subsequent model processing, reasoning analysis and the maintenance of graph structure stability. For example: Taking "gender" and "age stage" as axes to form a plane, each region corresponds to different populations and has real-world interpretability; or the spatial sparse and dense distribution of nodes in the graph is consistent with the state division, constituting interpretability.
[0058] The summary of objects, dominant factors, combinations of dominant factors, dominant feature factors, distinguishability, and distribution interpretability is shown in Table 2.
[0059] Table 2 Definition Table of Terms in the Specification
[0060]
[0061]
[0062]
[0063] The typical determination examples listed in Table 1 show how to extract dominant factors from the object set and select dominant feature factors to construct a two-dimensional plane structure in different actual fields for subsequent structure mapping and state encoding.
[0064] The plane embeddability determination method of the present invention is applicable to object sets with significant attribute dimension characteristics, that is, in these sets, multiple dominant factors with distinguishability, semantic representativeness and statistical regularity can be extracted from each object, and two dominant feature factors are selected as the basic axes for plane embedding.
[0065] Such object sets are very common in actual engineering and information systems. For example, entity systems or graph models containing data such as personnel, geography, equipment, documents, and behaviors are widely present in scenarios such as smart cities, medical platforms, graph neural networks, content management, and user behavior modeling.
[0066] The plane embedding model constructed by the method of the present invention through the above-mentioned dominant factors can guide modelers to select the most representative feature pairs at the initial stage of modeling and complete the initialization of the mapping of the two-dimensional plane structure.
[0067] (2) Uniqueness
[0068] In the process of node mapping and state assignment in the two-dimensional integer point lattice graph (structured grid graph) constructed by the method of the present invention, the state uniqueness of each object node within the entire range of the two-dimensional integer point lattice graph is ensured. This uniqueness stems from: the two-dimensional integer coordinate mapping of nodes has no overlap, and each node occupies a unique lattice point; the initial four-color states (S1 to S4) generated in combination with the parity of coordinates have structural derivability; if there is a state conflict, the "sequence adjustment mechanism" is used to reposition the node to ensure that adjacent states are different; each node finally obtains a unique identifier for the entire two-dimensional integer point lattice graph under the combination of "coordinate + state".
[0069] Therefore, in the method of the present invention, the state of each node not only exists uniquely, but also is traceable, verifiable, and replicable. Without relying on an external ID table or a central system of a centralized database (the traditional data storage method where a single control node uniformly assigns and manages identifiers), a complete unique identifier can be established within the structure of the two-dimensional integer point lattice graph.
[0070] (3) Compressible state expression ability
[0071] The method of the present invention takes into account both the expression ability and the storage compression ability in the state identification method, forming a compressible state expression method applicable to high-density graph structures. Its compressibility is reflected in the following aspects:
[0072] Basic four-color state compression model: Four types of states (S1: odd-odd, S2: odd-even, S3: even-odd, S4: even-even) can be efficiently represented by a 2-bit (00, 01, 10, 11) coding method;
[0073] Extended state subclasses are still compressed for expression: Even when entering the second-layer state mapping (such as S1 being refined into S1-1 to S1-4), an extended bit (such as the 3rd bit) can still be used to represent the subclasses without increasing the overall state space bit width;
[0074] Support for tensorized expression: The encoded state can be directly used as a node attribute to construct a tensor input matrix for the model training of the graph neural network GNN;
[0075] Support for state rotation and modulo-N compression mechanism: In the case of further increasing the number of state types, a compression logic of modulo operation and sparse mapping can be introduced to maintain a constant storage structure.
[0076] According to the above, the method of the present invention not only ensures the uniqueness of node states in the two-dimensional integer point lattice graph structure, but also provides a state expression system with high information density + high representation efficiency through methods of bit width compression, tensor adaptation, and semantic nesting, which is applicable to the structure expression and recognition modeling of large-scale data object sets under limited computing resources.
[0077] Map the undirected graph constructed from the object set into the two-dimensional rectangular coordinate system in the first quadrant to form an integer point lattice structure. The mapping process is as follows: Select a node with significant features as the starting point, and set its coordinates as (1, 1). Then, scan the subsequent nodes in the order from left to right and from bottom to top, and embed them into the coordinate lattice points in sequence to form a stable planar structure.
[0078] The basis for selecting the significant features of the starting point includes: the node with the largest degree of connection in the graph, the node with the highest weight value, the node with the strongest centrality index (such as Betweenness), or the logical main index node.
[0079] The scanning methods include: default scanning, from left to right and from bottom to top; snake-shaped scanning, Z-shaped scanning, spiral filling; depth-first DFS / or breadth-first BFS; weight-guided, semantics-first, or chronological arrangement; local block division scanning.
[0080] The above diverse scanning methods can be flexibly selected in combination with the significant features of the graph.
[0081] III. Initialization of Four Types of States
[0082] According to the parity of the X and Y values of each node in the plane rectangular coordinate system, the node states are divided into four types: odd-odd (odd, odd), odd-even (odd, even), even-odd (even, odd), and even-even (even, even). As Figure 2 shown, odd-odd means that both the coordinate values of X and Y are odd, and the same applies to odd-even, even-odd, and even-even. In this embodiment, the points in the two-dimensional integer point lattice graph are labeled with states: State 1 (X and Y values are odd and odd), State 2 (X and Y values are odd and even), State 3 (X and Y values are even and odd), and State 4 (X and Y values are even and even). The four types of states can also be labeled as four-color states for coloring operations (dyeing).
[0083] IV. Adjacency Detection
[0084] For all pairs of points with connection relationships in the two-dimensional integer point lattice graph, perform state conflict detection on the adjacent node states being different in sequence to ensure that the state classification between any adjacent nodes remains different, meeting the basic constraint requirement of "adjacent node states are different".
[0085] (1) Precondition for State Classification
[0086] For each node that has been mapped to the two-dimensional integer coordinate system, its state is determined by the parity combination of its coordinate point (x, y), and is divided into four state categories, denoted as S1 to S4 respectively:
[0087] State S1 (odd-odd): x is odd and y is odd;
[0088] State S2 (odd - even): x is odd and y is even;
[0089] State S3 (even - odd): x is even and y is odd;
[0090] State S4 (even - even): x is even and y is even.
[0091] The node is mapped to a certain coordinate point. According to the above four state categories, the corresponding state category identifier is obtained for subsequent detection and classification.
[0092] (2) Detection of adjacent node pair conflicts
[0093] As Figure 3 shown, in this embodiment, four adjacent nodes are selected to form a closed - loop unit area, where: the coordinate of node A is (m, n), node B is the right - hand adjacent node of A, with the corresponding coordinate (m + 1, n), node C is the upper - hand adjacent node of B, with the corresponding coordinate (m + 1, n + 1), node D is the left - hand adjacent node of C, with the corresponding coordinate (m, n + 1), and node D and node A form the last side of the closed loop. These four nodes form a basic detection unit, and their adjacency relationships are in turn:
[0094] During the process of adjacent detection, the following steps are used to determine whether there are state conflicts between nodes:
[0095] (a) Coordinate acquisition and state determination. According to the method of state initialization of the present invention, the state of each node is determined by the parity combination of its coordinates (x, y), corresponding to the following four - color states: State S1: x is odd, y is odd (odd - odd);
[0096] State S2: x is odd, y is even (odd - even);
[0097] State S3: x is even, y is odd (even - odd);
[0098] State S4: x is even, y is even (even - even).
[0099] According to Figure 3 the coordinate values of the nodes in the example, assume: node A(m, n) → (identified as) state S1, node B(m + 1, n) → state S2; node C(m + 1, n + 1) → state S3; node D(m, n + 1) → state S4. Calculate the parity combination of the coordinates of each node in turn to obtain its corresponding state.
[0100] (b) Conflict detection process. According to the principle of "adjacent states must be different", for all node pairs with connection relationships in the two - dimensional integer point lattice structure, the state conflict detection is carried out in turn:
[0101] Detection The states of S1 and S2 are different and the states are legal;
[0102] Detection The states of S2 and S3 are different and the states are legal;
[0103] Detection The states of S3 and S4 are different and the states are legal;
[0104] Detection The states of S4 and S1 are different and the states are legal.
[0105] If the same state appears between any pair of adjacent nodes, it is recorded as a state conflict, and this pair of nodes is marked and put into the conflict queue for subsequent sequence adjustment processing.
[0106] (c) Description of the adjacent detection feature. Any pair of adjacent nodes is used as a detection unit. The detection unit is the smallest closed loop for judging adjacent conflicts in the structured grid diagram of the present invention and has the following characteristics: The detection is based on the states deduced from the odd and even coordinates, and does not depend on other attributes; it has nothing to do with the topological structure and only depends on the node coordinate values and the coordinate values of its adjacent nodes; it can be extended to any pair of nodes with edge connection relationships in the integer point lattice diagram; when adding new nodes, local and rapid detection can be performed without full graph traversal.
[0107] Through Figure 3 The closed-loop unit area shown, the method of the present invention can simultaneously perform state conflict detection (state legality) detection of the state differences between node pairs (adjacent nodes) in a batch and parallel manner in a two-dimensional integer point lattice diagram, enabling the computer system to stably achieve the state difference allocation between nodes in a large-scale two-dimensional integer point lattice diagram structure. Figure 3 The structural unit shown provides a repeatable and extensible state conflict detection and legality judgment step, constituting the basic judgment path and execution standard for the present invention to achieve the state difference allocation between adjacent nodes, ensuring the controllability and scalability of all node state allocations in the overall graph structure.
[0108] (3) Conflict determination
[0109] After the adjacent detection step is completed, determine the conflict:
[0110] The states between any two adjacent nodes are represented as Si and Sj. If Si≠Sj, the states between the two adjacent nodes are different, marked as legal states, and the connection relationship between the two adjacent nodes is retained;
[0111] If Si = Sj, it is determined that there is a state conflict between the two adjacent nodes, which does not meet the state difference of the method of the present invention, record the conflict information and enter the sequence adjustment step;
[0112] Record conflict information: All point pairs with conflicting status are recorded as the basis for the subsequent "sequence adjustment" method;
[0113] Continue traversal: traverse point by point until all connection points have completed the detection process.
[0114] Typical examples of state conflicts. The following are some common state conflict situations in actual detection and comparison with legal situations:
[0115] Conflict example:
[0116] Node A (3,3) is marked as state S1, node B (5,5) is marked as state S1, and recorded as conflict;
[0117] Node C(2,4) is marked as state S4, node D(4,2) is marked as state S4, and recorded as conflict;
[0118] Legal examples:
[0119] Node E(1,2) is marked as state S2, node F(2,1) is marked as state S3, and the record is legal;
[0120] Node G(1,1) is marked as state S1, node H(2,2) is marked as state S4, and the record is legal;
[0121] The above legal and conflict judgments are completely dependent on the "coordinate parity state" and do not depend on node attributes or graph topology complexity.
[0122] Adjacency detection has nothing to do with the structure of the structured grid graph. The detection process completely relies on node coordinates and state judgment, and has nothing to do with the graph topology or semantics.
[0123] Adjacency detection can be done all at once after graph mapping is completed, or it can be done locally when nodes are added, that is, local and incremental (batch and incremental execution are supported).
[0124] When a node is added dynamically, the legitimacy of the state can be confirmed by simply checking its adjacent edges (adapting to dynamic evolution).
[0125] The adjacency detection and the order adjustment method are closely connected (coupled), and the conflict records will be directly used as the basis for the "order adjustment" (queue), and can also be adjusted locally to ensure the legitimacy of the overall state of the structured grid graph.
[0126] 5. Adjustment of ranking
[0127] During the state conflict detection process, when a state conflict occurs (a state conflict appears between adjacent nodes), that is, the four-color states of two adjacent nodes are the same, a position adjustment is performed. The subsequent nodes are re-searched for mapping positions in the order of right shift, up shift, and diagonal shift (right shift → up shift → diagonal shift) until the states of all adjacent nodes are different from each other. The node is embedded in a four-color state point that does not conflict with its adjacent nodes, that is, the states of other nodes connected to this node are different from the state of this node. In this way, the object point set in the two-dimensional integer point lattice diagram is divided into four categories.
[0128] Then, in the specified large category (state) in turn, the position adjustment step is performed again.
[0129] The method of "right shift → up shift → diagonal shift" for position adjustment:
[0130] Right shift: Try to map the node to the grid on the right side of the current coordinate, that is, from (x, y) → (x + 1, y);
[0131] Up shift: If there is still a conflict after the right shift, try to move one grid up from the original point, that is, from (x, y) → (x, y + 1);
[0132] Diagonal shift: If there is still a conflict after the up shift, try to move diagonally from the original point, that is, from (x, y) → (x + 1, y + 1);
[0133] If there is still a conflict after the above three moves, on the basis of the original point, continue to recursively adjust the next-level coordinate points in the above order, that is, in the order of "right shift → up shift → diagonal shift", increasing one coordinate value until a position that satisfies the condition that the states of all adjacent nodes are different from each other is found (repeated recursion).
[0134] Specific illustration: Suppose node A was originally intended to be mapped to (3, 2), but its adjacent node B has the same state as the state at this position, for example, state S2, resulting in a conflict: In the first step, try to right shift → coordinate (4, 2) and determine the state; if there is still a conflict, in the second step, try to up shift → coordinate (3, 3); if there is a conflict again, in the third step, try to diagonal shift → coordinate (4, 3). If the above positions all conflict with the states of adjacent nodes, continue to perform the next round of position adjustment attempts, starting from positions such as (5, 2), (3, 4), (5, 4)... until a position is found whose state is different from the states of all connected adjacent nodes, and then the adjustment can be completed.
[0135] As Figure 6 shown, from left to right and from bottom to top, in order, the adjustment is as follows:
[0136] Fix B11 to align with A11. From left to right, it is found that B51 conflicts with B11, so adjust B51 → B41 (A41)
[0137] Proceed to the next step. It is found that B14 conflicts with B32, so adjust B14 → B24(A24).
[0138] Proceed to the next step. It is found that B55 conflicts with B75, so adjust B55 → B54(A54).
[0139] Proceed to the next step. It is found that B36 conflicts with B54, so adjust B36 → B45(A45).
[0140] The adjustment is completed.
[0141] When the method of the present invention is implemented using a computer system, as Figure 4 shown, when a certain node is at the initial embedding position, such as coordinate A(m,n), and there is a state conflict with its adjacent nodes, the computer system will automatically adjust the position according to the "sequence adjustment". The adjustment steps are attempted in the following order of priority
[0142] First, shift one grid to the right (m+1,n);
[0143] If there is still a conflict, shift up (m,n+1);
[0144] If there is still a conflict, shift diagonally (m+1,n+1);
[0145] If all of the above three cases result in a conflict, continue to iterate to the next-level sequence position, such as m+2,n, m+1,n+2, until the node is mapped to a target position where the states of all adjacent nodes are different.
[0146] Each "four-color state" can be regarded as the first-layer state space. If further classification is required when the state density is too high, a second-layer sorting can be performed, and the current states are grouped by parity again to form a "second-layer state mapping". For example: classify the first layer into the second layer, and the second layer is divided into even and odd layers, and then scan odd and odd, odd and even in sequence within each layer.
[0147] Regard the "four-color states" (odd-odd, odd-even, even-odd, even-even) as the first-layer state space, and each node's initial state determination belongs to one of them.
[0148] When the number of nodes is large or the density within a certain state category is too high, resulting in frequent or failed sequence adjustments, the second-layer state mapping method can be entered to further subdivide the point set in the existing state classes.
[0149] If a node is determined to be in state S3 (even-odd): within the even-odd class, it can be further subdivided into four subclasses according to the parity of the X-axis + the parity of the Y-axis: even-odd - odd-odd, that is, x is even and y is odd, labeled as: secondary odd, even-odd - odd-even, even-odd - even-odd, even-odd - even-even; each subclass forms a second-layer sub-state space, and the sequence adjustment steps are performed again within this sub-space.
[0150] After the above-mentioned sequence adjustment and status conflict detection, if there are still status conflicts, it further enters the third-layer status classification step. The specific steps are as follows: on the basis of the original status, introduce additional feature dimensions, such as time tags, semantic categories, as status perturbation factors, then refine the status categories, and finally expand the status space to further resolve local conflicts.
[0151] Since the object point set is a finite set, the number of status conflicts is also finite, and the sequence adjustment and multi-layer status classification steps can be completed within a finite number of steps. Finally, each node in the undirected graph can be mapped to a unique coordinate point without status conflicts and is assigned a unique status identifier.
[0152] For any object node (designated point) in the undirected graph structure, after completing the status mapping by the method of the present invention, the position and status encoding of this point can be accurately located from the entire structure. This status encoding not only has uniqueness, but also carries the adjacency information and hierarchical path of this node in the entire structure, thereby obtaining all the structural features and semantic context of the object represented by this node.
[0153] Therefore, the method of the present invention not only realizes the global legal allocation of the status of all nodes, but also enables any node (designated point) to have a locatable, traceable, and verifiable status in the two-dimensional integer point lattice graph, meeting the requirements of accurate identification and tracking of individuals in the big data environment.
[0154] If the status category where the node is located has entered the second-layer status mapping and there are still status conflicts, it can further enter the "third-layer status mapping step". On the basis of the original coordinate parity classification, add additional dimensions (such as time tags, semantic tags, usage frequency, etc.) as "status perturbation factors" to make a refined subclass distinction of the status, so as to decouple the conflict-intensive areas and improve the status mutual exclusivity ability.
[0155] The third-layer status mapping step is a method for further dividing the node status using additional dimension information (such as time, semantics, tags, etc.) in the case of still existing conflicts after the first-layer (four-color status) and second-layer (coordinate parity subdivision) status classifications, which is used to expand the status space and solve the status conflicts in the high-density area.
[0156] The status perturbation factor is an additional status influence parameter used to guide the third-layer status mapping, including: time stamp, semantic category, affiliated tag, behavior frequency, which participate in the status classification as auxiliary features in the status dimension.
[0157] Finally, all the objects in the original point set have completed the mapping detection, obtaining a unique state identifier for conflict avoidance, which is the specific location of the specified point we want to obtain, and all the features of the specified point sample are obtained accordingly. By adjusting the order of the method of the present invention, it is ensured that the states between all adjacent nodes are different from each other, and finally a conflict-free state is completed.
[0158] Now, taking "all human individuals" as the object set, it is specifically described as follows:
[0159] The first-layer state mapping (four-color state). In this step, each object is initially used as the main axis with two dominant feature factors, gender (male / female) and age stage (juvenile / elderly), to construct a plane coordinate system, forming 4 basic state categories:
[0160] Age Male Female Teenager S1 S2 Elderly S3 S4
[0161] Corresponding to the four-color state division method:
[0162] S1: male + juvenile (odd-odd);
[0163] S2: female + juvenile (odd-even);
[0164] S3: male + elderly (even-odd);
[0165] S4: female + elderly (even-even).
[0166] The state classification of this layer constitutes the first-layer state mapping, which can be used to quickly realize the initial structure encoding.
[0167] The second-layer state mapping (subclass subdivision). When the density of objects in a certain state category (such as S3: male + elderly) is too high, resulting in state conflicts between nodes, more refined age information can be introduced for the second-layer classification. For example:
[0168] S3-1: 65 years old;
[0169] S3-2: 66 years old;
[0170] S3-3: 67 years old;
[0171] S3-4: 68 years old.
[0172] The "age value" here constitutes the second-layer sub-state mapping dimension under this state category. In the method of the present invention, the state is recorded in the form of a sequence number or an extended bit to maintain the overall encoding method.
[0173] Third - layer state mapping (perturbation factor guidance). If the state refinement in the second layer is still insufficient to completely resolve local state conflicts (for example, the individuals of 65 - year - old males are extremely dense in a certain city), it can further enter the "state perturbation factor" step. For example: geographical tags (residence), behavior tags (health level, occupation type), time tags (birth year or registration time). For example:
[0174] S3 - 1 - A: 65 years old, living in Beijing;
[0175] S3 - 1 - B: 65 years old, living in Guangzhou;
[0176] S3 - 1 - C: 65 years old, living in Shenzhen;
[0177] S3 - 1 - D: 65 years old, living in Shanghai.
[0178] This perturbation factor guides the state space to expand again, forming a third - layer state mapping structure, thereby further reducing the conflict probability of high - density nodes and ensuring the uniqueness and legality of the final node state.
[0179] The above example constructs the first - layer four - color state through "gender × age stage", refines the second layer by combining age values, and then introduces additional semantic perturbation factors to form the third layer, which can achieve high - precision, conflict - free structured state encoding and unique identity modeling for the entire human sample set, and has strong scalability, interpretability and application versatility.
[0180] The method of the present invention adopts the order adjustment and multi - layer state mapping method, embeds the original graph structure completely into the two - dimensional integer lattice space, and realizes the uniqueness and adjacency legality of all node states. This state identifier essentially constitutes a unique positioning flag for each object, and this method can be used as the input of neural network encoding, distributed search label, and clustering starting point.
[0181] The method of the present invention depends only on its adjacent nodes for the state value and detection of each node, without global data. The order adjustment is only carried out in local conflict pairs, and is suitable for distributed, parallel computing and edge deployment using a computer. The addition of new nodes and the removal of old nodes can be quickly reconstructed in the local node space, with the ability of dynamic evolution; the entire mapping process can be deployed in environments such as federated learning, the front - end of graph neural networks, and edge AI terminals, with good portability and scalability.
[0182] VI. Dynamic Evolution
[0183] For the newly added nodes in the undirected graph structure, according to the above steps two to five, the mapping method enters the two - dimensional integer point lattice graph, initializes the four types of states, performs adjacency detection, and if there is a state conflict, local order adjustment is carried out. For example Figure 7As shown, a new dynamic point B96 is added, and its adjacent connection lines are connected to points B98, B75, and B72. Since B96 conflicts with both points B78 and B72, point B96 is adjusted to B86, and the adjustment is completed. Node B96 is deleted, and the status of this node is released, without the need for global update of the two-dimensional integer point lattice graph.
[0184] Method for adding a new node: When a new node appears in an undirected graph structure, the node can be embedded into the two-dimensional integer point lattice graph according to the steps in Step 2. The specific steps are as follows:
[0185] Preselect coordinates: Based on the undirected graph structure and the scanning method, select an initial mapping position, such as at an adjacent node of a node; Status assignment: According to the parity of the selected coordinate point, preliminarily determine that the status of this node is one of S1 to S4; Adjacency detection: Detect whether there is a status conflict between this node and all its adjacent nodes; Perform local ranking adjustment. If there is no conflict, the mapping is completed. If there is a conflict, only perform ranking adjustment in the local graph where this node and the adjacent nodes are located; Other non-adjacent areas are not affected and there is no need for global re-mapping. If the local state space density is too high, a second-layer sub-state classification can be introduced, and according to the method of entering the second-layer state mapping, further ensure that the states are different from each other. The method of entering the second-layer state mapping allows local embedding of new nodes without global mapping reconstruction, and this method has extremely high scalability and distributed capabilities.
[0186] Node deletion method: When a certain node in the undirected graph structure is deleted, its corresponding node position in the two-dimensional integer point lattice graph is released, and status recycling is performed: The status identifier corresponding to this coordinate point is cleared, the connection relationship is disconnected, and the relevant connection edges are removed together, without global update: Since the status identifier is unique, deleting it will not affect the status of other nodes, and there is no need to perform any position migration or status change on the remaining structure.
[0187] In the method of the present invention, the status is divided into four categories, (odd-odd, odd-even, even-odd, even-even) are the basic statuses. When the node density increases or the status space is crowded, status compression and status expansion can be performed to adapt to undirected graph structures with different densities, scales, and semantic requirements.
[0188] The status compression method is: Represent the four categories of statuses (S1 to S4) using a binary 2-bit coding method (00, 01, 10, 11). In application scenarios such as graph neural networks, graph databases, and AI model inputs, this method can be directly used as a node attribute and input into a computer, supporting tensor processing and parallel computing, which is beneficial to improving the storage and computing efficiency of large-scale graph models.
[0189] When the node density significantly increases or the status space is crowded, the status can be expanded to five categories, six categories, or more categories to adapt to higher density structures.
[0190] The state expansion is to expand or further refine the existing state. The state expansion methods include the following categories: two - layer parity classification, such as further dividing S2 into S2 - 1 to S2 - 4 in S2; adding semantic or label dimensions, such as weight levels, time - period identifiers; introducing modulo - N state rotation methods, for example, taking the modulo of the state number. The expanded state still maintains unique identification and adjacency differentiation, which is beneficial for processing large - scale complex undirected graph structures.
[0191] The method of the present invention can be applied to finding a person or an object in a planar map, finding a specific document in a large - scale homogeneous document library with equal identity features, and finding individuals in other instance point sets with planar features.
[0192] The method of the present invention is applied to a computer. All instance point set scenarios with planar features or two - dimensional structures have distributed execution capabilities, support automatic state adjustment, and have high scalability and embeddability.
[0193] The method of the present invention is not only applicable to abstract graph structure modeling, but also widely applicable to actual object sets with two - dimensional embedding features, including: locating people or objects in a planar map, finding specific document individuals in a file system or document library, identifying key nodes or isolated points in a social network, encoding distribution modeling before image annotation and object recognition, urban - level space deployment planning, such as parking spaces, electricity meters, equipment points, equipment scheduling and task matching models in an intelligent manufacturing system, and quickly clustering and locating individual instances in a distributed AI search system.
[0194] Graph Neural Network (GNN) has been widely applied in many fields, such as: computer vision, social network analysis, bioinformatics modeling, and intelligent manufacturing system optimization. Using a computer, the method of the present invention can be widely applied in the GNN model as an input encoding method for node structure features, providing a stable state vector representation, thereby improving the training efficiency, structure perception ability, and result interpretability of the graph model. As Figure 8 shown, the steps are as follows:
[0195] (1) Construct the original graph structure. For a structured data set to be modeled, such as the relationships between objects in a computer image and a protein - protein interaction network, represent it as an undirected graph G(V, E) in a computer through a graph data modeling tool, such as the DGL or PyG graph neural network framework, where V is the set of nodes and E is the set of edges ( Figure 8 as shown in the left - hand part).
[0196] (2) Perform structure mapping and coordinate encoding. Apply the embedding method of the present invention to map the undirected graph G(V, E) to a two - dimensional planar integer lattice structure. Each node is embedded into an integer coordinate (x, y) in the first quadrant according to the structure rules (Figure 8 the middle part).
[0197] (3) Generate node state vectors. According to the parity combination of node coordinate points (such as x odd / y even), generate state classification codes for each node. This code can be in the form of a binary vector (such as S1 = 00, S2 = 01,...), and form the initial state vector of the node ( Figure 8 The right figure is the schematic diagram of the node state vector).
[0198] (4) Construct the input tensor structure. All node state vectors are arranged in sequence according to the node numbers (or the graph topological order) to form a multi-dimensional input tensor for subsequent input to the neural network. This tensor structure can be compressed for storage and has structural interpretability, supporting visualization and debugging.
[0199] (5) Input it into the GNN model training system. The generated node state tensor is used as the node attribute matrix and input into the graph neural network model, such as GCN, GAT, or GraphSAGE, to participate in training and inference as part of the node features.
[0200] Since the method of the present invention provides a state encoding method with different states between nodes and controllable structure, the GNN model can obtain a clearer structure-aware input during the training process, which helps to improve the convergence speed, classification performance, and interpretability of the model.
[0201] The method of the present invention has the following advantages:
[0202] (1) Distributed characteristics: All node mapping, detection, and adjustment can be independently completed locally,
[0203] (2) State sequence adjustment combined with local detection, which is applied in the computer to support dynamic adaptability,
[0204] (3) High scalability, density dynamic expansion, and multi-level expansion of the state space,
[0205] (4) Strong embedding ability, suitable for all instance point sets with two-dimensional embedding or planar features,
[0206] (5) Applied in the computer, with good computational visualization, and the structure, position, and state can all be intuitively presented, suitable for interactive computer systems and embedded AI terminals.
[0207] The two-dimensional integer point lattice graph state mapping method of the method of the present invention has the ability to accurately backtrack and locate (backtrack and identify) and quickly identify individual objects. This ability is based on unique coordinate embedding and multi-layer state encoding, allowing for the rapid reverse positioning of the mapping state and feature information of specific individuals from a large set of objects without global search. Specifically:
[0208] After the method completes the mapping of the undirected graph structure to the first quadrant of the plane rectangular coordinate system, forming a two-dimensional integer point lattice graph and state assignment, each object node obtains a globally unique coordinate position (x, y), and corresponds to a determined state code or extended state set, and the mathematical expression is: S ∈ {S1, S2, S3, S4}. At the same time, this coordinate position can be associated with the following structural information: node ID and label, the set of object nodes adjacent to it, the state class and subclass to which it belongs, and one-to-one correspondence with the vector position in the AI model structure tensor.
[0209] Therefore, by backtracking the following two information paths, accurate search and semantic recognition of the target individual can be achieved: Path 1: Reverse query the coordinate points from the state tensor, read the state vector of the target object (such as the GNN output node encoding), reverse solve the corresponding coordinate position, and obtain its object ID and associated context in the original dataset. Path 2: Infer the state label from the relationships in the graph structure, confirm the candidate object through clues such as adjacent nodes and upper and lower objects in the relationship graph, call the state position of the object in the embedded graph, and perform an exact match to quickly determine whether the object is the target individual. As shown in Table 3.
[0210] Table 3 Examples of Scenarios, Backtracking Targets, and Backtracking Methods
[0211]
[0212]
[0213] Backtracking recognition has the following advantages:
[0214] (1) Unique positioning: Coordinate + state dual identification ensures that each object can be uniquely found in space.
[0215] (2) Does not rely on global traversal: With the help of the encoding structure, the individual position can be directly reverse solved, with high efficiency.
[0216] (3) Supports AI integration: Can be directly aligned with the output results of the neural network, facilitating inference tracking.
[0217] (4) Strong structure visualization: The embedded graph and state structure can be graphically rendered, enhancing the interactive experience.
[0218] (5) Can be distributed: The backtracking of each node can be independently executed on the local or edge device.
[0219] As Figure 6 shown, a standard backtracking recognition method includes the following steps:
[0220] (1) Specify the object label or partial features.
[0221] (2) Retrieval status code or node context information;
[0222] (3) The corresponding coordinate position is reverse-searched;
[0223] (4) Obtain all states, adjacency relationships, and AI model vectors of the node;
[0224] (5) Judge the matching result.
[0225] If the match is successful (the state, context, label, and target features are exactly the same), the individual backtracking recognition is completed, and the unique identifier, state vector, and context information of the node are returned;
[0226] If the match is not successful, that is, there are situations such as fuzzy state, label matching failure, or coordinate offset, then the following steps are carried out:
[0227] a. Enter the fuzzy matching step, enter the feature tolerance matching or semantic extension method, and allow within a certain range, such as coordinate neighborhood, adjacent subclasses of the state, fuzzy label matching, to expand the retrieval range and re-identify possible target nodes.
[0228] b. Fall back to the upper-level state category for re-screening. If there is no effective match for the second-layer or third-layer state codes, fall back to the first-layer state category to which the node belongs, and reconstruct the candidate node set based on the dominant features.
[0229] c. Through the adjacency structure path comparison, if the candidate node states are similar but not exactly the same, compare their adjacency structures, such as who to connect to, the length of the connection path, and perform a structural-level matching auxiliary judgment with the target feature context.
[0230] d. Mark as "node to be confirmed" or "fuzzy matching result". If it still cannot be fully confirmed, mark the result as "candidate fuzzy node" or "pending manual verification" for further processing in the subsequent verification link.
[0231] On the basis of completing the backtracking recognition method, it is further extended to the extraction of all features (complete features) of the specified point sample, including the following steps:
[0232] (1) Obtain the embedded coordinates and state vector. According to the aforementioned backtracking path, locate the embedded coordinates (x, y) of the node in the two-dimensional plane integer point lattice diagram, and at the same time obtain the state (state classification code) of the node, such as the four-color states S1~S4 or their extended sub-states, and combine them to form the complete state vector of the node for structural positioning and semantic identification.
[0233] (2) Identify the state category and subclass information to which the node belongs. Based on the state vector of this node, determine its main state category (S1 - S4) and its subordinate subclasses (such as S1-1, S3-2), and combine perturbation factors (such as age group, city, time) to restore its classification path and semantic label, so as to obtain the position hierarchy and attribution of this node in the entire state system.
[0234] (3) Extract the adjacency structure information and path structure features, and further obtain all the direct adjacent nodes of this node in the two-dimensional integer point lattice graph structure, that is, all edge connection relationships, and calculate the topological structure information of the path length and connection depth between it and other key nodes (such as the central node, label node, target node). The topological structure information is used to judge the local density, propagation, and centrality two-dimensional integer point lattice graph characteristic attributes of this node in the graph.
[0235] (4) Generate the structural feature vector required for model training. The above state vector, adjacency relationship, path information, and the attribute data of the node itself (such as label, numerical feature) are combined to generate a complete input vector, which serves as the training feature dimension of this node in the AI model. This vector can be directly used as the training input of the graph neural network GNN model, constituting the basic attribute tensor of each node in the neural network, and supporting the execution of intelligent tasks such as two-dimensional integer point lattice graph classification, node prediction, clustering, or path recommendation.
[0236] The above process provides methods for object analysis, AI training interpretation, traceability analysis, etc. in a complex big data environment.
[0237] In the method of the present invention, "each node" is used to indicate that all nodes in the entire object set can be mapped and assigned a unique state; while "individual" and "unique individual" are used to indicate any target node (object) specified by the user, which can be accurately located and identified from the planar structure by the method of the present invention.
[0238] The method of the present invention has the advantages of efficient encoding, conflict avoidance, unique state, dynamic evolution, and individual traceability, and is applicable to multiple big data application scenarios with two-dimensional embedding structure or planar modeling characteristics. It can be applied to multiple types of computing environments, including local computing, distributed computing, edge deployment, and AI training platforms. As shown in Table 3.
[0239] Table 3 Typical application scenario examples of the method of the present invention
[0240]
[0241]
[0242] The method of the present invention is also adapted to be applied to the following several environments:
[0243] (1) Local deployment: Suitable for small-scale graph or local sub-graph tasks, it can run on a single machine and / or memory-based services.
[0244] (2) Edge computing deployment: Such as edge AI boxes for cameras, in-vehicle terminals, and hospital diagnostic assistance terminals.
[0245] (3) Distributed cluster deployment: Used for large-scale urban graphs, knowledge graph structures, or training systems.
[0246] (4) Integration with AI frameworks: It can be used as a GNN input module, feature compression module, and embedding layer preprocessing module. The GNN input module, feature compression module, and embedding layer preprocessing module are provided with: a graph structure preprocessing module, a coordinate embedding and state classification module, an adjacency detection and sequence adjustment module, a state tensor generation and output module, an individual backtracking and feature extraction module, and a dynamic update and distributed execution management module.
[0247] As a pre-judgment method, the method of the present invention can be applied to the judgment of large data sets. Four dominant factors can be extracted from a certain data set, and any two objects among them can form a data plane with structural distribution regularity.
[0248] Examples of typical object combinations:
[0249] (1) Urban parking spaces: longitude, latitude, frequency, category.
[0250] (2) Population data: gender, age, occupation, residential area.
[0251] (3) Medical information: tongue image features, medical history descriptions, body temperature changes, time stamps.
[0252] (4) File system: creation time, directory depth, content keywords, reference structure.
[0253] The method of the present invention is based on graph structure mapping, parity state classification, and sequence adjustment, and is used for distributed search and precise positioning of individuals in large data sets. This method maps the original object set structure to a two-dimensional first quadrant integer point lattice graph, and constructs a globally unique and locally adaptive coding method system through four-color state division and local adjustment. Compared with traditional methods, the present invention has the following innovation points and advantages: unique and traceable state: each individual object has a unique coordinate and state, supporting precise backtracking; no global adjustment is required for adding and deleting nodes; flexible state compression and expansion: it can be applied to 2-bit compression representation and can also be expanded into multiple types of states; strong distributed deployment ability: supporting edge AI deployment and automatic mapping of massive data; strong versatility: applicable to multiple fields such as matching plane maps, image modeling, graph neural networks, and file systems; the method of the present invention will make a positive contribution to the next-generation AI structure modeling, large-scale graph mining, and intelligent object positioning.
[0254] The significance of the method of the present invention lies in: constructing the maximum structural clarity with the minimum expression; achieving the largest-scale distribution mapping with the most concise state combination; supporting the evolution of the overall dynamic order with local automatic adjustment; using graph theory and mathematical limit methods to respond to the future structural dilemmas of AI.
[0255] The method of the present invention is not only applicable to the fields of graph data, AI models, structure recognition, and memory-computation integrated chips, but also a cognitive modeling method with generalizability, providing underlying guidance for the structural generation logic, state evolution order, and distributed collaborative recognition mechanism of future intelligent systems.
[0256] The computer system constructed by using the method of the present invention can be set on a personal computer (PC). The PC is connected to a database system, and the PC can also be connected to other intelligent terminals. The PCs and intelligent terminals within the network can respectively or simultaneously implement the method of the present invention.
Claims
1. A method for distributed searching of individuals in a large dataset, which is used for a set of two-dimensional data objects of limited scale, and includes the following steps: I. Establish an undirected graph In a set of objects with two-dimensional plane characteristics, each object is regarded as an original point, and an undirected graph structure is constructed accordingly. In this undirected graph, each node corresponds to an original point object, and the connection lines between nodes indicate that there is a direct adjacent relationship or semantic association relationship between two objects. The edges in the undirected graph satisfy the constraints of non-crossing and non-overlapping. II. Structure mapping Map the undirected graph structure into the first quadrant of the plane rectangular coordinate system to form a two-dimensional integer point lattice graph. III. Initialization of four types of states According to the parity of the X and Y values of each node in the plane rectangular coordinate system, the node states are divided into four categories: State 1: X and Y values are odd and odd; State 2: X and Y values are odd and even; State 3: X and Y values are even and odd; State 4: X and Y values are even and even. IV. Adjacency detection For all pairs of points with connection relationships in the two-dimensional integer point lattice graph, perform state conflict detection for adjacent node states to be different, so as to meet the requirement that adjacent node states are different. V. Sequence adjustment During the state conflict detection process, when there is a state conflict between adjacent nodes, perform sequence adjustment, and re-search for the mapping position of the subsequent nodes in the order of right shift, up shift, and diagonal shift until all adjacent node states are different.
2. The method for distributedly searching for individuals in a large dataset according to claim 1, wherein: When a new node is added to the undirected graph structure, it enters the two-dimensional integer point lattice graph according to the mapping method in steps II to V, initializes the four types of states, and performs adjacency detection. If there is a state conflict, perform local sequence adjustment.
3. The method for distributedly searching for individuals in a large dataset according to claim 1, wherein: When a node in the undirected graph structure is deleted, the corresponding node position in the two-dimensional integer point lattice graph is released, the state identifier corresponding to the coordinate point is cleared, and the relevant connection edges are removed together.
4. The method for distributedly finding individuals in a large dataset according to claim 1, characterized in that: When the node density increases or the state space is crowded, perform state expansion, and the state expansion includes: two-layer parity subclassification, adding semantic and label dimensions, and introducing the modulo N state rotation method.
5. The method for distributedly searching for individuals in a large dataset according to claim 1, characterized in that: When the node density increases or the state space is crowded, perform state compression, and represent the four states by the binary 2-bit coding method.
6. The method for distributedly finding individuals in a large dataset according to claim 1, characterized in that: The direct adjacent relationship in step I means that there is a clear structural or spatial proximity between objects; the semantic association relationship means the logical connection established between objects based on meaning, structure, or context.
7. The method for distributedly finding individuals in a large dataset according to claim 1, characterized in that: In step II, an object with significant characteristics is set as the starting point (A11) and mapped to the coordinate value (1, 1), and then other nodes are embedded into the two-dimensional integer grid points one by one along the X and Y axes of the plane rectangular coordinate system in the scanning manner from left to right and from bottom to top.
8. The method for distributedly finding individuals in a large dataset according to claim 1, characterized in that: Before mapping the undirected graph structure into the first quadrant of the plane rectangular coordinate system in step II, perform a plane embeddability determination, and the condition is: if each object in the object set has four representative basic attributes or description dimensions, and any two of them form a basic attribute or description dimension pair, it meets the plane embeddability condition.
9. The method for distributedly searching for individuals in a large dataset according to claim 1, characterized in that: In step IV, perform state conflict detection for adjacent node states to be different in sequence.
10. The method for distributedly searching for individuals in a large dataset according to claim 1, characterized in that: In the fourth step, state conflict detection is performed for adjacent nodes with different states, and at the same time, state detection between adjacent nodes is carried out in batches and in parallel.
Citation Information
Cited By
Super-computing-oriented high-expandability grid generation method and device, equipment and medium
CN120747418A