Personnel information big data screening method and system based on person and certificate identification
By detecting index structure adjustment events, analyzing the characteristics of spatiotemporal continuity breaks, selecting a parallel index construction mode, executing query tasks in parallel, and combining topological persistent homology analysis for partition load balancing, we solve the problems of query accuracy, response speed, and data consistency in dynamic data environments, and achieve zero-deviation query results in high-frequency data update scenarios.
Patent Information
- Application Number
- CN202511141118.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing technologies find it difficult to balance the query accuracy, response speed and data consistency of real-time screening in a dynamic data environment. Especially in scenarios with billions of data and high-frequency list updates, lightweight index structures reduce the accuracy of fuzzy retrieval, while high-precision index maintenance operations will block the query process or cause delayed results.
By detecting index structure adjustment events, analyzing the characteristics of spatiotemporal continuity breaks, selecting a parallel index construction mode, executing query tasks in parallel, and combining topological persistent homology analysis for partition load balancing, high-dimensional vectors are divided into hotspot subsets and non-hotspot subsets, and stored in independent memory to ensure the consistency of query results.
It achieves a balance between retrieval accuracy, response speed and data consistency in high-frequency data update scenarios. By detecting index structure adjustment events in real time and intelligently selecting the parallel index construction mode, it ensures the continuous execution of query tasks and zero deviation in query results.
Smart Images

Figure CN120743907A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of high-dimensional vector real-time query, and in particular to a personnel information big data screening method and system based on human evidence recognition. Background Art
[0002] In personnel security management based on identity verification, it is usually necessary to obtain personnel identification information through document recognition technology (such as ID card OCR) or biometric features (such as face), and perform real-time screening and matching in the massive background database. The database contains dynamically updated list data (such as key personnel database, behavioral feature vector database), and the screening process must meet the response requirements in seconds; existing technologies generally use approximate nearest neighbor search (ANN) algorithms to process high-dimensional vector similarity queries (such as behavioral association analysis based on facial features), or accelerate the accurate retrieval of structured identifiers (such as ID card numbers) through index optimization. However, when faced with data scale of billions and high-frequency list updates (such as real-time new risk personnel records), the system needs to continuously maintain an efficient query index structure.
[0003] Current technologies make it difficult to balance the query accuracy, response speed, and data consistency of real-time screening in a dynamic data environment. Specifically, the lightweight index structure adopted to meet real-time requirements will reduce the accuracy of fuzzy retrieval, while high-precision index maintenance operations (such as reconstruction and incremental updates) will block the query process or cause delayed results; at the same time, high-frequency data updates further amplify index maintenance overhead, resulting in mutual constraints among retrieval accuracy, system latency, and data freshness. This contradiction is particularly prominent in scenarios where it is necessary to continuously update a list library of billions and perform millisecond-level responses. Summary of the Invention
[0004] The present invention aims to solve the technical problems existing in the prior art and provides a method and system for screening big data of personnel information based on human-document identification.
[0005] The technical solution of the present invention to solve the above technical problems is as follows: The personnel information big data screening method based on human evidence recognition includes: S1. Receive structured personnel identification information and detect the real-time update operation type of the background database; S2. When the real-time update operation type is index structure adjustment, detect the spatiotemporal continuity break characteristics of the new data and historical data, and resolve the conflicting nodes in the index maintenance operation chain; S3. Select a parallel construction mode for the temporary index structure based on the spatiotemporal continuity break characteristics and the correlation strength of the conflicting nodes, and execute it in coordination with the current query task; S4. Perform spatial distortion evaluation on the high-dimensional vector distribution of the temporary index structure and calculate the vector neighbor relationship conflict ratio; S5. Determine the computational granularity of the topological persistent homology analysis based on the neighbor relationship conflict ratio. Perform partition load balancing on the temporary index structure based on the extracted key persistence features of the vector space, partition the high-dimensional vector into hot and non-hot subsets, and allocate independent memory storage nodes to the hot subsets. S6. Generate a verification sample set based on the structured personnel identification information, switch the query request to the temporary index structure to perform the verification query, and confirm the consistency with the original index result.
[0006] Furthermore, the step S1 includes: Obtain structured person identification information from the document image; Extract biometric data and associate it with structured person identification information; Transmit the associated structured personnel identification information to the update interface of the backend database; Monitor the real-time update operation log of the backend database and analyze the operation instruction types in the log; The real-time update operation type is marked as data content update or index structure adjustment according to the operation instruction type.
[0007] Furthermore, the step S2 includes: Extract the timestamp sequence and geographic location coordinate sequence of the newly added data; Extract the timestamp sequence and geographic location coordinate sequence of historical data; Calculate the sliding window standard deviation of the timestamp sequence of the newly added data and the timestamp sequence of the historical data; Calculate the center offset of the sliding window between the newly added data geographic location coordinate sequence and the historical data geographic location coordinate sequence; When the sliding window standard deviation exceeds the temporal continuity threshold or the sliding window center offset exceeds the spatial continuity threshold, it is marked as a spatiotemporal continuity break feature; Capture thread lock wait events in the index maintenance operation chain; The operation node that holds the thread lock for longer than the lock wait threshold is marked as a conflict node.
[0008] Furthermore, step S3 includes: The Pearson correlation coefficient between the numerical strength of the spatiotemporal continuity break feature and the numerical strength of the conflict nodes was calculated as the association strength; When the correlation strength is greater than the construction mode switching threshold, global shard construction is selected as the parallel construction mode of the temporary index structure; When the correlation strength is less than or equal to the build mode switching threshold, local incremental build is selected as the parallel build mode of the temporary index structure; Assign an independent query thread to the current query task; Allocate independent build threads for the selected parallel build mode; The query thread and the construction thread are executed alternately through the operating system kernel scheduler.
[0009] Furthermore, the step S4 includes: Get the original position coordinates and current position coordinates of each high-dimensional vector in the temporary index structure; Calculate the Euclidean distance between the original position coordinates and the current position coordinates of each high-dimensional vector as the position offset; When the position offset exceeds the spatial distortion threshold, the corresponding high-dimensional vector is marked as a distortion point; For each distorted point, find its neighbor relationship set in the original index; For each distorted point, find its neighbor relationship set in the temporary index structure; Compare the neighbor relationship sets of the distorted points in the original index and the temporary index structure; The ratio of the number of distorted points with inconsistent neighbor relationship sets to the total number of distorted points is counted as the vector neighbor relationship conflict ratio.
[0010] Furthermore, the step S5 includes: According to the neighbor relationship conflict ratio value, the preset granularity mapping table is queried to obtain the maximum edge length parameter of the Vietoris-Rips complex construction; The maximum edge length parameter is used to construct the Vietoris-Rips complex for the high-dimensional vector set in the temporary index structure; Compute the persistent homology group corresponding to the Vietoris-Rips complex and generate the corresponding persistent barcode; Extract the ring structure whose life cycle length exceeds the persistence threshold from the persistence barcode as the key persistence feature of the vector space; Divide the high-dimensional vectors contained in the ring structure into hotspot subsets; Divide the high-dimensional vectors that are not included in any ring structure into non-hotspot subsets; Allocate independent memory address spaces in the memory storage pool for the hotspot subset as independent memory storage nodes.
[0011] Furthermore, the preset granularity mapping table is queried according to the neighbor relationship conflict ratio value, and the maximum edge length parameter of the Vietoris-Rips complex construction is obtained in the following way: Establish a linear mapping function between the neighbor relationship conflict ratio value and the maximum edge length parameter; Input the current neighbor relationship conflict ratio value into the linear mapping function to calculate and output the maximum side length parameter; The linear mapping function satisfies that the maximum edge length parameter decreases monotonically as the neighbor relationship conflict ratio increases.
[0012] Furthermore, the persistent homology group corresponding to the Vietoris-Rips complex is calculated and the corresponding persistent barcode is generated by the following method: Convert the Vietoris-Rips complex to a boundary matrix; Perform matrix reduction algorithm on the boundary matrix to obtain the persistence interval; A life cycle barcode diagram is drawn based on the persistence interval as a persistence barcode.
[0013] Furthermore, step S6 includes: Extract the ID number field from the structured personnel identification information; Generate a verification sample set by uniformly sampling the hash value of the ID number field; Create an independent verification thread pool isolated from the current query task; Switch the query request of the independent verification thread pool to the temporary index structure; Execute query operations on the verification sample set through an independent verification thread pool; Compare the temporary index results returned by the query operation with the query results of the original index structure for the same verification sample set one by one; When the temporary index results of all verification samples completely match the original index results, the consistency verification is confirmed to pass.
[0014] On the other hand, the present invention provides a personnel information big data screening system based on person-document identification, which is used to implement the above-mentioned personnel information big data screening method based on person-document identification, including: Operation detection module, used to receive structured personnel identification information and detect the real-time update operation type of the background database; The conflict resolution module is used to detect the spatiotemporal continuity break characteristics of the newly added data and the historical data when the real-time update operation type is index structure adjustment, and resolve the conflict nodes in the index maintenance operation chain; A parallel construction module is used to select a parallel construction mode for the temporary index structure based on the spatiotemporal continuity break characteristics and the correlation strength of the conflicting nodes, and execute it in coordination with the current query task; The distortion assessment module is used to perform spatial distortion assessment on the high-dimensional vector distribution of the temporary index structure and calculate the conflict ratio of the vector neighbor relationship; The load distribution module is used to determine the computational granularity of topological persistent homology analysis based on the proportion of neighbor relationship conflicts. It performs partition load balancing operations on the temporary index structure based on the extracted key persistence features of the vector space, divides the high-dimensional vector into hot and non-hot subsets, and allocates independent memory storage nodes to the hot subsets. The result verification module is used to generate a verification sample set based on structured personnel identification information, switch the query request to the temporary index structure to perform the verification query, and confirm the consistency with the original index result.
[0015] The beneficial effects of the present invention are: 1. A collaborative optimization mechanism for dynamic index maintenance and real-time querying is established, effectively solving the problem of balancing retrieval accuracy, response speed, and data consistency in high-frequency data update scenarios. By detecting index structure adjustment events in real time and analyzing the characteristics of spatiotemporal continuity breaks, it intelligently selects a parallel index construction mode to accurately capture vector space distortion while ensuring the continuous execution of query tasks. Furthermore, hot data partitioning and isolation are achieved based on topological persistent homology analysis, allowing key vector sets to obtain independent memory resources. This ensures retrieval efficiency under high load from the underlying hardware level, and ultimately ensures zero-deviation in query results during dynamic data updates through a dual-index result verification mechanism.
[0016] 2. Spatiotemporal continuity detection and conflict node resolution predict the impact of index maintenance from the dual dimensions of data distribution and system resources, providing a quantitative basis for building model selection. Topological persistent homology analysis breaks through the limitations of traditional uniform partitioning and accurately captures stable topological features in the vector space through ring structure identification, ensuring a high degree of consistency between memory allocation and data hotspots. An isolated verification mechanism completes index switching security verification without interrupting business queries. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 The figure is a flow chart of the personnel information big data screening method based on human evidence recognition of the present invention.
[0018] Figure 2 It is a structural diagram of the personnel information big data screening system based on human and document identification of the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0020] Example 1
[0021] Figure 1 The present invention provides a method for screening big data of personnel information based on human evidence recognition, including: S1. Receive structured personnel identification information and detect the real-time update operation type of the background database; S2. When the real-time update operation type is index structure adjustment, detect the spatiotemporal continuity break characteristics of the new data and historical data, and resolve the conflicting nodes in the index maintenance operation chain; S3. Select a parallel construction mode for the temporary index structure based on the spatiotemporal continuity break characteristics and the correlation strength of the conflicting nodes, and execute it in coordination with the current query task; S4. Perform spatial distortion evaluation on the high-dimensional vector distribution of the temporary index structure and calculate the vector neighbor relationship conflict ratio; S5. Determine the computational granularity of the topological persistent homology analysis based on the neighbor relationship conflict ratio. Perform partition load balancing on the temporary index structure based on the extracted key persistence features of the vector space, partition the high-dimensional vector into hot and non-hot subsets, and allocate independent memory storage nodes to the hot subsets. S6. Generate a verification sample set based on the structured personnel identification information, switch the query request to the temporary index structure to perform the verification query, and confirm the consistency with the original index result.
[0022] in: S1. Receive structured personnel identification information and detect the real-time update operation type of the background database. The specific implementation includes: The document recognition process processes an input document image, which consists of a scanned copy or photograph of a person's ID or passport. This process uses optical character recognition technology to extract textual information from the document image. This textual information includes, but is not limited to, name, gender, document number, date of birth, and address fields, and combines these fields into structured personal identification information. While extracting textual information, biometric data is generated using a pre-trained deep neural network model. Specifically, this involves locating the facial region in the document image, scaling the facial region image to 112 pixels by 112 pixels, inputting it into a ResNet-34 deep neural network, extracting the 128-dimensional floating-point vector output by the penultimate layer of the network as the facial feature vector, and performing L2-norm normalization on this vector. Establish an association relationship between structured personnel identification information and facial feature vectors. The specific association method is: take the ID number field to calculate the MD5 hash value, intercept the first 8 characters as the basic identifier, append the last 6 digits of the millisecond value of the current system time, form a 16-bit string as a unique association identifier, and write the identifier into the metadata field of the structured personnel identification information and the header storage area of the facial feature vector at the same time.
[0023] The associated structured personnel identification information is transmitted via Hypertext Transfer Protocol to the backend database's update interface, which is a Representational State Transfer Architecture (RSTA) API endpoint. During transmission, the structured personnel identification information is encapsulated into a message body in JavaScript Object Notation format. The message body includes a timestamp field in Coordinated Universal Time (UTC). The timestamp is accurate to the millisecond level and is formatted as year-month-day hour:minute:second.millisecond (e.g., 2023-08-15 14:30:45.500). A standalone log monitoring process captures the backend database's operation logs in real time. The operation logs are stored in a circular overwrite manner on the database server's local storage device. The log monitoring process obtains new log entries through the operating system kernel's file change notification mechanism.
[0024] Log entries are text strings that include the operation time, operation type identifier, and operation content. The operation type identifier is a predefined string constant: when the identifier is INSERT, it represents an insert data operation, UPDATE represents an update data operation, DELETE represents a delete data operation, CREATE_INDEX represents an index creation operation, and DROP_INDEX represents an index deletion operation. When parsing the operation content, pattern matching rules are used to identify keywords: when the content contains the CREATE INDEX character sequence or the DROP INDEX character sequence, the operation instruction type is marked as an index structure adjustment operation; when the content contains the INSERT INTO character sequence or the UPDATE SET character sequence or the DELETE FROM character sequence, the operation instruction type is marked as a data content update operation. If the above pattern cannot be matched, it is marked as another operation type. The marking result and the corresponding log timestamp are written to the operation type record queue, which is a first-in-first-out data structure in memory with a capacity set to, for example, 1,000 records.
[0025] During the data transmission phase, if a Hypertext Transfer Protocol transmission fails, a retransmission mechanism is initiated: the initial retransmission interval is, for example, 2 seconds, and the interval doubles after each failure, reaching a maximum interval of, for example, 60 seconds, with a maximum retransmission limit of, for example, 5. The log monitoring process uses a sliding window reading mechanism, processing only the newly added 1024 bytes of log file content at a time. The document recognition process caches duplicate document images: the SHA-256 hash value of the image file is calculated. If an image with the same hash value is received within, for example, 10 minutes, the cached result is directly returned.
[0026] Structured personal identification information is stored as key-value pairs, with key names strictly corresponding to ID fields: name corresponds to the name field, id_number corresponds to the ID number field, gender corresponds to the gender field, birth_date corresponds to the birth date field, and address corresponds to the address field. All values are string types. When storing facial feature vectors, the header identification field is stored as a 4-byte prefix before the feature data. During the operation type marking process, when a single log entry matches both data operation and index operation keywords, it is preferentially marked as an index structure adjustment operation. The data consumption latency of the operation type record queue is controlled within 5 milliseconds, for example, to ensure real-time processing capabilities.
[0027] The file monitoring mechanism for log monitoring uses the inotify application programming interface in the Linux operating system to monitor log file size changes. A read operation is triggered when the file size increment exceeds, for example, 1 kilobyte. The pattern matching rules for CREATE INDEX are defined as: CREATE followed by one or more whitespace characters, followed by INDEX, case-insensitive. The pattern matching rules for DROP INDEX are defined as: DROP followed by one or more whitespace characters, followed by INDEX, also case-insensitive. Timestamps are generated using the international standard ISO 8601 extended format, with the time zone set to Coordinated Universal Time.
[0028] The unique identifier generation algorithm uses the RFC 1321 standard algorithm for MD5 hash calculation. The system time is obtained through the operating system's clock application programming interface, and the timestamp is the number of milliseconds since 00:00, January 1, 1970, Coordinated Universal Time. The deep neural network model uses the publicly available ResNet-34 architecture, with weight parameters pre-trained using the MS-Celeb-1M dataset. Face detection uses the Viola-Jones algorithm, and face alignment uses a similarity transformation to locate the center of each eye to a fixed coordinate position in the image.
[0029] S2. When the real-time update operation type is index structure adjustment, detect the spatiotemporal continuity break characteristics of the new data and historical data, and resolve the conflicting nodes in the index maintenance operation chain. The specific implementation includes: When the real-time update operation type in the operation type record queue is marked as index structure adjustment, the spatiotemporal continuity break detection process is initiated. New data records are extracted from the transaction log of the backend database. New data records are defined as the data set written after the most recent index structure adjustment operation. During extraction, records containing valid timestamp and geolocation coordinate fields are selected. Timestamp fields are stored in Coordinated Universal Time format with millisecond accuracy (for example, 2023-08-15 14:30:45.500), and geolocation coordinate fields are stored as longitude and latitude pairs in degrees (for example, longitude 116.4074, latitude 39.9042). For each new data record that meets the criteria, its timestamp and geolocation coordinates are added to the new data timestamp sequence and new data geolocation coordinate sequence, respectively. These two sequences are initialized as empty lists.
[0030] Synchronously extract historical data records. Historical data records are defined as data sets written within a specific time period (e.g., 24 hours) before the index structure adjustment. Use the same filtering criteria to obtain valid timestamps and coordinates. Store the timestamps and coordinates of the historical data in the historical data timestamp sequence and the historical data geographic location coordinate sequence, respectively. During the data extraction process, records missing timestamps or coordinates are automatically skipped to ensure that all elements in the sequence are complete and valid.
[0031] To check the continuity of a timestamp sequence, set the sliding window size to, for example, 100 records and the sliding step to, for example, 10 records. Perform a sliding window calculation on the timestamp sequence of newly added data and the timestamp sequence of historical data. The calculation process within each window is as follows: convert all timestamps within the window into a sequence of milliseconds from a base time (for example, 1970-01-01 00:00:00) and calculate the standard deviation of this sequence. The standard deviation calculation process involves taking the arithmetic mean of the sequence, calculating the difference between each value and the mean, summing the squares of the differences, dividing by the number of values, and taking the square root. The resulting standard deviation of the current window serves as the dispersion measure for that window.
[0032] To check the continuity of geographic coordinate sequences, a sliding window of the same size and step size is used. For each set of coordinate points within the window, the geometric center point of the newly added data coordinate point set and the historical data coordinate point set is calculated. The geometric center point calculation process is as follows: the arithmetic average of all longitude values in the window is calculated to obtain the center longitude, and the arithmetic average of all latitude values is calculated to obtain the center latitude. The spherical distance between the two center points is calculated as the center offset. The spherical distance is calculated using standard spherical trigonometry formulas: the longitude difference and latitude difference are converted to radians, and the formula is applied to calculate the great circle distance between the two points. The formula includes the Earth radius parameter.
[0033] When the standard deviation of the timestamp sequence in any window exceeds a temporal continuity threshold (e.g., 300 seconds) or the center offset of any window exceeds a spatial continuity threshold (e.g., 1000 meters), a spatiotemporal continuity break feature flag is generated. The thresholds are determined as follows: the temporal continuity threshold is the upper limit of the normal fluctuation range of the historical timestamp sequence (e.g., several times the average of the historical standard deviation), and the spatial continuity threshold is the typical spatial scale of the application scenario (e.g., a fraction of the diameter of a city block).
[0034] Synchronous monitoring of the index maintenance operation chain: Thread lock wait events in the index maintenance operation chain are captured through the database system interface. The index maintenance operation chain is generated by the database management system and contains the sequence of execution steps for index creation or deletion operations. Each operation step requests a thread lock at startup, and the monitoring system records the lock request time, lock grant time, and lock release time. The lock wait time is calculated as the difference between the lock grant time and the lock request time, in milliseconds.
[0035] When the lock wait time for an operation step exceeds a lock wait threshold (e.g., 500 milliseconds), the operation node is marked as a conflicting node. The lock wait threshold is set based on the performance metrics of the database instance (e.g., a multiple of the average response time). The identification information of all conflicting nodes and their corresponding lock wait time values are stored in a conflict node record list, an in-memory data structure used by subsequent steps.
[0036] During the sliding window calculation process, the window size and step size parameters are dynamically adjusted based on data traffic: when the number of new records per second exceeds a certain value (e.g., 1000), the window size is automatically expanded (e.g., to 200) and the step size is correspondingly increased (e.g., to 20). The Earth's radius used in spherical distance calculations is adjusted based on the coordinate region: a larger value (e.g., 6378137 meters) is used in the equatorial region, and a smaller value (e.g., 6356752 meters) is used in the polar regions. Leap second corrections are taken into account during timestamp conversion, using the international standard leap second table for calibration. During thread lock monitoring, if a distributed lock is encountered, the network delay compensation value is increased (e.g., 50 milliseconds).
[0037] The spatiotemporal continuity break feature is marked as a Boolean state value, set to true when the trigger condition is met and false otherwise. Each entry in the conflict node record list contains three fields: the operation node identifier, the lock wait time, and the operation type. The operation type corresponds to the subtype of the index structure adjustment operation (for example, an index node split operation). The detection process is executed periodically (for example, every 5 seconds). After each detection, the temporary sequence is cleared, but the conflict node record list is retained until the index maintenance operation chain ends.
[0038] Numerical sequence preprocessing in sliding window standard deviation calculations includes outlier filtering: if a timestamp value deviates from the window mean by more than three standard deviations, it is considered an outlier and removed. When calculating center offsets, if the number of coordinate points within a window falls below the minimum sample size (e.g., 10 points), the calculation for that window is skipped. Lock wait time records contain thread identifiers and operation phase identifiers to ensure that conflicting nodes can be accurately located. The index maintenance operation chain monitors all concurrently executed index operation tasks.
[0039] The temporal continuity threshold is set as follows: During system initialization, the standard deviation of a sliding window of historical data timestamp sequences is collected over a continuous 24-hour period. The mean and variance of these standard deviations are calculated, and the mean plus three times the variance is used as the temporal continuity threshold. The spatial continuity threshold is set as the distance obtained by multiplying the typical speed of movement in the application scenario (e.g., the average speed of urban traffic is 30 km / h) by the time window span (e.g., 5 minutes). The lock wait threshold is determined based on database stress testing results: the average lock wait time under 50% load multiplied by a safety factor (e.g., 2).
[0040] The default Earth radius in the spherical distance calculation formula is 6,371,000 meters, an approximate value recommended by the International Astronomical Union. When coordinate points are concentrated in a specific area, regional ellipsoid model parameters are used for precise calculations: the CGCS2000 ellipsoid is used in China, and the GRS80 ellipsoid is used in North America. Leap second correction data is obtained from the announcement documents published by the International Earth Rotation Service, and leap seconds are automatically added or subtracted during timestamp conversions. The network latency compensation value for distributed locks is determined by the 90th percentile of historical round-trip time measurements.
[0041] S3. Select a parallel construction mode for the temporary index structure based on the spatiotemporal continuity break characteristics and the correlation strength of the conflicting nodes, and execute it in coordination with the current query task. The specific implementation includes: After obtaining the spatiotemporal continuity break feature flag and the conflict node record list, the numerical strength of the spatiotemporal continuity break feature is calculated. The numerical strength is calculated as follows: when the spatiotemporal continuity break feature flag is true, the numerical strength is the larger of the maximum sliding window standard deviation or the maximum center offset of the trigger flag; when the flag is false, the numerical strength is set to zero. The quantitative strength of the conflict nodes is the total number of nodes in the conflict node record list. Both are dimensionless numbers. The numerical strength reflects the severity of the spatiotemporal break, while the quantitative strength reflects the intensity of competition for system resources.
[0042] The Pearson correlation coefficient between numerical strength and quantitative strength is calculated as the strength of association. The Pearson correlation coefficient calculation process includes the following steps: preparing a dataset containing at least 30 historical samples, each containing two variables: numerical strength and quantitative strength; calculating the arithmetic mean of each variable; calculating the deviation of each sample value from its respective mean; summing the products of the deviations of the two variables; calculating the square root of the sum of the squared deviations of the two variables; and dividing the sum of the products of the deviations by the product of the two square roots to obtain the correlation coefficient. The calculated result ranges from -1 to +1, with negative values indicating negative correlation and positive values indicating positive correlation.
[0043] When the association strength is greater than the build mode switching threshold (e.g., 0.7), global sharding is selected as the parallel build mode for the temporary index structure. Global sharding is performed by partitioning the data set to be built into several contiguous data blocks (e.g., by primary key range), assigning an independent build thread to each data block to create an index shard, and finally merging the shards to form a complete index. When the association strength is less than or equal to the build mode switching threshold (e.g., 0.7), local incremental build is selected. Local incremental build is performed by constructing an index subtree only for the newly added data, and then attaching the subtree to the existing index structure.
[0044] The threshold for building mode switching is set by measuring the execution efficiency of the two build modes at different correlation strengths during system testing, and using the correlation coefficient corresponding to the efficiency intersection as the threshold. The specific testing method is to perform, for example, 100 global shard builds and 100 local incremental builds at correlation strengths of 0.5, 0.6, 0.7, and 0.8, respectively. The average build time is recorded, and the point at which the global shard build time first falls below the local incremental build time is used as the threshold.
[0045] The process of allocating a separate query thread for the current query task involves requesting a new thread from the operating system's thread manager, setting the thread priority to interactive (e.g., real-time round-robin), and allocating dedicated processor resources (e.g., by binding a core to a processor set). Thread initialization parameters include a pointer to the query statement queue and the address of the result buffer. The query thread continuously monitors the query request queue and executes query statements when the queue is not empty.
[0046] The process of allocating independent build threads for the selected parallel build mode involves determining the number of threads based on the build mode type. In global sharded build mode, the number of threads equals the number of data shards (e.g., 8 threads for 8 shards); in local incremental build mode, a fixed number of threads, such as 2, is allocated. All build threads are set to background priority (e.g., batch processing), and processor resources are isolated from query threads. Build threads receive an index build task descriptor, which contains data address range and index type parameters.
[0047] The operating system kernel scheduler coordinates the execution of query and build threads by alternating them: the kernel scheduler uses a round-robin time-slicing strategy, setting the query thread's time slice to, for example, 10 milliseconds and the build thread's time slice to, for example, 20 milliseconds. When the query thread's time slice expires, the scheduler saves its execution state and switches to the build thread in the ready queue; when the build thread's time slice expires, the scheduler switches back to the query thread. On processors that support concurrent execution, the query and build threads are bound to different processing units.
[0048] Exception handling during the correlation strength calculation process includes: using a default correlation strength value (e.g., 0.5) when historical samples are insufficient; and automatically truncating calculated results to the boundary value when they exceed the [-1, 1] range. Safety constraints are added during build mode selection: local incremental build mode is enforced when system free memory falls below a threshold (e.g., 20% of total memory). Resource constraints are implemented during the thread allocation phase: a memory usage cap is set for build threads (e.g., no more than 40% of total system memory) and a minimum processor share is guaranteed for query threads (e.g., no less than 30%).
[0049] The Pearson correlation coefficient is calculated incrementally: each time a new sample is added, the sum, sum of squares, and sum of products of the two variables are updated, and the correlation coefficient is recalculated using the formula, avoiding the need to traverse the entire dataset. A checkpoint mechanism is implemented during thread execution: progress status is saved every 10,000 records processed, for example, to enable rapid recovery if the thread is preempted. The scheduler records timestamps when switching threads, generating a thread switch sequence log for performance analysis.
[0050] The data partitioning strategy for global sharding builds is dynamically adjusted based on the data type: numeric data is partitioned using equal-value domains, character data is partitioned by lexicographic range, and geographic coordinate data is partitioned by spatial grids. Subtree merge operations for local incremental builds use a lazy merge strategy: a merge is triggered when the subtree volume exceeds, for example, 50% of the parent tree node. Independent query threads use a connection pool management mechanism, with the maximum number of concurrent queries set based on the number of system processing units (e.g., 16 query threads for an 8-core processor). The dynamic priority increase mechanism for build threads is: when the build task delay exceeds, for example, 200 milliseconds, it is temporarily increased to the same priority as the query thread.
[0051] The build mode switching threshold is calibrated by continuously recording the relationship between actual association strength and build time during system operation, recalculating the efficiency crossover point and updating the threshold weekly. The thread switching time slice parameters are dynamically adjusted based on system load: when query response latency exceeds, for example, 100 milliseconds, the query thread time slice is increased to, for example, 15 milliseconds, while the build thread time slice is decreased to, for example, 15 milliseconds. Processor resource binding uses affinity settings to avoid cache invalidation caused by frequent thread migration between processor cores.
[0052] In calculating the Pearson correlation coefficient, the formula for calculating the sum of deviation products is as follows: for each sample, calculate the difference between its numerical strength and the mean, multiply it by the difference between the numerical strength and the mean, and then accumulate the product results for all samples. Square root calculations use the Newton iteration method, with the iteration termination condition being that the difference between two consecutive results is less than, for example, 0.0001. Shard merging in global sharding uses a multi-stage merging strategy: small shards are first merged in pairs, and then larger shards are gradually merged to reduce peak memory usage.
[0053] S4. Perform spatial distortion evaluation on the high-dimensional vector distribution of the temporary index structure and calculate the vector neighbor relationship conflict ratio. The specific implementation includes: After the temporary index structure is constructed, the spatial distortion assessment process is initiated. The original position coordinates and current position coordinates of each high-dimensional vector in the temporary index structure are obtained: the original position coordinates refer to the coordinate values in the original index structure, which are stored in the index metadata cache; the current position coordinates refer to the coordinate values in the temporary index structure, which are stored in the temporary memory database. The dimension of the high-dimensional vector is fixed to, for example, 128 dimensions, and each coordinate value is a single-precision floating-point number type (32 bits). The coordinate acquisition process traverses all vector entries through the index access interface, and simultaneously reads their original coordinate fields and current coordinate fields to ensure that the coordinate data is completely consistent.
[0054] The Euclidean distance between the original and current coordinates of each high-dimensional vector is calculated as the position offset. The Euclidean distance calculation process is as follows: for each dimension (from 1 to the total number of dimensions), the difference between the current coordinate value and the original coordinate value is calculated; this difference is squared; the squared values of all dimensions are accumulated; and the square root of the accumulated result is taken. The distance calculation formula is expressed as: distance equals the square root (the sum of the squares of the differences in all dimensions). The calculated result is a dimensionless number that reflects the degree of displacement of the vector in the feature space. When processing high-dimensional data, dimensionality scaling is implemented: when the number of dimensions exceeds 50, for example, the difference in each dimension is normalized by dividing it by the historical standard deviation of that dimension.
[0055] When the position offset exceeds the spatial distortion threshold, the corresponding high-dimensional vector is marked as a distorted point. The spatial distortion threshold is determined by calculating the statistical distribution of historical data vector offsets during system initialization and taking the 95th percentile value as the threshold (for example, 0.75). The dynamic adjustment mechanism involves recalculating the 95th percentile of the most recent offset after processing a specific number of vectors (for example, 10,000) and updating the threshold using a weighted average of the new and old values (for example, a weight of 0.3 for the new value and 0.7 for the original value). The marking operation writes the unique identifier of the distorted point into the distortion point record list, which is stored in a dynamic array.
[0056] For each distorted point, its neighbor set in the original index is searched. Using the original index's proximity search interface, the original coordinates of the distorted point are used as the query point. A fixed radius (e.g., a Euclidean distance of 1.0) is set for the search neighborhood. The identifiers of all vectors within this radius are returned as the neighbor set. The search radius is set by, for example, taking the average nearest neighbor distance in the original index as three. The search process uses a spatial index acceleration structure (e.g., a KD tree) and a timeout limit (e.g., 100 milliseconds). Upon a timeout, the radius is automatically increased by 10% and the search is retried.
[0057] For each distorted point, its neighbor relationship set in the temporary index structure is searched: using the same search radius and the current coordinates of the distorted point as the query point, the neighbor identifier set is retrieved through the temporary index structure's proximity search interface. The temporary index structure search uses the same acceleration mechanism, using the latest index data in the current build state. The neighbor relationship set is stored as an ordered integer array, where the array elements are the globally unique identifiers of the neighbor vectors.
[0058] Compare the neighbor relationship sets of the distorted points in the original index and the temporary index structure: for each distorted point, record the original neighbor set as set A and the temporary neighbor set as set B. Calculate the symmetric difference of the two sets (that is, elements that belong only to set A or only to set B). The difference calculation is accelerated using bitmaps: when the vector identifier range is continuous, the sets are represented by bitmaps, and the difference is calculated through bit operations. If the number of difference elements exceeds a certain proportion of the average size of the set (for example, 10%), the neighbor relationship is determined to be inconsistent; otherwise, it is considered consistent. Add a tolerance mechanism: if the difference elements are limited to newly added data vectors and the number does not exceed, for example, 5, it is still considered consistent.
[0059] The vector neighbor conflict ratio is calculated by counting the number of distorted points with inconsistent neighbor relationships relative to the total number of distorted points. The calculation process is as follows: traverse the distorted point list and perform a neighbor relationship comparison for each distorted point; record the cumulative number of inconsistent distorted points as M; obtain the total number of distorted points as N; and calculate the ratio value P as M divided by N multiplied by 100. The result P is a percentage value ranging from 0 to 100. Add confidence interval assessment: When the total number of distorted points is less than 50, for example, statistical methods are used to adjust the ratio value.
[0060] The spatial distortion threshold is initialized by collecting vector position offset data for a specific number of recent days (e.g., 7 days) and calculating the 95th percentile of this data. If the data volume is insufficient, a default value (e.g., 0.5) is used. The neighbor search process establishes a result cache: repeated queries for the same coordinates directly return cached results. Distortion point marking is batched: a threshold comparison is performed every time a specific number of vectors (e.g., 1000) are accumulated. Exception handling mechanisms include skipping vector calculations when coordinate data is missing and enabling disk caching when memory is insufficient.
[0061] Euclidean distance calculations utilize processor instruction optimizations: 128-bit registers are used to parallelize the squared difference calculations for each of the four dimensions. Neighborhood set comparisons are memory-constrained: when a single set exceeds a certain number of elements (e.g., 1000), a batch comparison strategy is used. The conflict ratio calculation results are stored as floating-point values for subsequent processing steps. Performance metrics are recorded throughout the calculation process, including the total number of processed vectors, the number of distortion points, and the comparison time, for system optimization.
[0062] Data validation has been added to position offset calculations: coordinate values outside the acceptable range (e.g., [-10, 10]) are marked as outliers and excluded from distortion point statistics. The neighbor search API is frequency-limited: the maximum number of queries per second must not exceed the system's processing capacity (e.g., 1000). The distortion point record list storage structure includes four fields: vector identifier, position offset, original coordinate pointer, and current coordinate pointer. The comparison results record the specific identifiers of the discrepant elements for subsequent analysis. After conflict ratio calculation is complete, an evaluation report is generated, including the ratio value, distortion point distribution histogram, and other statistical information.
[0063] S5. Determine the computational granularity of topology persistence coherence analysis based on the proportion of neighbor relationship conflicts. Perform partition load balancing on the temporary index structure based on the extracted key persistence features of the vector space. Divide the high-dimensional vector into hot and non-hot subsets, and allocate independent memory storage nodes to the hot subsets. The specific implementation includes: Based on the vector neighbor conflict ratio (denoted as P, ranging from 0 to 100) obtained in the previous steps, a preset granularity mapping table is consulted to obtain the maximum edge length parameter for the Vietoris-Rips complex construction. The granularity mapping table is implemented using a linear mapping function: ε = ab × P, where ε represents the maximum edge length parameter, a is the base edge length value (e.g., 2.0), and b is the scaling factor (e.g., 0.01). The function parameters are set based on system test data: on a historical dataset, ε = 1.8 when the conflict ratio is 10% and ε = 1.5 when the conflict ratio is 50%. The values of a and b are obtained through linear regression. This function strictly adheres to the property that the maximum edge length parameter decreases monotonically with increasing neighbor conflict ratio. ε reaches its maximum value (e.g., 2.0) when P = 0 and its minimum value (e.g., 1.0) when P = 100.
[0064] The Vietoris-Rips complex is constructed for the high-dimensional vector set in the temporary index structure using the calculated maximum edge length parameter ε. The construction process includes the following steps: calculating the Euclidean distance between all pairs of vectors to form a distance matrix; creating a set of 0-dimensional simplexes (each vector as a separate simplex); creating 1-dimensional edge simplexes for pairs of vectors with distances less than ε; creating 2-dimensional triangle simplexes for any three vectors whose pairwise distances are less than ε; and recursively constructing higher-dimensional simplexes up to a predefined dimensional limit (e.g., 3 dimensions). The complex construction uses an incremental algorithm: starting from a distance threshold of 0 and gradually increasing to the value of ε, the moment each simplex is first formed as the distance threshold increases is recorded as the generation time.
[0065] The constructed Vietoris-Rips complex is converted into a boundary matrix representation: the rows and columns of the matrix correspond to all simplexes (grouped and sorted by dimension), and the matrix elements take the values 0, +1, or -1, indicating whether the lower-dimensional simplex serves as a boundary for the higher-dimensional simplex. A matrix reduction algorithm is applied to the boundary matrix: starting from the lowest dimension, processing column by column, for the current column, search for the row containing the lowest nonzero element in the column; if the row is not the row of the current column, swap the two rows; use the row to eliminate the elements in the same column position in other rows; repeat until the matrix is reduced to lower triangular form. Record all row transformations, identify the generation and resolution times of the simplex corresponding to the nonzero columns, and generate a set of persistence intervals [birth, death), where birth is the generation time of the simplex and death is the resolution time.
[0066] Draw a persistence barcode chart based on a set of persistence intervals: establish a time axis coordinate system, layered by homology dimension (0D, 1D, 2D, etc.), and draw a horizontal line segment from birth to death for the corresponding persistence interval at each dimension layer. Extract ring structures from the barcode chart whose lifecycle length exceeds the persistence threshold: calculate the lifecycle length L = death - birth for each interval; set the persistence threshold T to the median value of all lifecycle lengths (for example, the sample median); when L is greater than T, record the homology class generation simplex corresponding to the interval (i.e., the core simplex of the ring structure). A dynamic threshold adjustment mechanism: When the memory demand of the hotspot subset exceeds the available system memory, T is automatically increased to, for example, the 75th percentile.
[0067] High-dimensional vectors contained in a ring structure are partitioned into hotspot subsets: all persistent intervals satisfying L > T are traversed to extract the core simplex that generates the homology class; the simplex structure is reverse-mapped to its original high-dimensional vectors; and all vectors associated with the ring structure are merged to form the hotspot subset. Vectors not associated with any ring structure are partitioned into non-hotspot subsets. This partitioning process uses a labeling method: each ring structure is assigned a unique label, which is propagated to all its member vectors, and finally partitioned by label status.
[0068] Allocate a separate memory storage node for the hotspot subset: Calculate the total memory requirement for the hotspot subset (number of vectors × vector dimensions × 4 bytes); apply for contiguous address space in the memory pool using the buddy system algorithm (allocate to the nearest power of 2); copy the hotspot subset vector data to this address space; and create a mapping table from the temporary index structure to this memory area. Memory allocation is subject to capacity constraints: If the calculated memory exceeds 80% of the system's available memory, select from the ring structures associated with the vectors in descending order of lifetime until the capacity is met.
[0069] Distance calculations in the Vietoris-Rips complex construction utilize spatial pruning optimization: a spatial index structure (e.g., a KD tree) is constructed to only compute distances between pairs of vectors within ε. The boundary matrix uses a sparse storage format, recording the row and column indices and sign values of nonzero elements. The matrix reduction algorithm incorporates numerical stability measures: when the absolute value of the pivot is less than, for example, 0.001, a column rotation is performed to find a valid pivot. A logical check is implemented after persistent barcode generation to ensure that the death time is always greater than the birth time. Otherwise, it is automatically corrected to birth + a minimum positive value (e.g., 0.001).
[0070] The linear mapping function's parameters are calibrated monthly: paired data on the conflict ratio P and optimal edge length ε from the last 30 days is collected, and new parameters a and b are fitted using the least squares method. Hotspot subset memory allocation employs a double-buffering strategy: source data remains accessible until the target address space is ready, and seamless migration is achieved through atomic pointer switching. Independent memory storage nodes record metadata blocks containing the starting address, length, access permissions, and associated ring structure identifiers. These metadata blocks are written to the memory management unit registers.
[0071] Dynamically set the upper limit of complex construction dimensions: Calculate the upper limit K = floor(log2(N)) based on the number of vectors N. For example, when N = 10,000, K = 13. Boundary matrix row and column sorting rules: Arrange in ascending order by the time of simplex generation, and in ascending order by dimension at the same time. Added exception handling for hotspot subset partitioning: When the number of vectors in a ring structure exceeds, for example, 1,000, it is automatically split into multiple sub-rings. Memory allocation failure handling mechanism: Attempt to compress the storage format of non-hotspot subsets (for example, 16-bit floating-point numbers) and retry allocation after freeing up space. All operations are audited and logged, including key parameter values and performance metrics such as execution time.
[0072] S6. Generate a verification sample set based on the structured personnel identification information, switch the query request to the temporary index structure to perform the verification query, and confirm the consistency with the original index result. The specific implementation includes: Based on the structured personnel identification information generated in the previous step, the verification sample set generation process is executed: the ID number field in the structured personnel identification information is extracted, which is a string type (for example, an 18-digit ID card number). The ID number field is hashed and a 128-bit hash value is generated using the MD5 hash algorithm. The hash value is converted to an integer H, and the sampling position is determined using the formula S equal to H divided by the total number of records N. Samples are uniformly selected at a fixed step size (for example, every 100 records) to form a verification sample set. The sample set size is set to a specific proportion of the total number of records (for example, 1%), and the minimum sample size is not less than, for example, 100 records. The sampling process ensures spatial distribution uniformity: the hash value space is divided into several intervals (for example, 100 intervals), and an equal number of samples are extracted from each interval.
[0073] Create an independent verification thread pool isolated from the current query task: Request a dedicated thread group from the operating system, with the number of threads set based on the number of processor cores (for example, 8 threads for an 8-core processor). Set up an independent memory address space, isolated from the main query task's memory space via the memory management unit. Configure a dedicated network connection channel and bind it to a separate port range (for example, 5000-5100). Set the thread priority to be lower than that of the main query task (for example, real-time priority for the main task and normal priority for the verification task) to avoid resource competition. Thread pool initialization parameters include the temporary index structure access address and the verification sample set pointer. The maximum number of thread pool threads should not exceed twice the number of system cores.
[0074] Switch query requests from the independent verification thread pool to a temporary index structure: Modify the thread pool's query routing configuration to set the target index address to the temporary index structure's memory address (e.g., starting at address 0x7F000000). Update the query connection pool's endpoint configuration to establish a dedicated TCP persistent connection to the temporary index structure. Load the temporary index structure's metadata descriptor, including the index type, vector dimensions, and search parameters. The switch operation uses an atomic update mechanism: a memory barrier instruction is used to switch the routing table pointer once ready. If the switch fails, it automatically falls back to the original index structure.
[0075] A separate verification thread pool executes queries on the verification sample set: the thread pool controller divides the verification sample set into several subsets (the number of subsets is equal to the number of threads), assigning each thread a subset. The thread execution process includes: obtaining the sample ID number; retrieving the corresponding biometric vector based on the ID number; performing a K-nearest neighbor search (K=1) using the biometric vector as the query input; and returning the identifier and similarity score of the most similar record. The query parameters remain the same as for the main query task: the Euclidean distance metric is used, the hierarchical navigable small-world graph search algorithm is used, and the query timeout is set to, for example, 200 milliseconds.
[0076] The temporary index results returned by the query operation are compared against the query results of the original index structure for the same verification sample set. The original index results are pre-calculated and cached in read-only memory before the temporary index is constructed. The comparison process includes three dimensions: the ID number consistency of the returned records must be an exact match; the biometric similarity difference threshold is set to, for example, 0.05 (normalized similarity range 0 to 1); and the record sorting consistency requires the order of the first K results to be exactly the same (when K = 1, it degenerates to a record match). The comparison execution unit records three indicators: the number of exact matches, the number of samples with a similarity difference exceeding the threshold, and the number of samples with inconsistent sorting.
[0077] When the temporary index results of all verification samples fully match the original index results, consistency verification is confirmed to have passed: a full match is defined as identical ID numbers, a similarity difference less than or equal to 0.05, and consistent sorting. Verification passes when the number of fully matched samples equals the total number of samples. Upon passing, a digital signature verification report is generated, including metrics such as the total number of samples, the number of matched samples, and the maximum similarity difference. If any mismatching samples exist, a secondary verification is performed: the original index is requeried to obtain the latest results, which are then compared again with the temporary index results. If the secondary verification still fails to match, the complete feature vectors of the difference samples are extracted for visual analysis.
[0078] Exception handling during sampling: When the ID number is missing, a composite hash value is generated using alternative fields (e.g., name plus date of birth). Hash algorithm conflict resolution: Open addressing is used to handle samples with the same hash value. Resource limits for thread pool creation: Memory usage is capped at 10% of total memory, and a single thread is forced to terminate if its CPU time exceeds, for example, 200 milliseconds.
[0079] Similarity calculation for result comparison: The difference in similarity scores between the original and temporary results is calculated as the absolute value of the original similarity minus the temporary similarity. Sorting consistency is checked using the Kendall Tau correlation coefficient: a coefficient greater than 0.9 is considered consistent. Verification reports are appended with a timestamp and hash checksum to prevent tampering. Difference sample analysis: The coordinate offset of the feature vector in the original and temporary indexes is calculated, and a histogram of the offset distribution is generated.
[0080] The verification sample set is stored as an encrypted cache file with a validity period of, for example, 24 hours. The thread pool execution process monitors resource usage indicators: CPU utilization, memory usage peak, and network throughput. The comparison result database records historical verification data and supports retrieval and analysis by time range. The final consistency judgment adds a manual review interface: when the automatic verification pass rate is between 99.5% and 100%, the manual review process is triggered. The system records a complete verification log, including detailed information such as the original results, temporary results, and comparison status of each sample. The log retention period is, for example, 30 days.
[0081] Dynamic adjustment mechanism for sample set size: When the total number of records exceeds, for example, 1 million, the sampling ratio automatically drops to 0.5%; when it falls below 100,000, it rises to 2%. Calibration method for the similarity difference threshold: Monthly statistical analysis of historical difference data is performed, and the 99th percentile value of the similarity difference distribution is used as the new threshold. Priority increase conditions for the verification thread pool: When the main query task is idle, the verification thread priority is temporarily increased to accelerate the verification process. Visual analysis of difference samples includes a 3D feature space projection, highlighting vector points whose offset exceeds the threshold.
[0082] Example 2
[0083] Figure 2 The present invention provides a structural diagram of a personnel information big data screening system based on person-document recognition, which includes: Operation detection module, used to receive structured personnel identification information and detect the real-time update operation type of the background database; The conflict resolution module is used to detect the spatiotemporal continuity break characteristics of the newly added data and the historical data when the real-time update operation type is index structure adjustment, and resolve the conflict nodes in the index maintenance operation chain; A parallel construction module is used to select a parallel construction mode for the temporary index structure based on the spatiotemporal continuity break characteristics and the correlation strength of the conflicting nodes, and execute it in coordination with the current query task; The distortion assessment module is used to perform spatial distortion assessment on the high-dimensional vector distribution of the temporary index structure and calculate the conflict ratio of the vector neighbor relationship; The load distribution module is used to determine the computational granularity of topological persistent homology analysis based on the proportion of neighbor relationship conflicts. It performs partition load balancing operations on the temporary index structure based on the extracted key persistence features of the vector space, divides the high-dimensional vector into hot and non-hot subsets, and allocates independent memory storage nodes to the hot subsets. The result verification module is used to generate a verification sample set based on structured personnel identification information, switch the query request to the temporary index structure to perform the verification query, and confirm the consistency with the original index result.
[0084] The calculations involved in the embodiments are all dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to actual conditions.
[0085] It should be noted that the present invention can be deployed on the device itself to realize embedded applications, and can also be run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.
[0086] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wireless or wired transmission. Wired transmission methods include optical fiber, twisted pair, coaxial cable, etc.; wireless transmission methods include infrared, microwave, etc. The computer-readable storage medium can be any available medium accessible by a computer, or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.
[0087] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0088] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0089] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0090] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0091] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0092] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0093] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for screening big data of personnel information based on human evidence recognition, characterized in that: include: S1. Receive structured personnel identification information and detect the real-time update operation type of the background database; S2. When the real-time update operation type is index structure adjustment, detect the spatiotemporal continuity break characteristics of the new data and historical data, and resolve the conflicting nodes in the index maintenance operation chain; S3. Select a parallel construction mode for the temporary index structure based on the spatiotemporal continuity break characteristics and the correlation strength of the conflicting nodes, and execute it in coordination with the current query task; S4. Perform spatial distortion evaluation on the high-dimensional vector distribution of the temporary index structure and calculate the vector neighbor relationship conflict ratio; S5. Determine the computational granularity of the topological persistent homology analysis based on the neighbor relationship conflict ratio. Perform partition load balancing on the temporary index structure based on the extracted key persistence features of the vector space, partition the high-dimensional vector into hot and non-hot subsets, and allocate independent memory storage nodes to the hot subsets. S6. Generate a verification sample set based on the structured personnel identification information, switch the query request to the temporary index structure to perform the verification query, and confirm the consistency with the original index result.
2. The method for screening big data of personnel information based on human evidence recognition according to claim 1 is characterized in that: The step S1 comprises: Obtain structured person identification information from the document image; Extract biometric data and associate it with structured person identification information; Transmit the associated structured personnel identification information to the update interface of the backend database; Monitor the real-time update operation log of the backend database and analyze the operation instruction types in the log; The real-time update operation type is marked as data content update or index structure adjustment according to the operation instruction type.
3. The method for screening big data of personnel information based on person-witness identification according to claim 1, characterized in that: The step S2 comprises: Extract the timestamp sequence and geographic location coordinate sequence of the newly added data; Extract the timestamp sequence and geographic location coordinate sequence of historical data; Calculate the sliding window standard deviation of the timestamp sequence of the newly added data and the timestamp sequence of the historical data; Calculate the center offset of the sliding window between the newly added data geographic location coordinate sequence and the historical data geographic location coordinate sequence; When the sliding window standard deviation exceeds the temporal continuity threshold or the sliding window center offset exceeds the spatial continuity threshold, it is marked as a spatiotemporal continuity break feature; Capture thread lock wait events in the index maintenance operation chain; The operation node that holds the thread lock for longer than the lock wait threshold is marked as a conflict node.
4. The method for screening big data of personnel information based on person-witness identification according to claim 1, characterized in that: The step S3 comprises: The Pearson correlation coefficient between the numerical strength of the spatiotemporal continuity break feature and the numerical strength of the conflict nodes was calculated as the association strength; When the correlation strength is greater than the construction mode switching threshold, global shard construction is selected as the parallel construction mode of the temporary index structure; When the correlation strength is less than or equal to the build mode switching threshold, local incremental build is selected as the parallel build mode of the temporary index structure; Assign an independent query thread to the current query task; Allocate independent build threads for the selected parallel build mode; The query thread and the construction thread are executed alternately through the operating system kernel scheduler.
5. The method for screening big data of personnel information based on person-document identification according to claim 1, characterized in that: The step S4 comprises: Get the original position coordinates and current position coordinates of each high-dimensional vector in the temporary index structure; Calculate the Euclidean distance between the original position coordinates and the current position coordinates of each high-dimensional vector as the position offset; When the position offset exceeds the spatial distortion threshold, the corresponding high-dimensional vector is marked as a distortion point; For each distorted point, find its neighbor relationship set in the original index; For each distorted point, find its neighbor relationship set in the temporary index structure; Compare the neighbor relationship sets of the distorted points in the original index and the temporary index structure; The ratio of the number of distorted points with inconsistent neighbor relationship sets to the total number of distorted points is counted as the vector neighbor relationship conflict ratio.
6. The method for screening big data of personnel information based on person-witness identification according to claim 1, characterized in that: The step S5 comprises: According to the neighbor relationship conflict ratio value, the preset granularity mapping table is queried to obtain the maximum edge length parameter of the Vietoris-Rips complex construction; The maximum edge length parameter is used to construct the Vietoris-Rips complex for the high-dimensional vector set in the temporary index structure; Compute the persistent homology group corresponding to the Vietoris-Rips complex and generate the corresponding persistent barcode; Extract the ring structure whose life cycle length exceeds the persistence threshold from the persistence barcode as the key persistence feature of the vector space; Divide the high-dimensional vectors contained in the ring structure into hotspot subsets; Divide the high-dimensional vectors that are not included in any ring structure into non-hotspot subsets; Allocate independent memory address spaces in the memory storage pool for the hotspot subset as independent memory storage nodes.
7. The method for screening big data of personnel information based on person-witness identification according to claim 6 is characterized in that: The preset granularity mapping table is queried based on the neighbor relationship conflict ratio value to obtain the maximum edge length parameter of the Vietoris-Rips complex construction in the following way: Establish a linear mapping function between the neighbor relationship conflict ratio value and the maximum edge length parameter; Input the current neighbor relationship conflict ratio value into the linear mapping function to calculate and output the maximum side length parameter; The linear mapping function satisfies that the maximum edge length parameter decreases monotonically as the neighbor relationship conflict ratio increases.
8. The method for screening big data of personnel information based on person-witness identification according to claim 6 is characterized in that: Computing the persistent homology group of the corresponding Vietoris-Rips complex and generating the corresponding persistent barcode is achieved by: Convert the Vietoris-Rips complex to a boundary matrix; Perform matrix reduction algorithm on the boundary matrix to obtain the persistence interval; A life cycle barcode diagram is drawn based on the persistence interval as a persistence barcode.
9. The method for screening big data of personnel information based on person-witness identification according to claim 1, characterized in that: The step S6 comprises: Extract the ID number field from the structured personnel identification information; Generate a verification sample set by uniformly sampling the hash value of the ID number field; Create an independent verification thread pool isolated from the current query task; Switch the query request of the independent verification thread pool to the temporary index structure; Execute query operations on the verification sample set through an independent verification thread pool; Compare the temporary index results returned by the query operation with the query results of the original index structure for the same verification sample set one by one; When the temporary index results of all verification samples completely match the original index results, the consistency verification is confirmed to pass.
10. A personnel information big data screening system based on person-evidence recognition, used to implement the personnel information big data screening method based on person-evidence recognition according to any one of claims 1 to 9, characterized in that: include: Operation detection module, used to receive structured personnel identification information and detect the real-time update operation type of the background database; The conflict resolution module is used to detect the spatiotemporal continuity break characteristics of the newly added data and the historical data when the real-time update operation type is index structure adjustment, and resolve the conflict nodes in the index maintenance operation chain; A parallel construction module is used to select a parallel construction mode for the temporary index structure based on the spatiotemporal continuity break characteristics and the correlation strength of the conflicting nodes, and execute it in coordination with the current query task; The distortion assessment module is used to perform spatial distortion assessment on the high-dimensional vector distribution of the temporary index structure and calculate the conflict ratio of the vector neighbor relationship; The load distribution module is used to determine the computational granularity of topological persistent homology analysis based on the proportion of neighbor relationship conflicts. It performs partition load balancing operations on the temporary index structure based on the extracted key persistence features of the vector space, divides the high-dimensional vector into hot and non-hot subsets, and allocates independent memory storage nodes to the hot subsets. The result verification module is used to generate a verification sample set based on structured personnel identification information, switch the query request to the temporary index structure to perform the verification query, and confirm the consistency with the original index result.
Citation Information
Patent Citations
Artificial intelligence digital science and technology platform based on big data analysis
CN120067166A
Meteorological metadata storage method and system based on machine learning
CN120104579A
Intelligent document retrieval and generation system based on metadata driving
CN120104624A
Efficient indexed data structures for persistent memory
US20220027349A1