Batch information retrieval and structured output method and system
The method and system improve the efficiency and security of multi-source data fusion by employing dynamic file parsing, distributed retrieval, and secure reporting to address parsing errors and compliance issues in heterogeneous data processing.
Patent Information
- Application Number
- CN202510812342.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-18
AI Technical Summary
In the multi-source data fusion processing, the existing technology has problems such as high file parsing error rate, low cross-modal data correlation efficiency and insufficient security compliance, resulting in low parsing accuracy, long response time and high compliance risks, which cannot meet the needs of public safety and financial anti-fraud scenarios.
Using dynamic file analysis engine, distributed search and query optimization and security compliance control, through custom configuration, dynamic identification file format, distributed computing resource allocation and improved PageRank algorithm, combined with OCR technology and digital watermark, it realizes accurate analysis of multi-source heterogeneous files, second-level batch search and compliant structured output.
It significantly improves the accuracy and consistency of multi-source heterogeneous file parsing, achieves 100,000-level data second-level response, reduces file parsing error rate and data leakage traceability failure rate, improves data processing efficiency and security, and meets the requirements of GDPR and other regulations.
Smart Images

Figure CN120316163A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of big data analysis and information security, and specifically to a method and system for batch information retrieval and structured output. Background Art
[0002] With the surging demand for multi-source data fusion in public security scenarios, the industry faces threefold challenges of accurate parsing of heterogeneous files, cross-dimensional correlation mining, and compliant output. The current technical system has significant defects: at the file parsing level, traditional systems rely on file extensions to identify formats (such as judging Excel files by.xlsx), and the misjudgment rate of abnormal files with inconsistent content and extensions (such as CSV data disguised as.txt) is as high as 7%. Moreover, the OCR extraction technology for scanned documents lacks semantic segmentation ability, resulting in frequent errors in splitting keyword fields such as ID numbers; at the batch retrieval level, existing distributed architectures use fixed thread pools to process heterogeneous tasks and do not dynamically allocate computing resources according to data types, resulting in a retrieval time of over 180 seconds for ten-thousand-level data. At the security compliance level, the generated reports lack invisible digital watermarks and cannot meet the requirements of regulations such as GDPR for operation traceability and dissemination control. The above problems lead to systematic defects in the existing solutions in terms of processing efficiency (response time > 3 minutes), parsing accuracy (field error rate > 5%), and compliance risks (data leakage traceability failure rate 32%), seriously restricting the value mining and secure application of multi-source data. Summary of the Invention
[0003] The technical task of the present invention is to address the above deficiencies and provide a method and system for batch information retrieval and structured output, which can significantly improve the fusion processing efficiency and security of multi-source data in scenarios such as public security investigation and financial anti-fraud.
[0004] The technical solution adopted by the present invention to solve its technical problems is as follows: A method for batch information retrieval and structured output, which realizes batch information retrieval and structured output based on the parsing of multi-source heterogeneous files, specifically including: Resource index construction and data processing: Based on the data provided by the big data platform, through custom configuration, individual resources and integrated resources are integrated into the search engine through configuration; and the field mapping information is stored in the configuration database to realize the mapping of whether a field is a batch query; Dynamic parsing engine and feature extraction: Dynamically identify the format type of the uploaded file through file header feature codes, and adopt streaming parsing and a hybrid classification model to extract key feature information including ID numbers and mobile phone numbers; Distributed batch retrieval and query optimization: Construct a distributed retrieval task scheduling framework, dynamically allocate computing resources, and generate a ranking of personnel intimacy in combination with an improved PageRank algorithm; Structured output and security compliance control: Automatically synthesize structured Word reports based on the template engine and integrate digital watermarks.
[0005] This method can solve key technical bottlenecks such as high format misjudgment rate, low cross-modal data association efficiency and insufficient security compliance in multi-source heterogeneous file parsing. Through dynamic file parsing engine, intelligent association analysis algorithm and security control mechanism, it can realize accurate parsing, second-level batch retrieval and compliant structured output of massive heterogeneous data (such as Excel, PDF, CSV, etc.), thereby significantly improving the efficiency and security of multi-source data fusion processing in scenarios such as public security investigation and financial anti-fraud.
[0006] Furthermore, the resource index construction and data processing include the following steps: Batch search field configuration: Based on the business scenarios in the public security field, first preset the batch search type, including ID number, license plate number, mobile phone number, case number, etc.; then configure the resources, covering the data source, table name, Chinese name, primary key field, and timestamp field of the resource table; then configure the search engine information, including the search engine type, engine address, index name, number of instances, shards, replica information, and log writing method; finally, configure the mapping field, clarify the index field name, field Chinese description, word segmentation rules, weight, and select the corresponding batch search type. After completing the above configuration, store the configuration content in the configuration database, and create the corresponding index in the search engine according to the configuration; Resource index extraction: read data records from the database (such as relational database MySQL, PostgreSQL, etc.) according to specific query conditions (set as Query Condition) through the corresponding database connection interface JDBC (set the read data record set to be DataSet={Record1, Record2,…, Recordn}); then, convert and map each data record according to the data format requirements of the above configuration; finally, write the converted and mapped data (TransformedDataSet) in batches or one by one to the specified index (Index) and document type (Type) of the search engine through the search engine's client interface (such as the Java High Level REST Client of the search engine ES), thereby completing the data extraction process from the database to the search engine, so as to realize efficient data retrieval and analysis functions in the search engine.
[0007] Furthermore, each data record is converted and mapped according to the data format requirements of the above configuration. Suppose the transformation function is TransformationFunction, the result after mapping is: TransformedDataSet = {TransformationFunction(Record1), TransformationFunction(Record2), …, TransformationFunction(Recordn)}.
[0008] Furthermore, the dynamic parsing engine and feature extraction specifically include: File format recognition: The file formats uploaded in the public security field include text files (such as.doc,.txt, etc.), image files (such as.jpg,.png, etc.), binary structured files (such as.xlsx, etc.), and PDF files, etc.; The parsing engine first preliminarily determines the file format through the file extension, file header information, or specific identifiers, including text files (such as.doc,.txt, etc.), image files (such as.jpg,.png, etc.), binary structured files (such as.xlsx, etc.), and PDF files; For example, the file header of a PDF file usually contains identifiers such as "%PDF-", and image files have specific image format markers. The parsing engine quickly classifies based on these features; Data extraction and conversion: For files of different formats, corresponding extraction methods are adopted; After receiving the file through the MultipartFile interface of Spring Boot, create a corresponding Workbook object according to the extension name, and obtain the Sheet worksheet; Traverse the rows and cells of the worksheet, extract data according to the cell type, and handle null values and exceptions at the same time; After extraction, convert the data according to business requirements, such as processing date formats, calculating numerical values, etc.; Finally, store the converted data in structures such as List or Map for subsequent business processing; Information parsing and extraction: In the text content, usually ID numbers, license plate numbers, mobile phone numbers, case numbers, etc. have specific coding rules; For information such as ID numbers and license plate numbers in images, optical character recognition (OCR) technology can be used to convert the text in the image into editable text, and then regular expressions are used to match and extract.
[0009] Furthermore, for files of different formats, corresponding extraction methods are adopted; For text files, directly read the text content, and then through operations such as character encoding conversion, unify it into a format convenient for processing; For image files, use image processing technology to integrate the OCR service to extract element information such as ID numbers in the image file; For PDF files, use a PDF parsing library to extract the text content and metadata; For excel files, use the Apache POI library to process the file.
[0010] Furthermore, for the distributed batch retrieval and query optimization, a large amount of data is stored dispersedly on multiple nodes, with each node responsible for a part of the data. For example, according to a certain hash algorithm or data range division, the data is sliced into different nodes. For instance, the data is allocated to different nodes according to the hash value of the ID card number, ensuring that the same or related data is stored in the same slice or adjacent nodes as much as possible for subsequent retrieval. When searching, a parallel retrieval method is adopted. When a batch retrieval request is initiated, the query is sent to multiple nodes storing relevant data in parallel. Multiple nodes perform search operations simultaneously, greatly accelerating the search speed. For example, when retrieving a batch of information such as ID card numbers and license plate numbers, different nodes can simultaneously search the data they store instead of searching each node sequentially. After each node completes the search, the results are returned to a coordination node or the client. The coordination node will perform operations such as merging, sorting, and collating on these results to finally obtain a complete result set that meets the requirements. The query optimization includes index optimization and query statement optimization. Regarding the index optimization: First, reasonably design the index structure. According to the fields and query patterns that are frequently queried, create appropriate indexes. This includes creating separate indexes for fields such as ID card numbers, license plate numbers, and contact information to improve query speed. For fields such as case numbers that may have specific format and range queries, design corresponding index structures to optimize the query. As the number of documents increases, the size of the index will also increase. To improve storage efficiency and retrieval speed, it is necessary to compress the inverted index. Technologies such as differential coding are used to compress the index, reducing the storage space and improving the read and write efficiency of the index between memory and disk. Assuming the original index size is S and the compression ratio is r (0 < r < 1), the size of the compressed index is ; Regarding the query statement optimization: First, for fields with exact matches, such as ID card numbers and license plate numbers, use precise term queries instead of match queries that may perform word segmentation. For example, when querying for an ID card number of 123456789012345678, using a term query can directly locate the accurate record, while a match query may perform word segmentation on the ID card number, resulting in a decrease in query efficiency.
[0011] Construct an efficient boolean query. Build a boolean query by reasonably combining clauses such as must (must be satisfied), should (should be satisfied), mustNot (must not be satisfied), etc., to accurately filter data. For example, to query records that simultaneously meet a specific ID number and a case number range, the must clause can be used to combine these two conditions. Use filtering conditions to add filtering conditions to the query to reduce the number of documents that need to be scanned. For instance, first filter out most irrelevant data through some fixed conditions such as area codes, and then conduct a more detailed query. Suppose the original query needs to scan N documents, and M irrelevant documents can be excluded through the filtering conditions. Then the number of documents that need to be scanned after optimization is ; Cache frequently used query results. When the same query is initiated again, directly obtain the results from the cache without having to execute the query operation again. Let the cache hit rate be p (0 ≤ p ≤ 1). If the query execution time without caching is T, then the average query time with caching is , where t is the time to obtain the results from the cache, usually ( ).
[0012] Furthermore, the structured output and security compliance control. The structured output includes the export of batch retrieval results and the export of personnel file information; The export of batch retrieval results: Structurally export the distributed search results; The export of personnel file information: The exported content of the personnel file includes information such as the basic information of the person, the analysis of co-residing persons, the analysis of traveling companions, and the intimacy ranking, etc.; The analysis of co-residing persons: Provided based on the big data platform, identify potential co-residing persons through address similarity calculation and association rule mining. First, clean the address data (such as removing special symbols and unifying the abbreviations of administrative regions), and extract key fields (such as the name of the community and the building number) for structured storage. Then use the edit distance (Levenshtein Distance) to quantify the address differences, and set a threshold to filter out records with high similarity; Verify the shortlisted similar addresses with multi-source data; For example, verify with auxiliary data such as water and electricity payment records and property registration information to improve the accuracy of judging co-residing relationships; Peer identification: Extract travel information with spatio-temporal overlap from traffic records (flights, high-speed rails, buses, etc.) or mobile signaling data, cluster the trajectory points that appear at the same location or within a specified location range within the same time period to generate a list of peer events; then perform feature modeling from two aspects of time window and spatial density, define a time tolerance (such as ±10 minutes) to offset device time errors, identify dense trajectory areas based on the DBSCAN algorithm, and filter out accidental peers; count multiple peer events and assign higher association weights to high-frequency peer combinations; Frequency weight calculation: Count multiple peer events and assign higher association weights to high-frequency peer combinations, e.g., peer strength = number of peer times × log(total travel duration).
[0013] Intimacy ranking: Integrate various relationship data between people, including friend relationships in social networks, colleague relationships at work, relative relationships in families, etc., remove invalid data and duplicate data, and convert the data into a format suitable for algorithm processing, such as constructing a graph structure, where nodes represent people and edges represent relationships between people; Improve on the basis of the traditional PageRank algorithm, considering more factors affecting the intimacy of people. For example, consider the strength of the relationship (e.g., a relationship with more frequent contact has a higher strength) and the type of relationship, etc., assign corresponding weights to the edges and nodes in the graph structure, apply the improved algorithm to the constructed graph structure, and update the PageRank value of each node through iterative calculation; During the calculation process, transfer and distribute the importance scores of nodes according to the weights of the edges and the connection relationships between nodes; Finally, sort the people according to the calculated PageRank values of each person node to generate an intimacy ranking of people; The higher the PageRank value of a person, the more prominent the person is in the intimacy ranking, indicating their intimacy with other people, and output the top 5 people in the intimacy ranking; Through the above calculations and analyses, finally output the basic information of the personnel file (basic information such as name, gender, contact information, etc.), associated information: co-residing personnel, list of peer personnel and their occurrence frequencies, and the top five intimate personnel (information such as name, ID number, contact information, intimacy, etc.) of the personnel; Security and compliance control: Automatically add a watermark to the export result, and the watermark content is the name and ID number of the currently logged-in user.
[0014] The present invention also claims to protect a batch information retrieval and structured output system, including: Resource index construction and data processing module, which is used to integrate a single resource and an integrated resource into a search engine through custom configuration based on the data provided by the big data platform; and store the field mapping information in the configuration database to realize the mapping of whether a field is a batch query; A dynamic parsing engine and an element extraction module are used to dynamically identify the format type of an uploaded file through file header feature codes, and adopt streaming parsing and a hybrid classification model to extract key element information including ID card numbers and mobile phone numbers; A distributed batch retrieval and query optimization module is used to build a distributed retrieval task scheduling framework, dynamically allocate computing resources, and generate a ranking of personnel intimacy in combination with an improved PageRank algorithm; A structured output and security compliance control module is used to automatically synthesize a structured Word report based on a template engine and integrate digital watermarks; The system realizes batch information retrieval and structured output through the above methods.
[0015] The present invention also claims to protect a batch information retrieval and structured output device, including: at least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the above method.
[0016] The present invention also claims to protect a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above method can be implemented.
[0017] Compared with the prior art, a batch information retrieval and structured output method and system of the present invention have the following beneficial effects: 1. The present invention dynamically identifies the format type of an uploaded file through various methods such as combining file header feature codes with file extensions. Compared with the traditional method of only relying on file extensions to identify formats, it can effectively reduce the misjudgment rate of abnormal files (such as CSV data disguised as.txt), and the file parsing error rate is lower than 0.3%. At the same time, for scanned documents, OCR technology integrating semantic segmentation capabilities is adopted, which greatly reduces the splitting errors of keyword fields such as ID card numbers, and significantly improves the accuracy and consistency of multi-source heterogeneous file parsing, providing a reliable basis for subsequent data processing and analysis.
[0018] 2. The present invention builds a distributed retrieval task scheduling framework, dynamically allocates computing resources according to data types, and adopts a parallel retrieval method, enabling the system to achieve a second-level response for 100,000-level data. Compared with the existing distributed architecture that uses a fixed thread pool to process heterogeneous tasks, resulting in a retrieval time of more than 180 seconds for 10,000-level data, the retrieval efficiency of the present invention has been greatly improved. In addition, through index optimization and query statement optimization, such as reasonably designing the index structure, adopting accurate term queries, and constructing efficient Boolean queries, etc., the data retrieval speed is further accelerated, meeting the requirements for rapid retrieval of massive data in scenarios such as public security investigation and enterprise background investigation.
[0019] 3. The present invention integrates association analysis technology, which can mine the potential relationship between multimodal data such as identity, communication, and trajectory. For example, in the analysis of co-residents, potential co-residents can be accurately identified through address similarity calculation and association rule mining, combined with auxiliary data such as water and electricity payment records; the identification of fellow travelers extracts time and space overlapping information from traffic records or mobile phone signaling data, and accurately finds fellow travelers through feature modeling and filtering; based on the improved PageRank algorithm, a variety of relationship data are integrated to generate a personnel intimacy ranking. These functions build a complete multi-dimensional association network, make up for the lack of cross-dimensional association analysis in traditional systems, and provide strong support for deep mining of data value.
[0020] 4. While outputting the structured data, the present invention automatically adds a watermark to the exported data, and the watermark content includes the name and ID number of the currently logged-in user. This method meets the requirements of GDPR and other regulations for operation traceability and dissemination control, and prevents data leakage to a certain extent. Even if a data leakage occurs, the watermark information can be used to trace the source of the data outflow, effectively reducing the failure rate of data leakage tracing and improving the compliance and security of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of a method for batch information retrieval and structured output provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The present invention will be further described below in conjunction with specific embodiments.
[0023] The embodiment of the present invention provides a method for batch information retrieval and structured output, which implements batch information retrieval and structured output based on multi-source heterogeneous file parsing, and specifically includes: 1. Resource index construction and data processing: Based on the data provided by the big data platform, through custom configuration, individual resources and integrated resources are integrated into the search engine through configuration; and the field mapping information is stored in the configuration database to realize whether the field is a batch query mapping. The following steps are included: (1) Batch search field configuration: Based on the business scenarios in the public security field, first preset the batch search type, including ID number, license plate number, mobile phone number, case number, etc.; then configure the resources, covering the data source, table name, Chinese name, primary key field, and timestamp field of the resource table; then configure the search engine information, including the search engine type, engine address, index name, number of instances, shards, replica information, and log writing method; finally, configure the mapping field, clarify the index field name, field Chinese description, word segmentation rules, weight, and select the corresponding batch search type. After completing the above configuration, store the configuration content in the configuration database, and create the corresponding index in the search engine based on the configuration.
[0024] (2) Resource index extraction: Read data records from a database (such as relational databases such as MySQL and PostgreSQL) based on specific query conditions (set as Query Condition) through the corresponding database connection interface JDBC (set the set of data records to be read as DataSet={Record1, Record2,…, Recordn}); then, transform and map each data record according to the data format requirements of the above configuration (set the transformation function as TransformationFunction, and the result after mapping is: TransformedDataSet={TransformationFunction(Record1), TransformationFunction(Record2),…, TransformationFunction(Recordn)}). Finally, write the transformed and mapped data (TransformedDataSet) in batches or one by one to the specified index (Index) and document type (Type) of the search engine through the search engine's client interface (such as the Java High Level REST Client of the search engine ES), thereby completing the data extraction process from the database to the search engine, so as to realize the functions of efficient data retrieval and analysis in the search engine.
[0025] 2. Dynamic parsing engine and feature extraction: Dynamically identify the format type of uploaded files through file header feature codes, and use streaming parsing and hybrid classification models to extract key element information including ID number and mobile phone number.
[0026] (1)File format recognition: The file formats uploaded in the field of public security include text files (such as.doc,.txt, etc.), image files (such as.jpg,.png, etc.), binary structured files (such as.xlsx, etc.), and PDF files, etc. The parsing engine first makes a preliminary judgment on the file format through the file extension, file header information, or specific identifiers. For example, the file header of a PDF file usually contains identifiers such as "%PDF-", and image files have specific image format markers. The parsing engine quickly classifies based on these features.
[0027] (2)Data extraction and conversion: For files of different formats, corresponding extraction methods are adopted. For text files, directly read the text content, and then through operations such as character encoding conversion, unify it into a format convenient for processing. For image files, use image processing technology to integrate OCR services to extract element information such as ID numbers in the image files. For PDF files, use a PDF parsing library to extract text content and metadata; for excel files, use the Apache POI library to process the files. After receiving the file through the MultipartFile interface of Spring Boot, create a corresponding Workbook object according to the file extension, and obtain the Sheet worksheet. Traverse the rows and cells of the worksheet, extract data according to the cell type, and handle null values and exceptions at the same time. After extraction, convert the data according to business requirements, such as processing date formats, calculating numerical values, etc. Finally, store the converted data in structures such as List or Map for subsequent business processing.
[0028] (3)Information parsing and extraction: In the text content, usually ID numbers, license plate numbers, mobile phone numbers, case numbers, etc. all have specific coding rules. For information such as ID numbers and license plate numbers in images, optical character recognition (OCR) technology can be used to convert the text in the image into editable text, and then regular expressions are used to match and extract. The following shows regular expressions of several types: ID number:
[0029] License plate number:
[0030] Mobile phone number:
[0031] Case number: .
[0032] 3. Distributed batch retrieval and query optimization: Build a distributed retrieval task scheduling framework, dynamically allocate computing resources, and generate a ranking of personnel intimacy by combining the improved PageRank algorithm.
[0033] The principle of distributed batch retrieval is as follows: A large amount of data is scattered and stored on multiple nodes, with each node responsible for a part of the data. For example, according to a certain hash algorithm or data range division, the data is sharded to different nodes. For instance, the data is allocated to different nodes according to the hash value of the ID card number, ensuring that the same or related data is stored in the same shard or adjacent nodes as much as possible for subsequent retrieval. When searching, a parallel retrieval method is adopted. When a batch retrieval request is initiated, the query is sent in parallel to multiple nodes storing the relevant data. Multiple nodes perform search operations simultaneously, greatly accelerating the search speed. For example, when retrieving a batch of information such as ID card numbers and license plate numbers, different nodes can simultaneously search the data they store instead of sequentially searching each node one by one. After each node completes the search, the results are returned to a coordination node or the client. The coordination node will perform operations such as merging, sorting, etc. on these results to finally obtain a complete result set that meets the requirements.
[0034] Query optimization includes index optimization and query statement optimization.
[0035] The aforementioned index optimization: First, reasonably design the index structure. According to the frequently queried fields and query patterns, create appropriate indexes. For example, create separate indexes for fields such as ID card numbers, license plate numbers, and contact information to improve the query speed. For fields like case numbers that may have specific format and range queries, design corresponding index structures to optimize the query. As the number of documents increases, the size of the index will also increase accordingly. To improve storage efficiency and retrieval speed, it is necessary to compress the inverted index. Technologies such as differential coding are used to compress the index, reducing the storage space and improving the read and write efficiency of the index between memory and disk. Assuming the original index size is S and the compression ratio is r (0 < r < 1), the size of the compressed index is .
[0036] The aforementioned query statement optimization: First, for fields with exact matches, such as ID card numbers and license plate numbers, use precise term queries instead of match queries that may perform word segmentation. For example, when querying for an ID card number of 123456789012345678, using a term query can directly locate the accurate record, while a match query may perform word segmentation on the ID card number, resulting in a decrease in query efficiency.
[0037] Build efficient Boolean queries by reasonably combining clauses including must (must be satisfied), should (should be satisfied), mustNot (must not be satisfied), etc. to build Boolean queries and accurately filter data; for example, to query records that meet both a specific ID number and a case number range, you can use the must clause to combine these two conditions. Use filtering conditions to add filtering conditions to the query to reduce the number of documents that need to be scanned. For example, first filter out most of the irrelevant data through some fixed conditions such as area codes, and then perform more detailed queries. Assuming that the original query needs to scan N documents, and the filtering conditions can exclude M irrelevant documents, then the number of documents that need to be scanned after optimization is ; Cache frequently used query results. When the same query is initiated again, the results are directly obtained from the cache without executing the query operation again. Suppose the cache hit rate is p (0≤p≤1), and if the query execution time is T when there is no cache, then the average query time when there is a cache is , where t is the time to retrieve the result from the cache, usually ( ).
[0038] 4. Structured output and security compliance control: Automatically synthesize structured Word reports based on the template engine and integrate digital watermarks.
[0039] Structured output includes export of batch search results and personnel file information; (1) Batch retrieval result output: Export the distributed search results from the previous step in a structured manner.
[0040] (2) Export of the personnel file information: The exported personnel file information includes basic information of the personnel, analysis of co-resident personnel, analysis of peer personnel, intimacy ranking and other information.
[0041] Co-resident analysis: Based on the big data platform, potential co-residents are identified through address similarity calculation and association rule mining. First, the address data is cleaned (such as removing special symbols and unifying administrative division abbreviations), and key fields (such as community name and building number) are extracted for structured storage; then the edit distance (Levenshtein Distance) is used to quantify address differences, and a threshold is set to filter out high-similarity records. Multi-source data verification is performed on the filtered similarity shortcuts; for example, verification is performed in combination with auxiliary data such as water and electricity payment records and property registration information to improve the accuracy of co-resident relationship judgment.
[0042] Peer identification: Extract travel information with spatio-temporal overlap from traffic records (flights, high-speed rails, buses, etc.) or mobile phone signaling data. Cluster the trajectory points that appear at the same or adjacent locations within the same time period to generate a list of peer events. Then, perform feature modeling from two aspects: time window and spatial density. Define a time tolerance (e.g., ±10 minutes) to offset device time errors, identify dense trajectory areas based on the DBSCAN algorithm, and filter out accidental peers. Count multiple peer events and assign higher association weights to high-frequency peer combinations; Frequency weight calculation: Count multiple peer events and assign higher association weights to high-frequency peer combinations. For example, peer strength = number of peer events × log(total travel duration).
[0043] Intimacy ranking: Integrate various relationship data between people, including friend relationships in social networks, colleague relationships at work, and kinship relationships in families, etc. Remove invalid and duplicate data, and convert the data into a format suitable for algorithm processing, such as constructing a graph structure, where nodes represent people and edges represent the relationships between people. Improve on the traditional PageRank algorithm, considering more factors that affect the intimacy of people. For example, not only consider the existence or non-existence of relationships, but also consider the strength of relationships (set the intimacy of a single relationship as 1 and accumulate according to the number of times. For example, set the intimacy of one train peer as 1, and the intimacy of 2 train peers as 2), the type of relationship (set different intimacies according to different relationship types. For example, co-resident: 2, marriage: 2, peer: 1, same-car violation: 1), etc. Assign corresponding weights to the edges and nodes in the graph structure, apply the improved algorithm to the constructed graph structure, and update the PageRank value of each node through iterative calculation; During the calculation process, transfer and distribute the importance scores of nodes according to the weights of the edges and the connection relationships between nodes. Finally, rank the people according to the PageRank values of each person node calculated, generate the intimacy ranking of people. The higher the PageRank value of a person, the more forward they are in the intimacy ranking, indicating their intimacy with other people. Output the top 5 people in the intimacy ranking.
[0044] Through the above calculations and analyses, finally output the basic information of the personnel file (basic information such as name, gender, contact information, etc.), associated information: co-resident personnel, list of peer personnel and their occurrence frequencies, and the top five intimate personnel (information such as personnel name, ID number, contact information, intimacy, etc.).
[0045] (3)Security and compliance control: Automatically add a watermark to the exported result, and the watermark content is the name and ID number of the currently logged-in user.
[0046] This method integrates a dynamic file parsing engine, a distributed resource scheduling algorithm, and correlation analysis technology, and is particularly suitable for processing data sources with mixed formats such as Excel tables, scanned documents, and spatio-temporal trajectories. It can achieve accurate extraction of key elements such as ID numbers and mobile phone numbers, construction of multi-dimensional correlation networks, and automatic generation of compliance reports, solving the technical problems of poor parsing consistency of multi-format files, low batch retrieval efficiency, and lack of cross-dimensional correlation analysis in traditional systems. It can achieve second-level response for data of up to 100,000 levels, with a file parsing error rate lower than 0.3%, and can be widely applied in fields such as public security investigation and enterprise background investigation, significantly improving the depth of data utilization and compliance security.
[0047] An embodiment of the present invention also provides a batch information retrieval and structured output system, which realizes batch information retrieval and structured output through the batch information retrieval and structured output method described in the above embodiment.
[0048] The system includes: 1. A resource index construction and data processing module, which is used to, based on the data provided by the big data platform, through custom configuration, integrate single resources and integrated resources into the search engine through configuration; and store the field mapping information in the configuration database to realize the mapping of whether the field is a batch query. It includes: (1) Batch retrieval field configuration: Based on the business scenarios in the field of public security, first preset batch retrieval types, including ID numbers, license plate numbers, mobile phone numbers, case numbers, etc.; then configure the resources, covering the data source, table name, Chinese name, primary key field, and timestamp field of the resource table; subsequently, configure the search engine information, including the search engine type, engine address, index name, number of instances, sharding, replica information, and log writing method; finally, configure the mapping fields, clarify the index field name, field Chinese description, word segmentation rule, weight, and select the corresponding batch retrieval type. After completing the above configuration, store the configuration content in the configuration database and create corresponding indexes in the search engine according to the configuration.
[0049] (2)Resource Index Extraction: From a database (such as relational databases MySQL, PostgreSQL, etc.), according to specific query conditions (set as Query Condition), read data records through the corresponding database connection interface JDBC; then, convert and map each data record according to the data format requirements configured above (set the conversion function as TransformationFunction, and the result after mapping is: TransformedDataSet={TransformationFunction(Record1),TransformationFunction(Record2),…,TransformationFunction(Recordn)}). Finally, through the client interface of the search engine (such as the Java High Level REST Client of the search engine ES), write the converted and mapped data (TransformedDataSet) into the specified index (Index) and document type (Type) of the search engine in batches or one by one, thereby completing the extraction process of data from the database to the search engine to achieve functions such as efficient retrieval and analysis of data in the search engine.
[0050] 2. Dynamic Parsing Engine and Feature Extraction Module, used to dynamically identify the format type of the uploaded file through the file header feature code, and adopt a streaming parsing and hybrid classification model to extract key feature information including ID card numbers and mobile phone numbers. It includes: (1)File Format Identification: The file formats uploaded in the public security field include text files (such as.doc,.txt, etc.), image files (such as.jpg,.png, etc.), binary structured files (such as.xlsx, etc.), and PDF files, etc. The parsing engine first preliminarily judges the file format through the file extension, file header information or specific identifiers. For example, the file header of a PDF file usually contains identifiers such as "%PDF-", and image files have specific image format markers. The parsing engine quickly classifies based on these features.
[0051] (2) Data extraction and transformation: For files in different formats, corresponding extraction methods are adopted. For text files, the text content is directly read, and then through operations such as character encoding conversion, it is unified into a format convenient for processing. For image files, image processing technology is used to integrate OCR services to extract element information such as ID numbers in the image files. For PDF files, a PDF parsing library is used to extract text content and metadata; for excel files, the Apache POI library is used to process the files. After receiving the files through the MultipartFile interface of Spring Boot, a corresponding Workbook object is created based on the file extension, and the Sheet worksheet is obtained. The rows and cells of the worksheet are traversed, data is extracted according to the cell type, and null values and exceptions are processed at the same time. After extraction, the data is transformed according to business requirements, such as processing date formats, calculating numerical values, etc. Finally, the transformed data is stored in structures such as List or Map for subsequent business processing.
[0052] (3) Information parsing and extraction: In the text content, usually ID numbers, license plate numbers, mobile phone numbers, case numbers, etc. have specific coding rules. For information such as ID numbers and license plate numbers in images, optical character recognition (OCR) technology can be used to convert the text in the image into editable text, and then regular expressions are used to match and extract.
[0053] 3. Distributed batch retrieval and query optimization module, which is used to build a distributed retrieval task scheduling framework, dynamically allocate computing resources, and generate a ranking of personnel intimacy in combination with the improved PageRank algorithm.
[0054] The principle of distributed batch retrieval is as follows: A large amount of data is stored dispersedly on multiple nodes, and each node is responsible for a part of the data. For example, according to a certain hash algorithm or data range division, the data is sharded to different nodes. For example, the data is allocated to different nodes according to the hash value of the ID number, ensuring that the same or related data is stored in the same shard or adjacent nodes as much as possible for subsequent retrieval. When searching, a parallel retrieval method is adopted. When a batch retrieval request is initiated, the query will be sent to multiple nodes storing relevant data in parallel. Multiple nodes perform search operations simultaneously, greatly accelerating the search speed. For example, to retrieve a batch of information such as ID numbers and license plate numbers, different nodes can search the data they store simultaneously instead of searching each node sequentially. After each node completes the search, the results are returned to a coordination node or the client. The coordination node will perform operations such as merging, sorting, etc. on these results to finally obtain a complete result set that meets the requirements.
[0055] Query optimization includes index optimization and query statement optimization.
[0056] The index optimization described above: First, reasonably design the index structure and create appropriate indexes based on frequently queried fields and query patterns. For example, separate indexes are created for fields such as ID number, license plate number, and contact information to increase query speed. For fields such as case number that may have specific formats and range queries, design corresponding index structures to optimize queries. As the number of documents increases, the size of the index will also increase. In order to improve storage efficiency and retrieval speed, it is necessary to compress the inverse index, and use techniques such as differential encoding to compress the index, reduce storage space, and improve the read and write efficiency of the index between memory and disk. Assuming that the original index size is S and the compression ratio is r (0 <r<1),则压缩后索引大小为 .
[0057] The query optimization is as follows: First, for fields with exact matches, such as ID card number and license plate number, use precise term query instead of match query that may perform word segmentation. For example, when querying the ID card number of 123456789012345678, the term query can directly locate the accurate record, while the match query may perform word segmentation on the ID card number, resulting in reduced query efficiency.
[0058] Build efficient Boolean queries by reasonably combining clauses including must (must be satisfied), should (should be satisfied), mustNot (must not be satisfied), etc. to build Boolean queries and accurately filter data; for example, to query records that meet both a specific ID number and a case number range, you can use the must clause to combine these two conditions. Use filtering conditions to add filtering conditions to the query to reduce the number of documents that need to be scanned. For example, first filter out most of the irrelevant data through some fixed conditions such as area codes, and then perform more detailed queries. Assuming that the original query needs to scan N documents, and the filtering conditions can exclude M irrelevant documents, then the number of documents that need to be scanned after optimization is ; Cache frequently used query results. When the same query is initiated again, the results are directly obtained from the cache without executing the query operation again. Suppose the cache hit rate is p (0≤p≤1), and if the query execution time is T when there is no cache, then the average query time when there is a cache is , where t is the time to retrieve the result from the cache, usually ( ).
[0059] 4. Structured output and security compliance control module, which is used to automatically synthesize structured Word reports based on the template engine and integrate digital watermarks.
[0060] Structured output includes export of batch search results and personnel file information; Batch retrieval result output: Export the distributed search results from the previous step in a structured manner.
[0061] Export of the personnel file information: The exported content of the personnel file includes information such as the basic information of the personnel, analysis of co-residing personnel, analysis of fellow travelers, and intimacy ranking.
[0062] Analysis of co-residing personnel: Provided by the big data platform, potential co-residing personnel are identified through address similarity calculation and association rule mining. First, clean the address data (such as removing special symbols and standardizing administrative division abbreviations), and extract key fields (such as community name and building number) for structured storage; then use the Levenshtein Distance to quantify the address differences and set a threshold to filter records with high similarity. Conduct multi-source data verification on the selected similar records; for example, combine auxiliary data such as water and electricity payment records and property registration information for verification to improve the accuracy of co-residing relationship judgment.
[0063] Identification of fellow travelers: Extract travel information with spatio-temporal overlap from traffic records (flights, high-speed rails, buses, etc.) or mobile signaling data. Cluster the trajectory points that appear at the same or adjacent locations within the same time period to generate a list of fellow travel events. Then perform feature modeling from two aspects: time window and spatial density, define a time tolerance (such as ±10 minutes) to offset device time errors, identify dense trajectory areas based on the DBSCAN algorithm, and filter out accidental fellow travel. Count multiple fellow travel events and assign higher association weights to high-frequency fellow travel combinations; Frequency weight calculation: Count multiple fellow travel events and assign higher association weights to high-frequency fellow travel combinations, such as: Fellow travel intensity = Number of fellow travel times × log(Total travel duration).
[0064] Intimacy Ranking: Integrate various relationship data among personnel, including friend relationships in social networks, colleague relationships at work, and kinship relationships in families, etc. Remove invalid and duplicate data, and convert the data into a format suitable for algorithm processing, such as constructing a graph structure, where nodes represent personnel and edges represent relationships between personnel. Improve on the basis of the traditional PageRank algorithm, considering more factors that affect the intimacy of personnel. For example, not only consider the existence of relationships, but also consider the strength of relationships (set the intimacy of a relationship as 1 and accumulate it according to the number of times. For example, the intimacy of traveling by train once is set as 1, and the intimacy of traveling by train twice is 2), the type of relationship (set different intimacies according to different relationship types. For example, living together: 2, marriage: 2, traveling together: 1, being caught for the same traffic violation: 1), etc. Assign corresponding weights to the edges and nodes in the graph structure, apply the improved algorithm to the constructed graph structure, and update the PageRank value of each node through iterative calculation; during the calculation process, according to the weights of the edges and the connection relationships between nodes, transfer and distribute the importance scores of the nodes. Finally, sort the personnel according to the calculated PageRank values of each personnel node, generate the intimacy ranking of personnel. The higher the PageRank value of a person, the more forward they are in the intimacy ranking, indicating their intimacy with other personnel, and output the top 5 personnel in the intimacy ranking.
[0065] Through the above calculations and analyses, finally output the basic information of the personnel file (basic information such as name, gender, contact information, etc.), associated information: list of co-residing personnel, list of traveling companions and the frequency of occurrence, and the top five personnel with the highest intimacy (information such as personnel name, ID number, contact information, intimacy, etc.).
[0066] Security and Compliance Control: Automatically add a watermark to the export result, and the content of the watermark is the name and ID number of the currently logged-in user.
[0067] The embodiment of the present invention also provides a batch information retrieval and structured output device, including: at least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the batch information retrieval and structured output method described in the above embodiment.
[0068] An embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the batch information retrieval and structured output method described in the above embodiment is implemented. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium.
[0069] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0070] Embodiments of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0071] In addition, it should be clear that not only can the functions of any one of the above embodiments be implemented by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0072] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part and all of the actual operations, thereby implementing the functions of any one of the above embodiments.
[0073] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.
Claims
1. A method for batch information retrieval and structured output, characterized in that Implementing batch information retrieval and structured output based on multi-source heterogeneous file parsing, specifically including: Resource index construction and data processing: Based on the data provided by the big data platform, through custom configuration, individual resources and integrated resources are integrated into the search engine through configuration; and the field mapping information is stored in the configuration database to achieve the mapping of whether the field is for batch query; Dynamic parsing engine and feature extraction: Dynamically identify the format type of the uploaded file through the file header feature code, and use streaming parsing and hybrid classification models to extract key feature information including ID card numbers and mobile phone numbers; Distributed batch retrieval and query optimization: Build a distributed retrieval task scheduling framework, dynamically allocate computing resources, and generate a ranking of personnel intimacy in combination with the improved PageRank algorithm; Structured output and security compliance control: Automatically synthesize a structured Word report based on the template engine and integrate digital watermarks.
2. The method for batch information retrieval and structured output according to claim 1, characterized in that The resource index construction and data processing include the following steps: Batch retrieval field configuration: Based on the business scenarios in the field of public security, first preset the batch retrieval types, including ID card numbers, license plate numbers, mobile phone numbers, and case numbers; then configure the resources, covering the data source, table name, Chinese name, primary key field, and timestamp field of the resource table; subsequently configure the search engine information, including the search engine type, engine address, index name, number of instances, sharding, replica information, and log writing method; finally configure the mapping fields, clarify the index field name, field Chinese description, word segmentation rule, weight, and select the corresponding batch retrieval type. After completing the above configuration, store the configuration content in the configuration database and create corresponding indexes in the search engine according to the configuration; Resource index extraction: Read data records from the database according to specific query conditions through the corresponding database connection interface JDBC; then, convert and map each data record according to the configured data format requirements; finally, write the converted and mapped data into the specified index and document type of the search engine in batches or one by one through the client interface of the search engine, thus completing the extraction process of data from the database to the search engine.
3. The method for batch information retrieval and structured output according to claim 2, wherein The conversion and mapping of each data record according to the above configured data format requirements Let the conversion function be TransformationFunction, and the result after mapping is: TransformedDataSet={TransformationFunction(Record1),TransformationFunction(Record2),…,TransformationFunction(Recordn)}.
4. A method for batch information retrieval and structured output according to claim 1, characterized in that, The dynamic parsing engine and feature extraction specifically include: File format identification: The parsing engine first preliminarily judges the file format through the file extension, file header information, or specific identifiers, including text files, image files, binary structured files, and PDF files; Data extraction and conversion: Use corresponding extraction methods for files of different formats; after receiving the file with the help of Spring Boot's MultipartFile interface, create the corresponding Workbook object according to the extension name and obtain the Sheet worksheet; traverse the rows and cells of the worksheet, extract data according to the cell type, and handle null values and exceptions; after extraction, convert the data according to business requirements, and finally store the converted data in a List or Map structure; Information parsing and extraction: Use optical character recognition technology to convert the text in the image into editable text, and then use regular expressions to match and extract it.
5. A method for batch information retrieval and structured output according to claim 4, characterized in that The corresponding extraction methods are adopted for files of different formats; For text files, the text content is directly read, and then the character encoding conversion operation is performed to unify it into a format that is easy to process; for image files, image processing technology is used to integrate OCR services to extract element information in image files; For PDF files, use the PDF parsing library to extract text content and metadata; For excel files, Apache POI library is used to process the files.
6. A method for batch information retrieval and structured output according to claim 1, characterized in that The distributed batch retrieval and query optimization stores a large amount of data in multiple nodes, with each node responsible for a portion of the data; a parallel retrieval method is used during the search, and when a batch retrieval request is initiated, the query is sent in parallel to multiple nodes storing related data; multiple nodes perform search operations simultaneously; after each node completes the search, the result is returned to a coordination node or client; the coordination node performs operations including merging, sorting and sorting on the results, and finally obtains a complete result set that meets the requirements; Query optimization includes index optimization and query statement optimization. The index optimization: create indexes according to frequently queried fields and query patterns, including: creating separate indexes for ID card numbers, license plate numbers, and contact information fields; designing corresponding index structures for fields with specific format and range queries; compressing inverted indexes. Let the original index size be S and the compression ratio be r, where 0 < r < 1, then the size of the compressed index is ; The query statement optimization: first, for the fields with exact matches, use precise term queries instead of match queries that perform word segmentation; Build efficient Boolean queries by combining must, should, and mustNot clauses to accurately filter data; add filtering conditions to the query to reduce the number of documents that need to be scanned; cache frequently used query results.
7. A method for batch information retrieval and structured output according to claim 1, characterized in that, The structured output and security compliance control, structured output includes batch search result export and personnel file information export; The batch search result output: exporting the distributed search results in a structured manner; Export of personnel file information: The personnel file export content includes basic information of personnel, analysis of co-resident personnel, analysis of peer personnel, and intimacy ranking information; The co-resident analysis: based on the big data platform, potential co-residents are identified through address similarity calculation and association rule mining; first, the address data is cleaned and key fields are extracted for structured storage; Then, the edit distance is used to quantify the address differences and a threshold is set to filter out high-similarity records; multi-source data verification is performed on the filtered similarity shortcuts; Peer identification: Extract travel information with spatio-temporal overlap from traffic records or mobile signaling data, cluster trajectory points that appear at the same location or within a specified location range during the same time period, and generate a list of peer events; Then, perform feature modeling from two aspects: time window and spatial density. Define a time tolerance to offset device time errors, identify dense trajectory regions based on the DBSCAN algorithm, and filter out accidental peers; Count multiple peer events, and assign high correlation weights to high-frequency peer combinations; Intimacy ranking: Integrate various relationship data between people, including friend relationships in social networks, colleague relationships at work, and relative relationships in families, remove invalid and duplicate data, convert the data into a format suitable for algorithm processing, and construct a graph structure, where nodes represent people and edges represent relationships between people; Assign corresponding weights to the edges and nodes in the graph structure, and update the PageRank value of each node through iterative calculation; During the calculation process, transfer and distribute the importance scores of nodes according to the weights of the edges and the connection relationships between nodes; Finally, sort the people according to the calculated PageRank value of each person node to generate a person intimacy ranking; The higher the PageRank value of a person, the higher their position in the intimacy ranking, indicating their intimacy with other people, and output the top 5 people in the intimacy ranking; Through the above calculations and analyses, finally output the basic information and associated information of the personnel file: co-residing personnel, list of peer personnel and their occurrence frequencies, and the top five intimate personnel; Security and compliance control: Automatically add a watermark to the exported result, and the watermark content is the name and ID number of the currently logged-in user.
8. A batch information retrieval and structured output system, characterized in that Including: Resource index construction and data processing module, which is used to integrate single resources and integrated resources into the search engine through custom configuration based on the data provided by the big data platform; And store the field mapping information in the configuration database to implement the mapping of whether the field is a batch query; Dynamic parsing engine and element extraction module, which is used to dynamically identify the format type of the uploaded file through the file header feature code, and extract key element information including ID numbers and mobile phone numbers using streaming parsing and a hybrid classification model; Distributed batch retrieval and query optimization module, which is used to build a distributed retrieval task scheduling framework, dynamically allocate computing resources, and generate a person intimacy ranking in combination with an improved PageRank algorithm; Structured output and security compliance control module, which is used to automatically synthesize a structured Word report based on a template engine and integrate digital watermarks; The system realizes batch information retrieval and structured output through the method described in any one of claims 1 to 7.
9. An apparatus for batch information retrieval and structured output, characterized in that Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.
10. A computer-readable medium, characterized in that, Computer instructions are stored on the computer-readable medium, and when the computer instructions are executed by the processor, the method described in any one of claims 1 to 7 can be realized.
Citation Information
Patent Citations
Cloud search system based on distributed search engine
CN118093986A
Knowledge graph platform, query method and device thereof and electronic equipment
CN118261243A
Method for realizing efficient indexing and retrieval of mass data
CN118503489A
Case investigation auxiliary system based on multi-source data association analysis
CN118643465A
Optimization method for index retrieval of digital library
CN119128045A
Cited By
Intelligent printing system and method based on template engine
CN120723187A
Batch data importing method and device supporting offline and online data correction
CN121255823A