A big data-based multi-source data intelligent analysis management method and system

By using big data intelligent analysis and management methods, the problems of outdated equipment and data integration faced by grassroots public security organs in handling electronic evidence have been solved. This has enabled efficient repair of multi-source data and accurate identification of abnormal events, thereby improving investigation efficiency and risk prevention and control capabilities.

CN121144267BActive Publication Date: 2026-05-05ECONOMIC CRIME INVESTIGATION DETACHMENT OF GUANGZHOU PUBLIC SECURITY BUREAU
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ECONOMIC CRIME INVESTIGATION DETACHMENT OF GUANGZHOU PUBLIC SECURITY BUREAU
Filing Date
2025-09-08
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

When handling electronic evidence, grassroots public security organs face problems such as outdated equipment, low data recovery success rate, lack of integration mechanism for multi-source data, low efficiency of manual analysis and easy to miss key clues, making it difficult to meet the needs of investigation.

Method used

We adopt a multi-source data intelligent analysis and management method based on big data. Through big data acquisition, analysis and risk prevention modules, we realize the anomaly detection and repair of data files, unify time format and geographic coordinates, construct the main behavior timeline, trajectory map and spatiotemporal event map, and combine multi-dimensional anomaly judgment logic to improve the standardization and automation of anomaly event identification.

Benefits of technology

It significantly improved the integrity and utilization of multi-source data, shortened the analysis cycle, reduced the misjudgment rate, improved the efficiency of event response, and achieved precise risk prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121144267B_ABST
    Figure CN121144267B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-source data intelligent analysis and management method and system based on big data, belonging to the field of big data analysis technology. The method of this invention includes big data acquisition, big data analysis, and risk prevention and control. First, through big data acquisition, after collecting data files, anomaly detection is performed to identify repairable abnormal files, and these repairable abnormal files are repaired. Second, through big data analysis, heterogeneous data is integrated from the cloud database based on the target event. Data integration is achieved through unified time format, geographical coordinates, and identity identifiers to determine whether the target event is abnormal. If abnormal, its behavior type is determined. Finally, through risk prevention and control, based on the status data of various regions and behavior types, the regions and behavior types with risk exceeding the threshold are analyzed. Early warning is issued by detecting the amount of sensitive chat segments of the target, thereby achieving precise risk prevention and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data analytics, specifically to a method and system for intelligent analysis and management of multi-source data based on big data. Background Technology

[0002] Electronic evidence, carrying comprehensive information about a suspect's clothing, food, housing, transportation, and entertainment, has become a "digital key" for reconstructing crime scenes in criminal investigations. However, current electronic evidence processing faces prominent problems such as difficulty in recovery, analysis, and utilization. A multi-source data intelligent analysis and management method and system based on big data is needed to integrate professional knowledge and improve work efficiency.

[0003] Grassroots public security organs generally face the problems of outdated technology and equipment when handling electronic evidence. Traditional methods for processing electronic evidence are ill-suited for damaged Office documents, ZIP compressed files, or abnormally encoded data, resulting in low success rates in repair and failing to meet investigative needs. For example, when storage media suffers physical damage, grassroots police officers can only report the electronic devices through multiple levels of management, wasting a lot of time and missing the golden period for case solving. At the same time, facing massive amounts of electronic data, grassroots units lack efficient analysis tools. With the advent of the big data era, the amount of electronic data involved in cases has exploded. Traditional analysis methods relying on manual review are inefficient and seriously hinder the progress of case solving. Currently, the amount of data such as chat logs and transfer records obtained by police officers often reaches hundreds of gigabytes. Manual screening is not only inefficient but also prone to missing crucial clues.

[0004] Currently, most technical departments can only extract data from evidence to ensure the legality of data sources, but they lack the extraction and analysis of key data within the evidence. On one hand, while multiple departments have case investigation outlines, this professional knowledge, such as criminal investigation outlines, exists primarily in text form and has not been effectively transformed into operational technical models. For example, in economic crime cases, criminals often use multiple false identities, complex fund transfer channels, and covert communication methods to commit crimes. The electronic data involved includes call records, transfer information, and web browsing history from multiple sources, with intricate relationships between them. Police officers struggle to quickly apply this experience and knowledge to electronic evidence analysis, leading to deviations in the direction of the investigation. On the other hand, multi-source police data is scattered across different systems, lacking an integration mechanism. Household registration information, case files, and surveillance videos are independent, requiring grassroots police officers to frequently switch systems when retrieving and analyzing them. Furthermore, different data formats are incompatible, making it impossible to form a complete chain of evidence. Summary of the Invention

[0005] To address the aforementioned technical shortcomings, the present invention aims to provide a method and system for intelligent analysis and management of multi-source data based on big data.

[0006] To solve the above technical problems, the present invention adopts the following technical solution: The present invention provides a multi-source data intelligent analysis and management method based on big data, including the following steps: Step 1, Big Data Acquisition: Collect data files, perform anomaly detection on the data files to obtain each abnormal data file and each normal readable data file, analyze each abnormal data file to obtain each repairable abnormal data file, repair each repairable abnormal data file to obtain each repaired abnormal data file, and summarize and store each repaired abnormal data file and each normal readable data file into a big data cloud database.

[0007] Step 2: Big Data Analysis: Based on the set target event, integrate the big data cloud library to obtain various heterogeneous data of the target event. Analyze the various heterogeneous data of the target event to determine whether the target event is abnormal. When the target event is abnormal, analyze the data file of the target event in the big data cloud library to obtain the behavior type of the target event.

[0008] Step 3: Risk prevention and control. Upload the behavior type of the target event to the big data cloud database to obtain the status data of each behavior type in each region. Analyze the status data of each behavior type in each region to carry out risk prevention and control.

[0009] Preferably, the repair process for each repairable abnormal data file is as follows: skip the central directory for each first-class repairable file, directly parse the filename and compression method of the header information of the file entity fragment, extract the file content, and thus repair each first-class repairable file.

[0010] The entity corruption in each of the second-category repairable files is marked as "unrecoverable". The file name and compression method of the header information of the file entity fragments are parsed, and the non-"unrecoverable" file content is extracted. This is used to repair each of the second-category repairable files.

[0011] For each type of third-category repairable file, the verification error is ignored. The original byte stream of the file is extracted first, and then the readable content is restored through subsequent text recognition or format conversion. This is used to repair each type of third-category repairable file.

[0012] Each of the fourth category of repairable files automatically switches its encoding format and re-parses based on the anomaly detection results. If the encoding format cannot be matched, the garbled characters are replaced by a character mapping table to repair each of the fourth category of repairable files.

[0013] Each of the fifth category of repairable files uses a local reconstruction algorithm: when a Word document has missing XML tags causing content corruption, the algorithm will fill in the missing tags based on the consistency of the format of the complete tags before and after; when an Excel spreadsheet cell is damaged, the algorithm will infer and repair the abnormal cell content based on the data type of adjacent cells, thereby repairing each of the fifth category of repairable files.

[0014] Preferably, the integration of the big data cloud database is carried out in the following specific process: the heterogeneous data of the target event includes the behavioral sequence of the target event, associated locations, and the behavior of each group.

[0015] Extract cloud database files of target events from big data cloud databases. For time formats of different data sources, convert them into Unix timestamps using time parsing algorithms. For different geographic data formats, convert them into a unified coordinate system using geocoding technology. For identity identifiers such as ID card numbers, mobile phone numbers, device MAC addresses, and social media accounts, use hash desensitization to retain uniqueness while protecting privacy. Then, establish a feature code mapping table to associate multi-dimensional identifiers of the same subject.

[0016] Based on the set target events, extract the time, address, and line of each target event. Based on the same identity feature code, aggregate the behavioral data of the subject at different time points to construct a subject behavior timeline and obtain the behavioral sequence of the target events. Based on the same identity feature code, aggregate the occurrence records of the subject at different geographical coordinates to construct a subject trajectory map and obtain the associated locations of the target events. Based on the same timestamp, aggregate the behaviors of all subjects in the same geographical area at that time point to construct a spatiotemporal event map, identify group behaviors at the same time and place, and obtain the group behaviors of the target events.

[0017] On the other hand, the present invention provides a multi-source data intelligent analysis and management system based on big data, including the following modules: a big data acquisition module, used to collect data files, perform anomaly detection on the data files to obtain each abnormal data file and each normal readable data file, analyze each abnormal data file to obtain each repairable abnormal data file, repair each repairable abnormal data file to obtain each repaired abnormal data file, and summarize and store each repaired abnormal data file and each normal readable data file into a big data cloud database.

[0018] The big data analytics module is used to integrate the big data cloud database based on the set target event, obtain the heterogeneous data of the target event, analyze the heterogeneous data of the target event, determine whether the target event is abnormal, and when the target event is abnormal, analyze the data file of the target event in the big data cloud database to obtain the behavior type of the target event.

[0019] The risk prevention and control module is used to upload the behavior type of the target event to the big data cloud database, thereby obtaining the status data of each behavior type in each region, analyzing the status data of each behavior type in each region, and carrying out risk prevention and control.

[0020] The beneficial effects of this invention are as follows: 1. This invention first acquires data through big data, collects data files, performs anomaly detection, identifies repairable abnormal files, and repairs the repairable abnormal files. 2. Through big data analysis, it integrates heterogeneous data from the cloud database based on the target event, and achieves data integration through unified time format, geographical coordinates, and identity identifiers to determine whether the target event is abnormal. If abnormal, it determines its behavior type. Finally, through risk prevention and control, based on the status data of various regions and behavior types, it analyzes and obtains the regions and behavior types where the risk exceeds the threshold, and provides early warning prompts by detecting the amount of sensitive chat segments of the target, thus achieving precise risk prevention and control.

[0021] 2. This invention addresses fragmented ZIP files, garbled documents, and structurally damaged Office files containing damaged data. It designs a classification and repair strategy, which, through a process of "anomaly detection - repairability analysis - targeted repair," effectively recovers repairable anomaly data, improves the utilization of damaged data, and transforms potentially unusable invalid data into valuable analytical evidence, significantly enhancing the integrity and utilization rate of multi-source data.

[0022] 3. This invention effectively solves the heterogeneity problem of multi-source data in terms of time, space, and subject by unifying heterogeneous time formats to Unix timestamps through time parsing algorithms, converting multi-source geographic data into a unified coordinate system through geocoding technology, and hash-based de-identification to associate multi-dimensional identity identifiers of the same subject. Based on this, it constructs a "subject behavior timeline," a "trajectory map," and a "spatiotemporal event map," realizing multi-dimensional associations of "time-geography-behavior-subject," providing a structured and traceable data foundation for anomaly event analysis, and reducing the inefficiency and error rate of traditional manual association.

[0023] 4. This invention improves the standardization and automation of abnormal event identification by using a multi-dimensional abnormal judgment logic of "behavioral sequence deviation + associated location risk matching + gang behavior pattern analysis" combined with a preset "normal behavior baseline". Compared with the traditional judgment method that relies on human experience, it reduces misjudgment and omission caused by subjective judgment, while shortening the abnormal event analysis cycle and significantly improving event response efficiency. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of the implementation steps of the method of the present invention.

[0026] Figure 2 This is a schematic diagram of the system structure connection of the present invention. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] according to Figure 1 As shown, the present invention provides a multi-source data intelligent analysis and management method based on big data, including the following steps: Step 1, Big Data Acquisition: Collect data files, perform anomaly detection on the data files to obtain each abnormal data file and each normal readable data file, analyze each abnormal data file to obtain each repairable abnormal data file, repair each repairable abnormal data file to obtain each repaired abnormal data file, and summarize and store each repaired abnormal data file and each normal readable data file into a big data cloud database.

[0029] In one specific embodiment, the anomaly detection of the data files is carried out as follows: the data files include each ZIP file and each document.

[0030] Scan all binary data related to the ZIP format in the storage medium, extract the file header identifier and central directory identifier to obtain the file header identifier and central directory identifier of each ZIP file. Parse the starting address of the file data recorded in the central directory in the storage medium and check whether this address is continuous with the actual storage address of the file header. If the address corresponding to the offset of a ZIP file matches the actual position of the file header and the file data is not truncated by other data segments, it is judged as normal, and thus each normal ZIP file is obtained. If the address corresponding to the offset of a ZIP file does not match the actual position of the file header, or the file data is truncated by other data segments, it is judged as abnormal, and thus each abnormal ZIP file is obtained.

[0031] The documents are divided into Word documents and Excel documents.

[0032] Check the closure of XML tags in each Word document. If a Word document does not contain any unclosed or illegal tags, it is considered normal, and this process yields all normal Word documents. If a Word document contains unclosed or illegal tags, it is considered structurally corrupted, and this process yields all abnormal Word documents. Check the integrity of the stream object in an Excel document. If the stream object in an Excel document is complete and its length matches the specified value, it is considered normal, and this process yields all normal Excel documents. If the stream object in an Excel document is missing or its length is abnormal, it is considered abnormal, and this process yields all abnormal Excel documents.

[0033] By summing up all normal ZIP files, normal Word documents, and normal Excel documents, we obtain all normal readable data files. By summing up all abnormal ZIP files, abnormal Word documents, and abnormal Excel documents, we obtain all abnormal data files.

[0034] In one specific embodiment, the analysis of each abnormal data file is carried out as follows: If the central directory of an abnormal ZIP file is damaged but the file entity fragments are intact, this type of file is recorded as a first type of repairable file, and so on, to obtain each first type of repairable file; if part of the file entity of an abnormal ZIP file is damaged, this type of file is recorded as a second type of repairable file, and so on, to obtain each second type of repairable file; if an abnormal ZIP file does not match the CRC check, this type of file is recorded as a third type of repairable file, and so on, to obtain each third type of repairable file.

[0035] Summarize all the abnormal Word documents and all the abnormal Excel documents to obtain the abnormal documents. If an abnormal document has an encoding error, record it as a fourth type of repairable file, and so on to obtain all fourth type repairable files. If an abnormal document has a damaged structure, record it as a fifth type of repairable file, and so on to obtain all fifth type repairable files.

[0036] By summing up all the first-class repairable files, the second-class repairable files, the third-class repairable files, the fourth-class repairable files, and the fifth-class repairable files, we can obtain each repairable abnormal data file.

[0037] In one specific embodiment, the repair process for each repairable abnormal data file is as follows: skip the central directory for each first-type repairable file, directly parse the filename and compression method of the header information of the file entity fragment, extract the file content, and thus repair each first-type repairable file.

[0038] The entity corruption in each of the second-category repairable files is marked as "unrecoverable". The file name and compression method of the header information of the file entity fragments are parsed, and the non-"unrecoverable" file content is extracted. This is used to repair each of the second-category repairable files.

[0039] For each type of third-category repairable file, the verification error is ignored. The original byte stream of the file is extracted first, and then the readable content is restored through subsequent text recognition or format conversion. This is used to repair each type of third-category repairable file.

[0040] Each of the fourth category of repairable files automatically switches its encoding format and re-parses based on the anomaly detection results. If the encoding format cannot be matched, the garbled characters are replaced by a character mapping table to repair each of the fourth category of repairable files.

[0041] Each of the fifth category of repairable files uses a local reconstruction algorithm: when a Word document has missing XML tags causing content corruption, the algorithm will fill in the missing tags based on the consistency of the format of the complete tags before and after; when an Excel spreadsheet cell is damaged, the algorithm will infer and repair the abnormal cell content based on the data type of adjacent cells, thereby repairing each of the fifth category of repairable files.

[0042] Step 2: Big Data Analysis: Based on the set target event, integrate the big data cloud library to obtain various heterogeneous data of the target event. Analyze the various heterogeneous data of the target event to determine whether the target event is abnormal. When the target event is abnormal, analyze the data file of the target event in the big data cloud library to obtain the behavior type of the target event.

[0043] In one specific embodiment, the integration of the big data cloud database is carried out as follows:

[0044] The heterogeneous data of the target event includes the behavioral sequence of the target event, associated locations, and the behavior of various groups. The cloud database files of the target event are extracted from the big data cloud database. For the time format of different data sources, they are uniformly converted into Unix timestamps through time parsing algorithms. For different geographical data formats, they are converted into a unified coordinate system through geocoding technology. For identity identifiers such as ID card number, mobile phone number, device MAC address, and social media account, uniqueness is preserved and privacy is protected through hash desensitization. Then, a feature code mapping table is established to associate multi-dimensional identifiers of the same subject.

[0045] Based on the set target events, extract the time, address, and line of each target event. Based on the same identity feature code, aggregate the behavioral data of the subject at different time points to construct a subject behavior timeline and obtain the behavioral sequence of the target events. Based on the same identity feature code, aggregate the occurrence records of the subject at different geographical coordinates to construct a subject trajectory map and obtain the associated locations of the target events. Based on the same timestamp, aggregate the behaviors of all subjects in the same geographical area at that time point to construct a spatiotemporal event map, identify group behaviors at the same time and place, and obtain the group behaviors of the target events.

[0046] In one specific embodiment, the analysis of the heterogeneous data of the target event is carried out as follows: each target subject is obtained from the target event, and the number of times each target subject performs each action, the number of times each location is recorded, and the number of times the group is associated are extracted from the big data cloud database.

[0047] Actions of each target subject at each timestamp that exceed the preset number of actions are recorded as normal actions. If the action sequence of a target event does not conform to the normal action of a target subject at a certain timestamp, it indicates that the target subject at that timestamp is an abnormal time subject, and that timestamp is an abnormal timestamp. If the number of abnormal timestamps is greater than the preset number of abnormal timestamps or the number of abnormal time subjects at a certain timestamp is greater than the number of abnormal subjects, it indicates that the target event is an abnormal event.

[0048] Locations where the number of times each target entity is located at a location at each timestamp exceeds the preset number of times are recorded as normal locations. These locations are then aggregated to obtain the normal areas for each target entity at each timestamp. If the associated location of a target event is a normal area for a target entity at a certain timestamp, and this associated location belongs to a risk area address in the database, it indicates that the target event is an abnormal event.

[0049] If the number of times a target entity is associated with a group at a certain timestamp is greater than the preset number of group associations, it indicates that the timestamp is an abnormal timestamp of the group. If the number of abnormal entities in an abnormal timestamp of a group is greater than the preset value, or if the number of entities in the normal area of ​​the target entity in the abnormal timestamp of the group is greater than the preset number of entities, it indicates that the target event is an abnormal event.

[0050] In one specific embodiment, the analysis of the data file of the target event in the big data cloud database is carried out as follows: The data file of the target event consists of chat segments and chat words of the target event. The feature data of sensitive words and sensitive text segments of each behavior type are obtained from the database. The similarity between the feature data of the chat segments of the target event and the feature data of the sensitive text segments of each behavior type is calculated to obtain the text segment similarity of each behavior type of the target event. The similarity between the feature data of the chat words of the target event and the feature data of the sensitive words of each behavior type is calculated to obtain the word segment similarity of each behavior type of the target event. According to a preset weighting factor, the text segment similarity and word segment similarity of each behavior type of the target event are weighted and calculated to obtain the sensitivity similarity of each behavior type of the target event. The behavior type with the highest sensitivity similarity is selected as the behavior type of the target event, thereby obtaining the behavior type of the target event.

[0051] Step 3: Risk prevention and control. Upload the behavior type of the target event to the big data cloud database to obtain the status data of each behavior type in each region. Analyze the status data of each behavior type in each region to carry out risk prevention and control.

[0052] In one specific embodiment, the analysis of the status data of each behavior type in each region is carried out as follows: the status data of each behavior type in each region includes the number of analyses and the number of cases closed for each behavior type in each region and time period, and the total number of analyses and the total number of cases closed for each behavior type in each region are obtained by summarizing them.

[0053] Divide the total number of analyses for each behavior type in each region by the preset analysis frequency threshold to obtain the analysis incidence rate for each behavior type in each region. Divide the total number of case closures for each behavior type in each region by the preset case closure frequency threshold to obtain the case closure incidence rate for each behavior type in each region. Based on the weighting factors preset by the staff, the analysis incidence rate and case closure incidence rate for each behavior type in each region are weighted and calculated to obtain the risk index for each behavior type in each region.

[0054] It should be noted that the analysis frequency threshold and case closure frequency threshold are the maximum values ​​of the number of analysis sessions and case closure sessions for each behavior type. These are generally set by staff based on historical maximum values. The specific values ​​are set by staff, for example, the analysis frequency threshold is 8 and the case closure frequency threshold is 6. The weighting factors preset by staff are related to the case closure rate. The higher the case closure rate, the greater the weighting factor of the case closure occurrence rate. The lower the case closure rate, the greater the weighting factor of the analysis occurrence rate. The specific values ​​are set by staff, for example, the weighting factor of the case closure occurrence rate is 0.3 and the weighting factor of the analysis occurrence rate is 0.7. The case closure rate is calculated as the number of case closures divided by the number of case analyses.

[0055] The number of analyses and the number of cases closed for each behavior type in each region and time period are plotted on a time axis to obtain trend curves for the number of analyses and the number of cases closed for each behavior type in each region. Image recognition technology is used to obtain the slope of the number of analyses and the slope of the number of cases closed for each behavior type in each region. Based on the weighting factors preset by the staff, the slopes of the number of analyses and the slopes of the number of cases closed for each behavior type in each region are weighted and calculated to obtain the risk trend index for each behavior type in each region.

[0056] It should be noted that the weighting factors in the weighted calculation of the slope of the number of analyses and the slope of the number of cases closed are related to the frequency of cases. The higher the frequency of cases, the larger the weighting factor of the slope of the number of analyses, and the lower the frequency of cases, the larger the weighting factor of the slope of the number of cases closed. The specific values ​​are set by the staff. For example, the weighting factor of the slope of the number of analyses is 0.4 and the weighting factor of the slope of the number of cases closed is 0.6.

[0057] In one specific embodiment, the risk prevention and control process is as follows: If the risk index of a certain behavior type in a certain region is greater than a preset risk index threshold or the risk trend index is greater than a preset risk trend index threshold, it indicates that the behavior type in that region is risky. Risk prevention and control is carried out on this type of behavior in that region by detecting the chat segments of each object in that region to obtain the number of sensitive chat segments of this type of behavior for each object in that region. When the number of sensitive chat segments of this type of behavior for each object in that region is greater than a preset sensitive chat segment number threshold, an early warning is issued to that object.

[0058] It should be noted that the risk index threshold and risk trend index threshold are the maximum values ​​of the risk index and risk trend index under normal circumstances. When these thresholds are exceeded, it indicates that there is a risk of the corresponding behavior type in the region. The specific values ​​are set by the staff. For example, the risk index threshold is 1.3 and the risk trend index threshold is 1.6. The sensitive chat message quantity threshold is the maximum value of the sensitive chat message quantity in normal chat. When the sensitive chat message quantity is greater than the threshold, it indicates that there is a risk of the specific behavior type of the chat message. The specific value is set by the staff.

[0059] according to Figure 2 As shown, the present invention provides a multi-source data intelligent analysis and management system based on big data, including the following modules: big data acquisition module, big data analysis module, risk prevention and control module, and database.

[0060] The big data analysis module is connected to the big data acquisition module and the risk prevention and control module, respectively. The big data acquisition module, the big data analysis module, and the risk prevention and control module are all connected to the database.

[0061] The big data acquisition module is used to collect data files, perform anomaly detection on the data files, obtain each abnormal data file and each normal readable data file, analyze each abnormal data file to obtain each repairable abnormal data file, repair each repairable abnormal data file to obtain each repaired abnormal data file, and summarize and store each repaired abnormal data file and each normal readable data file into the big data cloud database.

[0062] The big data analytics module is used to integrate the big data cloud database based on the set target event, obtain the heterogeneous data of the target event, analyze the heterogeneous data of the target event, determine whether the target event is abnormal, and when the target event is abnormal, analyze the data file of the target event in the big data cloud database to obtain the behavior type of the target event.

[0063] The risk prevention and control module is used to upload the behavior type of the target event to the big data cloud database, thereby obtaining the status data of each behavior type in each region, analyzing the status data of each behavior type in each region, and carrying out risk prevention and control.

[0064] The database is used to store the addresses of risk areas, the feature data of sensitive words and phrases of each behavior type, and the feature data of sensitive text segments of each behavior type.

[0065] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this specification, they should all fall within the protection scope of the present invention.

Claims

1. A multi-source data intelligent analysis and management method based on big data, characterized in that, Includes the following steps: Step 1: Big Data Acquisition: Collect data files, perform anomaly detection on the data files to obtain each abnormal data file and each normal readable data file, analyze each abnormal data file to obtain each repairable abnormal data file, repair each repairable abnormal data file to obtain each repaired abnormal data file, and summarize and store each repaired abnormal data file and each normal readable data file into the big data cloud database. Step 2, Big Data Analysis: Based on the set target event, integrate the big data cloud library to obtain various heterogeneous data of the target event. Analyze the various heterogeneous data of the target event to determine whether the target event is abnormal. When the target event is abnormal, analyze the data file of the target event in the big data cloud library to obtain the behavior type of the target event. Step 3: Risk prevention and control. Upload the behavior type of the target event to the big data cloud database to obtain the status data of each behavior type in each region. Analyze the status data of each behavior type in each region to carry out risk prevention and control. The analysis of status data for each behavior type in each region is carried out in the following specific process: the status data for each behavior type in each region includes the number of analyses and the number of cases closed for each behavior type in each region and time period, and the total number of analyses and the total number of cases closed for each behavior type in each region are obtained by summarizing them. Divide the total number of analyses for each behavior type in each region by the preset analysis number threshold to obtain the analysis occurrence rate for each behavior type in each region. Divide the total number of case closures for each behavior type in each region by the preset case closure number threshold to obtain the case closure occurrence rate for each behavior type in each region. Based on the weighting factors preset by the staff, the analysis occurrence rate and case closure occurrence rate for each behavior type in each region are weighted and calculated to obtain the risk index for each behavior type in each region. The number of analyses and the number of cases closed for each behavior type in each region and time period are plotted on a time axis to obtain trend curves for the number of analyses and the number of cases closed for each behavior type in each region. Image recognition technology is used to obtain the slope of the number of analyses and the slope of the number of cases closed for each behavior type in each region. Based on the weighting factors preset by the staff, the slopes of the number of analyses and the slopes of the number of cases closed for each behavior type in each region are weighted and calculated to obtain the risk trend index for each behavior type in each region.

2. The method for intelligent analysis and management of multi-source data based on big data according to claim 1, characterized in that, The anomaly detection process for the data files is as follows: The data files include each ZIP file and each document; Scan all binary data related to the ZIP format in the storage medium, extract the file header identifier and central directory identifier to obtain the file header identifier and central directory identifier of each ZIP file, parse the starting address of the file data recorded in the central directory in the storage medium, check whether the address is continuous with the actual storage address of the file header, if the address corresponding to the offset of a ZIP file matches the actual position of the file header and the file data is not truncated by other data segments, it is judged as normal, and thus obtain each normal ZIP file; if the address corresponding to the offset of a ZIP file does not match the actual position of the file header, or the file data is truncated by other data segments, it is judged as abnormal, and thus obtain each abnormal ZIP file; The documents are divided into Word documents and Excel documents; Check the closure of XML tags in each Word document. If a Word document does not contain any unclosed or illegal tags, it is considered normal, and this process is used to obtain normal Word documents. If a Word document contains unclosed or illegal tags, it is considered structurally corrupted, and this process is used to obtain abnormal Word documents. Check the integrity of the stream object in an Excel document. If the stream object in an Excel document is not missing and its length matches, it is considered normal, and this process is used to obtain normal Excel documents. If the stream object in an Excel document is missing or its length is abnormal, it is considered abnormal, and this process is used to obtain abnormal Excel documents. By summing up all normal ZIP files, normal Word documents, and normal Excel documents, we obtain all normal readable data files. By summing up all abnormal ZIP files, abnormal Word documents, and abnormal Excel documents, we obtain all abnormal data files.

3. The method for intelligent analysis and management of multi-source data based on big data according to claim 2, characterized in that, The analysis of each abnormal data file is described in the following process: If the central directory of an abnormal ZIP file is damaged but the file fragments are intact, this type of file is recorded as a first-class repairable file, and so on. If part of the file entity of an abnormal ZIP file is damaged, this type of file is recorded as a second-class repairable file, and so on. If an abnormal ZIP file does not match the CRC check, this type of file is recorded as a third-class repairable file, and so on. Summarize all abnormal Word documents and all abnormal Excel documents to obtain all abnormal documents. If an abnormal document has an encoding error, record it as a fourth type of repairable file, and so on to obtain all fourth type repairable files. If an abnormal document has a damaged structure, record it as a fifth type of repairable file, and so on to obtain all fifth type repairable files. By summing up all the first-class repairable files, the second-class repairable files, the third-class repairable files, the fourth-class repairable files, and the fifth-class repairable files, we can obtain each repairable abnormal data file.

4. The method for intelligent analysis and management of multi-source data based on big data according to claim 3, characterized in that, The repair process for each recoverable abnormal data file is as follows: For each type of first-class repairable file, skip the central directory, directly parse the file name and compression method in the header information of the file entity fragment, extract the file content, and use this to repair each type of first-class repairable file; The entity corruption in each of the second-class repairable files is marked as "unrecoverable". The file name and compression method of the header information of the file entity fragments are parsed, and the non-"unrecoverable" file content is extracted to repair each of the second-class repairable files. For each type of third-category repairable file, check errors are ignored. The original byte stream of the file is extracted first, and then the readable content is restored through subsequent text recognition or format conversion. This is used to repair each type of third-category repairable file. Each of the fourth category of repairable files automatically switches its encoding format and re-parses based on the anomaly detection results. If the encoding format cannot be matched, the garbled characters are replaced by a character mapping table to repair each of the fourth category of repairable files. Each of the fifth category of repairable files uses a local reconstruction algorithm: when a Word document has missing XML tags causing content corruption, the algorithm will fill in the missing tags based on the consistency of the format of the complete tags before and after; when an Excel spreadsheet cell is damaged, the algorithm will infer and repair the abnormal cell content based on the data type of adjacent cells, thereby repairing each of the fifth category of repairable files.

5. The method for intelligent analysis and management of multi-source data based on big data according to claim 1, characterized in that, The integration of the big data cloud database is described in the following process: The heterogeneous data of the target event includes the behavioral sequence of the target event, associated locations, and the behavior of each group; Extract cloud database files of target events from big data cloud databases, convert time formats from different data sources to Unix timestamps using time parsing algorithms, convert different geographic data formats to a unified coordinate system using geocoding technology, and de-identify identity identifiers such as ID card numbers, mobile phone numbers, device MAC addresses, and social media accounts using hash desensitization to retain uniqueness while protecting privacy. Then, establish a feature code mapping table to associate multi-dimensional identifiers of the same subject. Based on the set target events, extract the time, address, and line of each target event. Based on the same identity feature code, aggregate the behavioral data of the subject at different time points to construct a subject behavior timeline and obtain the behavioral sequence of the target events. Based on the same identity feature code, aggregate the occurrence records of the subject at different geographical coordinates to construct a subject trajectory map and obtain the associated locations of the target events. Based on the same timestamp, aggregate the behaviors of all subjects in the same geographical area at that time point to construct a spatiotemporal event map, identify group behaviors at the same time and place, and obtain the group behaviors of the target events.

6. The method for intelligent analysis and management of multi-source data based on big data according to claim 5, characterized in that, The analysis of the heterogeneous data of the target event is carried out in the following specific process: Extract each target entity from the target event, and extract the number of each target entity's actions, the number of times in each location, and the number of times they are associated with a group from the big data cloud database at each time stamp; Actions of each target subject at each timestamp that exceed the preset number of actions are recorded as normal actions. If the action sequence of a target event does not conform to the normal action of a target subject at a certain timestamp, it indicates that the target subject at that timestamp is an abnormal time subject and that timestamp is an abnormal timestamp. If the number of abnormal timestamps exceeds the preset number of abnormal timestamps or the number of abnormal time subjects at a certain timestamp exceeds the number of abnormal subjects, it indicates that the target event is an abnormal event. Locations where the number of times each target entity is at a location at each time stamp exceeds the preset number of times are recorded as normal locations. These locations are then aggregated to obtain the normal areas for each target entity at each time stamp. If the associated location of a target event is a normal area for a target entity at a certain time stamp, and this associated location belongs to a risk area address in the database, it indicates that the target event is an abnormal event. If the number of times a target entity is associated with a group at a certain timestamp is greater than the preset number of group associations, it indicates that the timestamp is an abnormal timestamp of the group. If the number of abnormal entities in an abnormal timestamp of a group is greater than the preset value, or if the number of entities in the normal area of ​​the target entity in the abnormal timestamp of the group is greater than the preset number of entities, it indicates that the target event is an abnormal event.

7. The method for intelligent analysis and management of multi-source data based on big data according to claim 6, characterized in that, The analysis of the data files of the target events in the big data cloud database is carried out in the following specific process: The data file for the target event consists of chat segments and chat phrases. Sensitive phrases and sensitive text segments for each behavior type are retrieved from the database. The similarity between the feature data of the chat segments and the feature data of the sensitive text segments for each behavior type is calculated to obtain the text segment similarity for each behavior type. Similarly, the similarity between the feature data of the chat phrases and the feature data of the sensitive phrases for each behavior type is calculated to obtain the phrase similarity for each behavior type. Based on a preset weighting factor, the text segment similarity and phrase similarity for each behavior type are weighted to obtain the sensitivity similarity for each behavior type. The behavior type with the highest sensitivity similarity is selected as the behavior type of the target event, thus obtaining the behavior type of the target event.

8. The method for intelligent analysis and management of multi-source data based on big data according to claim 1, characterized in that, The risk prevention and control process is as follows: If the risk index of a certain behavior type in a certain region is greater than the preset risk index threshold or the risk trend index is greater than the preset risk trend index threshold, it indicates that the behavior type in that region is risky. Risk control measures are implemented for this type of behavior in that region. This is achieved by detecting the number of sensitive chat messages of this type of behavior for each individual in that region. When the number of sensitive chat messages of this type of behavior for each individual in that region exceeds the preset sensitive chat message quantity threshold, an early warning is issued to that individual.

9. A smart analysis and management method system applying the multi-source data intelligent analysis and management method based on big data as described in any one of claims 1-8, characterized in that, Includes the following modules: The big data acquisition module is used to collect data files, perform anomaly detection on the data files, obtain each abnormal data file and each normal readable data file, analyze each abnormal data file to obtain each repairable abnormal data file, repair each repairable abnormal data file to obtain each repaired abnormal data file, and summarize and store each repaired abnormal data file and each normal readable data file into the big data cloud database. The big data analysis module is used to integrate the big data cloud library according to the set target event, obtain the heterogeneous data of the target event, analyze the heterogeneous data of the target event, determine whether the target event is abnormal, and when the target event is abnormal, analyze the data file of the target event in the big data cloud library to obtain the behavior type of the target event. The risk prevention and control module is used to upload the behavior type of the target event to the big data cloud database, thereby obtaining the status data of each behavior type in each region, analyzing the status data of each behavior type in each region, and carrying out risk prevention and control.

Citation Information

Patent Citations

  • ChatGPT-based global perception early warning method and system, medium and electronic equipment

    CN118411811A

  • Asset risk tracing method and device

    CN119205351A

  • Drug risk monitoring method based on multi-source data fusion

    CN120234773A