A multi-regional data compliance adaptation engine intelligent system

The intelligent system of multi-regional data compliance adaptation engine solves the problem of automatically identifying and filtering user information based on encryption requirements in different regions, achieving efficient and secure encryption processing and improving the accuracy and security of the system.

CN121167752BActive Publication Date: 2026-03-31ZHONGHE YUNKE INFORMATION TECH GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies fail to automatically identify user information requiring encryption based on the encryption requirements of different regions, and fail to accurately filter user information for encryption.

Method used

Design a multi-regional data compliance adaptation engine intelligent system, including information collection, information extraction, feature analysis, information processing and encryption modules. By collecting regional information and user information, the system classifies and filters image and text information using an encrypted database, and performs encryption processing in combination with regional encrypted information.

Benefits of technology

It improves the efficiency and security of encryption, accurately identifies and processes user information, and enhances the accuracy and security of the multi-regional data compliance adaptation engine intelligent system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167752B_ABST
    Figure CN121167752B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and more particularly to a multi-territory data compliance adaptation engine intelligent system. The system comprises: an information collection module for collecting territorial information and user information; an information extraction module comprising a picture classification unit and a text extraction unit; a feature analysis module comprising a picture screening unit for screening pictures to be encrypted and a paragraph analysis unit for determining sensitive information; an information processing module for determining whether non-sensitive information is encrypted or updating an encrypted database; and an encryption module comprising a first encryption unit for encrypting pictures to be encrypted and sensitive information, and a second encryption unit for encrypting non-sensitive information determined to be encrypted. The present application utilizes the above-mentioned modules to cooperate with each other, effectively improving the information recognition efficiency, and further improving the accuracy of the multi-territory data compliance adaptation engine intelligent system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a multi-regional data compliance adaptation engine intelligent system. Background Technology

[0002] In recent years, a large number of companies have begun to develop overseas businesses, with their overseas markets spanning the globe. This has led to a sharp increase in the challenges and complexity of data compliance and security. If a new solution is developed for each new country, it will seriously hinder the company's overseas expansion and bring great complexity to the system. By setting up a solution and storing the data compliance requirements of each country in a configured manner, and selecting the appropriate solution through dynamic routing when in use, the system does not need to focus on compliance issues, but can instead focus on more valuable work such as the rapid implementation of system business. It can encrypt and store personal information in accordance with the privacy regulations of sensitive personal data in different countries / regions.

[0003] Chinese Patent Application Publication No. CN116108502A discloses a method, system, device, and medium for generating and decrypting secure electronic documents. This application provides a method for generating secure electronic documents. It obtains electronic signature data in a unified format by identifying and authenticating electronic signature-related data. Based on the signature base data, biometric information, supplementary evidence data, business additions, and image steganography of the entire signature process, a signature image is obtained, which incorporates handwriting biometric data and message characteristics. Based on the heterogeneous characteristics of the multimodal signature image's low-level features and the high-level semantic correlation, a fused signature electronic document is formed, resulting in fused data. The electronic signature handwriting characteristics are used to encrypt a key using a national cryptographic symmetric encryption method, yielding an encryption key based on the original handwriting electronic signature characteristics. This encryption key is then used to encrypt and fuse the signed electronic document, resulting in a secure electronic document. This enables the flow of "multi-faceted evidence" across systems and regions with electronic documents, meeting the needs of biometric data comparison and authentication, and big data tracing in prosecutorial case handling scenarios.

[0004] However, the above method has the following problems: it fails to automatically identify user information that requires encryption based on the encryption requirements of different regions, and it fails to accurately filter user information for encryption. Summary of the Invention

[0005] To address this, the present invention provides a multi-regional data compliance adaptation engine intelligent system to overcome the problems in existing technologies that fail to automatically identify user information requiring encryption based on encryption requirements in different regions and fail to accurately perform encryption screening of user information.

[0006] To achieve the above objectives, the present invention provides a multi-regional data compliance adaptation engine intelligent system, comprising:

[0007] The information collection module is used to collect geographic information and user information. The geographic information includes user location information and geographic encryption information, and the user information includes image information and text information.

[0008] The information extraction module, which is connected to the information acquisition module, includes an image classification unit for dividing the image information into regular images and non-regular images based on an encrypted database, and a text extraction unit for scanning the text information or the text in the non-regular images to generate corresponding feature paragraphs.

[0009] The feature analysis module is connected to the information acquisition module and the information extraction module respectively, and includes an image filtering unit for filtering out the images to be encrypted from the regular images based on the user location information, and a paragraph analysis unit for comparing the feature paragraphs with preset paragraphs to generate paragraph similarity to determine sensitive information.

[0010] An information processing module, which is connected to the information acquisition module and the feature analysis module respectively, is used to combine non-sensitive information and the regional encryption information to generate a sensitive characterization value to determine whether the non-sensitive information is encrypted, or to update the encryption database.

[0011] The encryption module, which is connected to the information acquisition module, the feature analysis module, and the information processing module, includes a first encryption unit for encrypting images to be encrypted and sensitive information based on the regional encryption information, a second encryption unit for encrypting non-sensitive information determined to be encrypted, and a third encryption unit for determining non-sensitive information as non-sensitive information and encrypting it based on the boundary value analysis of the non-sensitive information.

[0012] Furthermore, the image classification unit is used to divide the image information into regular images and unregular images, wherein,

[0013] Used to scan the image information to obtain the image location and image size;

[0014] This is used to compare the image information sequentially with preset images in the encrypted database based on the image position and the image size, and determine whether the image information is a regular image or an unconventional image based on the comparison result.

[0015] Furthermore, the text extraction unit is used to scan text information or text in the unconventional image to generate several text paragraphs, wherein,

[0016] Used to generate several digital feature values ​​based on the encrypted database;

[0017] Used to scan each of the aforementioned text segments and generate corresponding numerical representation values;

[0018] This is used to determine whether the corresponding text paragraph is a feature paragraph based on the comparison result of the numerical feature value and the numerical representation value.

[0019] Furthermore, the image filtering unit is used to obtain the user location information and combine it with the encrypted database to determine a number of preset encrypted images and their corresponding blank matching degrees, and to filter out the images to be encrypted from the regular images based on the blank matching degrees.

[0020] Furthermore, the paragraph analysis unit includes:

[0021] A character distance subunit, which is connected to the text extraction unit, is used to determine several sensitive words in the feature paragraph based on the encrypted database and to calculate the character distance between adjacent sensitive words in the feature paragraph.

[0022] A sensitive information subunit, connected to the character distance subunit, is used to select several preset paragraphs that include the same number of sensitive words, and to calculate several paragraph similarities by sequentially comparing the character distance with the sensitive character distances corresponding to each preset paragraph to determine whether the feature paragraph is sensitive information.

[0023] Furthermore, the sensitivity indicator value is determined by the number of sensitive words and the encryption level of the sensitive words, wherein,

[0024] Based on the geographical encryption information, the encryption level of the sensitive words is determined to include a first encryption level, a second encryption level, and a third encryption level.

[0025] Furthermore, the information processing module compares the sensitive characterization value with a preset sensitive value, and determines whether the non-sensitive information should be encrypted based on the comparison result. The preset sensitive value is negatively correlated with the sensitive word encryption level.

[0026] Furthermore, the information processing module is used to update the encrypted database, wherein,

[0027] If it is determined that the non-sensitive information is encrypted, then it is determined that the non-sensitive information will be updated in the encrypted database.

[0028] Furthermore, the second encryption unit compares the boundary value of the encrypted non-sensitive information with a preset boundary value, and determines the reason why the feature segment corresponding to the non-sensitive information is non-sensitive information based on the comparison result, which is either a scanning problem of the information extraction module or an update of the regional encryption information.

[0029] The preset boundary value is positively correlated with the total number of characters in the non-sensitive information.

[0030] Furthermore, the information collection module matches the user location information with the encrypted regional information and calls the encrypted database corresponding to the encrypted regional information.

[0031] Compared with existing technologies, the beneficial effects of this invention are as follows: By collecting regional and user information, and retrieving corresponding regional encrypted information based on user location information, this invention classifies user image information into regular and non-regular images based on the encrypted database corresponding to the regional encrypted information. This involves classifying user image information, selecting official and valid ID documents for regular images, and other non-official and non-regular images. This allows for accurate and targeted identification and processing of image information. Since the relative positions of information in standardized and unified official ID documents are fixed, the areas requiring encryption can be directly ignored or masked without scanning, improving both the efficiency and security of subsequent encryption. This effectively enhances the efficiency of image information processing and improves the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0032] Furthermore, this invention determines whether the text information and the text in the unconventional images are encrypted by combining the numerical features of the text information and the text in the unconventional images. Since much identity information is composed of numbers, improving the sensitivity to numbers in the text plays an important role in identifying information that needs to be encrypted. At the same time, it can perform preliminary screening of the text information to be encrypted. The image screening unit combines the acquired user location information to calculate the blank matching degree of regular images in the encrypted database. By screening the user's regular images through blank matching degree, it can quickly filter out the images that need to be encrypted from the regular images, further improving the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0033] Furthermore, this invention uses an encrypted database to statistically analyze sensitive words related to encryption in the user's location, filters sensitive words in feature paragraphs, counts the number of characters in the text to determine the relative positions of each sensitive word in the feature paragraph, and determines whether the feature paragraph is sensitive information based on the encryption properties of the preset paragraph with the closest calculated character distance. By combining the number of sensitive words and the encryption level of the sensitive words, it determines whether non-sensitive information should be encrypted. This improves the accuracy of determining whether information should be encrypted and further enhances the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0034] Furthermore, when non-sensitive information is determined to require encryption, this invention updates the non-sensitive information to the encryption database, thereby improving the encryption database. For feature segments corresponding to non-sensitive information, the reasons for their non-sensitive status are analyzed, and problems in the information identification process and the failure to update encryption rules in different regions in a timely manner are rectified. During information processing, when a certain type of non-sensitive information is determined to require encryption, it is included in the encryption database for management. This operation first directly improves the content system of the encryption database and enriches the information coverage of the database. On this basis, for feature segments determined to be non-sensitive information, by deeply analyzing the specific reasons and characteristics for their non-sensitive status, the accuracy of information identification can be optimized in reverse. At the same time, targeted rectification and optimization are implemented to address deviations in the information identification process and the problem of lagging updates to encryption rules in different regions. This not only effectively ensures the stable operation of the encryption database but also significantly improves its overall security, further enhancing the accuracy of the multi-regional data compliance adaptation engine intelligent system. Attached Figure Description

[0035] Figure 1 This is a structural block diagram of the intelligent system for multi-regional data compliance adaptation engine of the present invention;

[0036] Figure 2 This is a logic diagram for determining whether a text paragraph is a characteristic paragraph in an embodiment of the present invention;

[0037] Figure 3 This is a logic diagram for determining whether a feature segment is sensitive information in an embodiment of the present invention;

[0038] Figure 4 This is a logic diagram illustrating why the feature segment corresponding to the encrypted non-sensitive information is considered non-sensitive information in an embodiment of the present invention. Detailed Implementation

[0039] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0040] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0041] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0042] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0043] Please see Figure 1 The diagram shown is a structural block diagram of the multi-regional data compliance adaptation engine intelligent system of the present invention. An embodiment of the present invention provides a multi-regional data compliance adaptation engine intelligent system, comprising:

[0044] The information collection module is used to collect geographic information and user information. Geographic information includes user location information and geographic encryption information, while user information includes image information and text information.

[0045] The information extraction module, which is connected to the information acquisition module, includes an image classification unit for classifying image information into regular and non-regular images based on an encrypted database, and a text extraction unit for scanning text information or text in non-regular images to generate corresponding feature paragraphs.

[0046] The feature analysis module is connected to the information acquisition module and the information extraction module respectively. It includes an image filtering unit for filtering out images to be encrypted from regular images based on user location information, and a paragraph analysis unit for comparing feature paragraphs with preset paragraphs to generate paragraph similarity to determine sensitive information.

[0047] The information processing module is connected to the information acquisition module and the feature analysis module respectively. It is used to combine non-sensitive information and regional encrypted information to generate sensitive characterization values ​​to determine whether non-sensitive information should be encrypted, or to update the encrypted database.

[0048] The encryption module is connected to the information acquisition module, the feature analysis module, and the information processing module, and includes a first encryption unit for encrypting images and sensitive information based on regional encryption information, a second encryption unit for encrypting non-sensitive information determined to be encrypted, and a third encryption unit for determining non-sensitive information as non-sensitive information and encrypting it based on the boundary value analysis of the non-sensitive information.

[0049] Specifically, the image classification unit is used to divide image information into regular images and non-regular images, where,

[0050] Used to scan image information to obtain image location and image size;

[0051] It is used to compare image information with preset images in an encrypted database sequentially based on image location and image size, and determine whether the image information is a regular image or an unconventional image based on the comparison results.

[0052] Understandably, the image information scanning process does not scan the content of the image, but only obtains the image position and size. The scanned image position and size are compared with preset images in the encrypted database. If the image position and size are consistent with the image position and size in the preset images in the encrypted database, the image information is determined to be a regular image; otherwise, the image information is an unconventional image.

[0053] Specifically, this invention collects regional and user information, retrieves corresponding encrypted regional information based on user location information, and classifies user image information into regular and non-regular images based on the encrypted database corresponding to the encrypted regional information. This involves categorizing user image information, selecting officially unified and valid ID documents for regular images, and other non-official and non-regular images. This allows for accurate and targeted identification and processing of image information. Because the relative positions of information in standardized and unified official ID documents are fixed, the system can directly avoid scanning or blocking the areas requiring encryption by not recognizing the information. This improves both the efficiency and security of subsequent encryption, effectively enhancing the efficiency of image information processing and improving the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0054] Please see Figure 2 As shown, this is a logic diagram for determining whether a text segment is a feature segment according to an embodiment of the present invention. The text extraction unit is used to scan text information or text in unconventional images to generate several text segments, wherein...

[0055] Used to generate several digital feature values ​​based on an encrypted database;

[0056] Used to scan each text segment and generate corresponding numerical representation values;

[0057] It is used to determine whether a corresponding text paragraph is a feature paragraph based on the comparison results of numerical feature values ​​and numerical representation values.

[0058] It is understandable that the text extraction unit scans text information or text in unconventional images to generate several text segments and performs preprocessing such as noise reduction, tilt correction and contrast enhancement to make the scanned document clearer. This is existing technology for those skilled in the art and will not be described in detail here.

[0059] Understandably, several digital feature values ​​are generated based on the encrypted database. The digital sequence to be encrypted is extracted from the encrypted database, and the number of digits and the proportion of characters occupied by the digits in the digital sequence are counted. The digital feature value = ax (number of digits) + bx (proportion of characters occupied by digits), where a and b are weighting values. Since the proportion of characters occupied by digits better reflects the characteristics of the digital sequence, a is generally taken as 0.3 and b as 0.7. Only the numerical values ​​are taken when calculating the above formula.

[0060] Understandably, scanning each text segment reveals a sequence of numbers within that segment. The number of numbers in each segment and the percentage of characters represented by those numbers are then calculated. The numerical representation value is calculated as: c x number of numbers + d x percentage of characters represented by those numbers. Here, c and d are weighted values. Since the percentage of characters represented by those numbers better reflects the characteristics of the text segment sequence, c is typically set to 0.3 and d to 0.7. Only the numerical values ​​are used in the above calculation.

[0061] Specifically, the comparison between the numerical feature value and the numerical representation value determines whether the corresponding text paragraph is a feature paragraph. If the numerical feature value and the numerical representation value are the same, then the corresponding text paragraph is determined to be a feature paragraph.

[0062] If the numerical feature value is different from the numerical representation value, then the corresponding text paragraph is determined not to be a feature paragraph.

[0063] In one specific embodiment, the encrypted database contains three sets of numeric sequences to be encrypted. The first set contains 11 digits with a character ratio of 0.84; the second set contains 10 digits with a character ratio of 1; and the third set contains 6 digits with a character ratio of 1. Let 'a' be 0.3 and 'b' be 0.7. The numeric feature value for the first set is 3.9, for the second set is 1, and for the third set is 2.5. Each text segment is then filtered to extract three sets of numeric sequences. The first set of text segment numeric sequences... The first group of text segments has 6 digits and a digit ratio of 0.5. The second group of text segments has 10 digits and a digit ratio of 1. The third group of text segments has 7 digits and a digit ratio of 1. Let c be 0.3 and d be 0.7. The numerical representation value of the first group of text segments is 2.2, the numerical representation value of the second group of text segments is 1, and the numerical representation value of the third group of text segments is 2.8. Since the numerical feature value of the second group of text segments is the same as the numerical representation value of the third group of text segments, the second group of text segments is determined to be the feature segment.

[0064] Specifically, the image filtering unit is used to obtain user location information and combine it with an encrypted database to determine several preset encrypted images and their corresponding blank matching degree, and to filter out images to be encrypted from regular images based on the blank matching degree.

[0065] It is understandable that the whitespace fit is the minimum distance (in mm) between the edge of the largest image and the nearest text in a regular image, and only the numerical value is taken when calculating.

[0066] Specifically, the blank space matching degree of a regular image is compared with the blank space matching degree of a preset encrypted image in the encrypted database. If they are the same, the regular image is determined to be an image to be encrypted; if they are different, the regular image is determined not to be an image to be encrypted.

[0067] Specifically, this invention determines whether text information and text in unconventional images are encrypted by combining numerical features of text information and text in unconventional images. Since much identity information is composed of numbers, improving the sensitivity to numbers in text plays an important role in identifying information that needs to be encrypted. At the same time, it can perform preliminary screening of text information to be encrypted. The image screening unit combines the acquired user location information to calculate the blank matching degree of regular images in the encrypted database. By screening the user's regular images based on the blank matching degree, it can quickly filter out images that need to be encrypted from regular images, further improving the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0068] Please see Figure 3 As shown, this is a logic diagram for determining whether a feature paragraph is sensitive information according to an embodiment of the present invention. The paragraph analysis unit includes:

[0069] The character distance subunit, which is connected to the text extraction unit, is used to determine several sensitive words in the feature paragraph based on the encrypted database and to calculate the character distance between adjacent sensitive words in the feature paragraph.

[0070] The sensitive information subunit, which is connected to the character distance subunit, is used to select several preset paragraphs that include the same number of sensitive words, and to calculate several paragraph similarities by sequentially comparing the character distance with the sensitive character distances corresponding to each preset paragraph to determine whether the feature paragraph is sensitive information.

[0071] Understandably, based on the encrypted database, several sensitive words in the characteristic paragraphs are identified. The character distances between adjacent sensitive words in the characteristic paragraphs are calculated, and each character distance is compiled into a character distance list. Several preset paragraphs containing the same sensitive words are selected, and the sensitive character distances between adjacent sensitive words in the preset paragraphs are calculated. Each sensitive character distance is compiled into a sensitive character distance list. The similarity between the character distance list and each sensitive character distance list is calculated. The paragraph similarity MP between the character distance list YB=(YB1, YB2, ..., YBj, ..., YBm) and the sensitive character distance list EB=(EB1, EB2, ..., EBj, ..., EBm) is calculated; where j=1, 2, ..., m; paragraph similarity: MP=∑mj=1YBj*EBj / (sqrt(∑mj=1(YBj)2)*sqrt(∑mj=1(EBj)2)).

[0072] Specifically, the character distance is sequentially calculated and compared with the distance of sensitive characters corresponding to each preset paragraph to generate several paragraph similarities. These similarities are then compared with preset similarities, and the comparison results are used to determine whether the feature paragraph contains sensitive information.

[0073] If the paragraph similarity is greater than or equal to the preset similarity, the feature paragraph is determined to be sensitive information;

[0074] If the paragraph similarity is less than the preset similarity, the feature paragraph is determined to be non-sensitive information;

[0075] In one specific embodiment, a preset similarity of 0.9 is set. If the paragraph similarity of 0.94 is greater than the preset similarity, the feature paragraph is determined to be sensitive information.

[0076] If the paragraph similarity is 0.72, which is less than the preset similarity, the feature paragraph is determined to be non-sensitive information.

[0077] It is understandable that the preset similarity is negatively correlated with the number of sensitive words contained in the feature paragraph. This is because the more sensitive words there are, the greater the character spacing, and the larger the variation in character spacing values. Therefore, the preset similarity is negatively correlated with the number of sensitive words contained in the feature paragraph. Preferably, the preset similarity value ranges from 0.8 to 0.9.

[0078] Specifically, the sensitivity indicator value is determined by the number of sensitive words and the encryption level of those words, where,

[0079] The encryption level of sensitive words is determined based on the regional encryption information, including the first encryption level, the second encryption level, and the third encryption level. The encryption level of the first encryption level is higher than that of the second encryption level, and the encryption level of the second encryption level is higher than that of the third encryption level.

[0080] It is understandable that the sensitivity index = α x number of sensitive words + β, where β is the weighted value corresponding to different encryption levels of sensitive words, α is the weighted value of the number of sensitive words, and α + β = 1. Since the encryption level of sensitive words has a significant impact on the sensitivity of sensitive information, different encryption levels of sensitive words correspond to different weighted values. Generally, the first encryption level β is 0.8, the second encryption level β is 0.7, and the third encryption level β is 0.6.

[0081] Specifically, the information processing module compares the sensitive characterization value with the preset sensitive value, and determines whether non-sensitive information should be encrypted based on the comparison result. The preset sensitive value is negatively correlated with the encryption level of sensitive words.

[0082] Understandably, given the same number of sensitive words, the higher the encryption level, the greater the impact of the number of sensitive words on the sensitivity indicator value. Therefore, the preset sensitivity value is negatively correlated with the encryption level. Preferably, the preset sensitivity value is 1.2 when the encryption level is the first level; 1.3 when the encryption level is the second level; and 1.4 when the encryption level is the third level.

[0083] Specifically, if the sensitive characterization value is greater than or equal to the preset sensitive value, then non-sensitive information is determined to be encrypted;

[0084] If the sensitivity value is less than the preset sensitivity value, then non-sensitive information will not be encrypted.

[0085] In one specific embodiment, a preset sensitivity value of 1.2 is set. If the sensitivity value of 1.6 is greater than the preset sensitivity value, then non-sensitive information is determined to be encrypted.

[0086] If the sensitivity value is 1.0, which is less than the preset sensitivity value, then non-sensitive information will not be encrypted.

[0087] Specifically, this invention uses an encrypted database to statistically analyze sensitive words related to encryption in the user's region, filters sensitive words in characteristic paragraphs, counts the number of characters in the text to determine the relative positions of each sensitive word in the characteristic paragraph, determines whether the characteristic paragraph is sensitive information based on the encryption properties of the preset paragraph with the closest character distance, and determines whether non-sensitive information should be encrypted by combining the number of sensitive words and the encryption level of the sensitive words. This improves the accuracy of determining whether information should be encrypted and further enhances the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0088] Specifically, the information processing module is used to update the encrypted database, wherein...

[0089] If non-sensitive information is determined to be encrypted, then the non-sensitive information will be updated in the encrypted database.

[0090] Please see Figure 4 As shown, it is a logic diagram of the reason why the feature segment corresponding to the encrypted non-sensitive information is non-sensitive information in an embodiment of the present invention. The second encryption unit compares the boundary value of the encrypted non-sensitive information with the preset boundary value. According to the comparison result, it determines that the reason why the feature segment corresponding to the encrypted non-sensitive information is non-sensitive information is a scanning problem of the information extraction module or an update of the regional encryption information.

[0091] Specifically, the boundary value for non-sensitive information is the percentage of blank areas within the total area of ​​the non-sensitive information. Updates to geo-encrypted information require precise selection of specific text during the conversion of non-sensitive information into preset sensitive information, leading to an increase in the percentage of blank areas and thus a larger preset boundary value. Conversely, issues such as scanning angle deviations during the information extraction module's scanning process can reduce the percentage of blank areas.

[0092] Specifically, if the boundary value of non-sensitive information is less than or equal to the preset boundary value, the reason for determining that the feature segment is non-sensitive information is that the regional encryption information has been updated.

[0093] If the boundary value of non-sensitive information is greater than the preset boundary value, the reason for determining that the feature paragraph is non-sensitive information is a scanning problem of the information extraction module;

[0094] In one specific embodiment, a preset boundary value of 70% is set. If the boundary value of non-sensitive information is 52%, which is less than the preset boundary value, the reason why the feature segment is determined to be non-sensitive information is that the regional encryption information has been updated.

[0095] If the boundary value of non-sensitive information is 78%, which is greater than the preset boundary value, then the reason why the feature paragraph is determined to be non-sensitive information is a scanning problem of the information extraction module.

[0096] Among them, the preset boundary value is positively correlated with the total number of characters in non-sensitive information.

[0097] It is understandable that the greater the total number of characters in non-sensitive information, the higher the probability that it contains information that needs to be encrypted. Therefore, the preset boundary value is positively correlated with the total number of characters in non-sensitive information. Preferably, the preset boundary value ranges from 60% to 80%.

[0098] Specifically, the information collection module matches user location information with encrypted geographic information and calls the encrypted database corresponding to the encrypted geographic information.

[0099] Specifically, when this invention determines that non-sensitive information needs encryption, it updates the non-sensitive information to the encryption database, thus improving the encryption database. For feature segments corresponding to non-sensitive information, it analyzes the reasons why they are non-sensitive and rectifys problems in the information identification process and the failure to update encryption rules in different regions in a timely manner. During information processing, when a certain type of non-sensitive information is determined to require encryption, it is included in the encryption database for management. This operation first directly improves the content system of the encryption database and enriches the information coverage of the database. On this basis, for feature segments determined to be non-sensitive information, by deeply analyzing the specific reasons and characteristics for their non-sensitivity, the accuracy of information identification can be optimized in reverse. At the same time, targeted rectification and optimization are implemented to address deviations in the information identification process and the problem of lagging updates to encryption rules in different regions. This not only effectively ensures the stable operation of the encryption database but also significantly improves its overall security and further enhances the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0100] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-territory data compliance adaptation engine intelligent system, characterized in that, The application relates to a sensitive information encryption method and device. The information collection module is used for collecting regional information and user information, the regional information comprises user position information and regional encryption information, and the user information comprises picture information and text information; The information extraction module is connected with the information collection module and comprises a picture classification unit used for classifying the picture information into regular pictures and irregular pictures based on an encryption database, and a text extraction unit used for scanning the text information or the text in the irregular pictures to generate corresponding feature paragraphs; The feature analysis module is connected with the information collection module and the information extraction module and comprises a picture screening unit used for screening out pictures to be encrypted in the regular pictures based on the user position information, and a paragraph analysis unit used for comparing the feature paragraphs with preset paragraphs to generate paragraph similarity to determine sensitive information; The information processing module is connected with the information collection module and the feature analysis module and used for combining non-sensitive information and the regional encryption information to generate sensitive representation values to determine whether the non-sensitive information is encrypted or not, or updating the encryption database; The encryption module is connected with the information collection module, the feature analysis module and the information processing module and comprises a first encryption unit used for encrypting pictures to be encrypted and sensitive information based on the regional encryption information, and a second encryption unit used for encrypting non-sensitive information determined to be encrypted and determining corresponding feature paragraphs as non-sensitive information based on boundary values of the non-sensitive information and reasons for encryption; The information collection module corresponds the user position information and the regional encryption information and calls the encryption database corresponding to the regional encryption information.

2. The multi-territory data compliance adaptation engine intelligent system of claim 1, wherein, The picture classification unit is used for classifying the picture information into regular pictures and irregular pictures, wherein, image positions and image sizes are obtained by scanning the picture information; the picture information is compared with preset pictures in the encryption database based on the image positions and the image sizes, and the picture information is determined as the regular pictures or the irregular pictures according to comparison results.

3. The multi-territory data compliance adaptation engine intelligent system of claim 2, wherein, The text extraction unit is used for scanning text information or text in the irregular pictures to generate a plurality of text paragraphs, wherein, a plurality of digital feature values are generated based on the encryption database; a corresponding digital representation value is generated by scanning each text paragraph; whether a corresponding text paragraph is a feature paragraph is determined according to comparison results of the digital feature values and the digital representation values.

4. The multi-territory data compliance adaptation engine intelligent system of claim 2, wherein, The picture screening unit is used for obtaining the user position information, combining the encryption database to determine a plurality of preset encryption pictures and corresponding blank fitting degrees, and screening out pictures to be encrypted in the regular pictures based on the blank fitting degrees.

5. The multi-territory data compliance adaptation engine intelligent system of claim 3, wherein, The paragraph analysis unit comprises: a character distance subunit connected with the text extraction unit and used for determining a plurality of sensitive words in the feature paragraphs based on the encryption database, and counting character distances between adjacent sensitive words in the feature paragraphs. The sensitive information subunit is connected with the character distance subunit, and is configured to select a plurality of preset paragraphs including the same number of sensitive words, and generate a plurality of paragraph similarities by sequentially calculating the character distance and sensitive character distance corresponding to each of the preset paragraphs to determine whether the feature paragraph is sensitive information.

6. The multi-region data compliance adaptation engine intelligent system of claim 5, wherein, The sensitive representation value is determined by the number of sensitive words and the encryption level of the sensitive words. The sensitive word encryption level determined based on the regional encryption information includes a first encryption level, a second encryption level, and a third encryption level.

7. The multi-region data compliance adaptation engine intelligent system of claim 6, wherein, The information processing module compares the sensitive representation value with a preset sensitive value, and determines whether the non-sensitive information is encrypted according to a comparison result. The preset sensitive value is negatively correlated with the sensitive word encryption level.

8. The multi-territory data compliance adaptation engine intelligent system of claim 7, wherein, The information processing module is configured to update the encryption database. If it is determined that the non-sensitive information is encrypted, it is determined that the non-sensitive information is updated to the encryption database.

9. The multi-territory data compliance adaptation engine intelligent system of claim 8, wherein, The second encryption unit compares the boundary value of the non-sensitive information to be encrypted with a preset boundary value, and determines that the feature paragraph corresponding to the non-sensitive information is non-sensitive information because the information extraction module scans the question or the regional encryption information is updated according to a comparison result. The preset boundary value is positively correlated with the total number of characters of the non-sensitive information.

Citation Information

Patent Citations

  • Method, system and device for generating and decrypting secure electronic file and medium

    CN116108502A

  • Network database user information encryption system and method

    CN115618398A

  • Information security management and monitoring system based on big data

    CN120408711A