Multi-region data compliance adaptation engine intelligent system

By classifying and analyzing user information through a multi-regional data compliance adaptation engine intelligent system, the problem of automatic identification and screening of encryption requirements in different regions is solved, achieving efficient and secure encryption processing and improving the accuracy and security of the system.

CN121167752AActive Publication Date: 2025-12-19ZHONGHE YUNKE INFORMATION TECH GRP CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511188254.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-12-19
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing technologies fail to automatically identify user information requiring encryption based on the encryption requirements of different regions, and fail to accurately filter user information for encryption.

Method used

This invention provides a multi-regional data compliance adaptation engine intelligent system, including an information collection module, an information extraction module, a feature analysis module, an information processing module, and an encryption module. By collecting regional information and user information, the system uses an encrypted database to classify and analyze the features of image and text information, filters out sensitive information to be encrypted, and performs encryption processing.

Benefits of technology

It improves the efficiency and security of encryption, enhances the accurate identification and processing of user information, and improves the accuracy and security of the multi-regional data compliance adaptation engine intelligent system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167752A_ABST
    Figure CN121167752A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a multi-region data compliance adaptation engine intelligent system. The system comprises an information acquisition module used for acquiring regional information and user information; the information extraction module comprises a picture classification unit and a character extraction unit; the feature analysis module comprises a picture screening unit used for screening pictures to be encrypted and a paragraph analysis unit used for judging sensitive information; the information processing module is used for judging whether the non-sensitive information is encrypted or not, or updating the encrypted database; and the encryption module comprises a first encryption unit used for encrypting the to-be-encrypted picture and the sensitive information, and a second encryption unit used for encrypting the non-sensitive information judged to be encrypted. According to the invention, the above modules cooperate with each other, so that the accuracy of the multi-region data compliance adaptation engine intelligent system is further improved while the information identification efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a multi-region data compliance adaptation engine intelligent system. BACKGROUND

[0002] In recent years, a large number of companies have begun to develop overseas business, and the data compliance and security challenges and complexity brought by the overseas countries all over the world have risen sharply. If a new solution is developed for each new country, it will seriously hinder the company's overseas progress and bring great complexity to the system. By setting up a solution to store the data compliance requirements of each country through configuration, and selecting the appropriate solution through dynamic routing when using, the system does not need to pay attention to compliance issues, so as to focus on more valuable work such as the rapid implementation of system business, and can encrypt the personal information according to the personal sensitive data privacy regulations of different countries / regions.

[0003] Chinese patent application publication No. CN116108502A discloses a secure electronic file generation and decryption method, system, device and medium. The application provides a secure electronic file generation method, which obtains unified format electronic signature data by identifying and authenticating electronic signature related data; obtains a signature picture with increased handwriting biological feature data and text feature data according to the signature base data, biological feature information, auxiliary data, business addition, and data usage image steganography of the whole signature process; forms a signature electronic file based on the multi-modal signature picture bottom feature heterogeneity and high-level semantic correlation characteristics to obtain fusion data; encrypts the key of the national secret symmetric encryption using the electronic signature handwriting feature to obtain an encryption key based on the original handwriting electronic signature feature; and encrypts and fuses the signature electronic file using the encryption key to obtain a secure electronic file. The "multi-evidence" can be transferred with the electronic file across systems and regions, meeting the needs of biological feature data comparison and identification and big data traceability in the prosecution scene.

[0004] However, the above method has the following problems: it cannot automatically identify user information with encryption needs according to encryption requirements in different regions, and it cannot accurately encrypt and screen user information. SUMMARY

[0005] Therefore, the present application provides a multi-region data compliance adaptation engine intelligent system to overcome the problem that the prior art cannot automatically identify user information with encryption needs according to encryption requirements in different regions, and cannot accurately encrypt and screen user information.

[0006] To achieve the above purpose, the present application provides a multi-region data compliance adaptation engine intelligent system, comprising:

[0007] The information collection module is used to collect geographic information and user information. The geographic information includes user location information and geographic encryption information, and the user information includes image information and text information.

[0008] The information extraction module, which is connected to the information acquisition module, includes an image classification unit for dividing the image information into regular images and non-regular images based on an encrypted database, and a text extraction unit for scanning the text information or the text in the non-regular images to generate corresponding feature paragraphs.

[0009] The feature analysis module is connected to the information acquisition module and the information extraction module respectively, and includes an image filtering unit for filtering out the images to be encrypted from the regular images based on the user location information, and a paragraph analysis unit for comparing the feature paragraphs with preset paragraphs to generate paragraph similarity to determine sensitive information.

[0010] An information processing module, which is connected to the information acquisition module and the feature analysis module respectively, is used to combine non-sensitive information and the regional encryption information to generate a sensitive characterization value to determine whether the non-sensitive information is encrypted, or to update the encryption database.

[0011] The encryption module, which is connected to the information acquisition module, the feature analysis module, and the information processing module, includes a first encryption unit for encrypting images to be encrypted and sensitive information based on the regional encryption information, a second encryption unit for encrypting non-sensitive information determined to be encrypted, and a third encryption unit for determining non-sensitive information as non-sensitive information and encrypting it based on the boundary value analysis of the non-sensitive information.

[0012] Furthermore, the image classification unit is used to divide the image information into regular images and unregular images, wherein,

[0013] Used to scan the image information to obtain the image location and image size;

[0014] This is used to compare the image information sequentially with preset images in the encrypted database based on the image position and the image size, and determine whether the image information is a regular image or an unconventional image based on the comparison result.

[0015] Furthermore, the text extraction unit is used to scan text information or text in the unconventional image to generate several text paragraphs, wherein,

[0016] Used to generate several digital feature values ​​based on the encrypted database;

[0017] Used to scan each of the aforementioned text segments and generate corresponding numerical representation values;

[0018] This is used to determine whether the corresponding text paragraph is a feature paragraph based on the comparison result of the numerical feature value and the numerical representation value.

[0019] Furthermore, the image filtering unit is used to obtain the user location information and combine it with the encrypted database to determine a number of preset encrypted images and their corresponding blank matching degrees, and to filter out the images to be encrypted from the regular images based on the blank matching degrees.

[0020] Furthermore, the paragraph analysis unit includes:

[0021] A character distance subunit, which is connected to the text extraction unit, is used to determine several sensitive words in the feature paragraph based on the encrypted database and to calculate the character distance between adjacent sensitive words in the feature paragraph.

[0022] A sensitive information subunit, connected to the character distance subunit, is used to select several preset paragraphs that include the same number of sensitive words, and to calculate several paragraph similarities by sequentially comparing the character distance with the sensitive character distances corresponding to each preset paragraph to determine whether the feature paragraph is sensitive information.

[0023] Furthermore, the sensitivity indicator value is determined by the number of sensitive words and the encryption level of the sensitive words, wherein,

[0024] Based on the geographical encryption information, the encryption level of the sensitive words is determined to include a first encryption level, a second encryption level, and a third encryption level.

[0025] Furthermore, the information processing module compares the sensitive characterization value with a preset sensitive value, and determines whether the non-sensitive information should be encrypted based on the comparison result. The preset sensitive value is negatively correlated with the sensitive word encryption level.

[0026] Furthermore, the information processing module is used to update the encrypted database, wherein,

[0027] If it is determined that the non-sensitive information is encrypted, then it is determined that the non-sensitive information will be updated in the encrypted database.

[0028] Furthermore, the second encryption unit compares the boundary value of the encrypted non-sensitive information with a preset boundary value, and determines the reason why the feature segment corresponding to the non-sensitive information is non-sensitive information based on the comparison result, which is either a scanning problem of the information extraction module or an update of the regional encryption information.

[0029] The preset boundary value is positively correlated with the total number of characters in the non-sensitive information.

[0030] Furthermore, the information collection module matches the user location information with the encrypted regional information and calls the encrypted database corresponding to the encrypted regional information.

[0031] Compared with existing technologies, the beneficial effects of this invention are as follows: By collecting regional and user information, and retrieving corresponding regional encrypted information based on user location information, this invention classifies user image information into regular and non-regular images based on the encrypted database corresponding to the regional encrypted information. This involves classifying user image information, selecting official and valid ID documents for regular images, and other non-official and non-regular images. This allows for accurate and targeted identification and processing of image information. Since the relative positions of information in standardized and unified official ID documents are fixed, the areas requiring encryption can be directly ignored or masked without scanning, improving both the efficiency and security of subsequent encryption. This effectively enhances the efficiency of image information processing and improves the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0032] Furthermore, this invention determines whether the text information and the text in the unconventional images are encrypted by combining the numerical features of the text information and the text in the unconventional images. Since much identity information is composed of numbers, improving the sensitivity to numbers in the text plays an important role in identifying information that needs to be encrypted. At the same time, it can perform preliminary screening of the text information to be encrypted. The image screening unit combines the acquired user location information to calculate the blank matching degree of regular images in the encrypted database. By screening the user's regular images through blank matching degree, it can quickly filter out the images that need to be encrypted from the regular images, further improving the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0033] Furthermore, this invention uses an encrypted database to statistically analyze sensitive words related to encryption in the user's location, filters sensitive words in feature paragraphs, counts the number of characters in the text to determine the relative positions of each sensitive word in the feature paragraph, and determines whether the feature paragraph is sensitive information based on the encryption properties of the preset paragraph with the closest calculated character distance. By combining the number of sensitive words and the encryption level of the sensitive words, it determines whether non-sensitive information should be encrypted. This improves the accuracy of determining whether information should be encrypted and further enhances the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0034] Furthermore, when non-sensitive information is determined to require encryption, this invention updates the non-sensitive information to the encryption database, thereby improving the encryption database. For feature segments corresponding to non-sensitive information, the reasons for their non-sensitive status are analyzed, and problems in the information identification process and the failure to update encryption rules in different regions in a timely manner are rectified. During information processing, when a certain type of non-sensitive information is determined to require encryption, it is included in the encryption database for management. This operation first directly improves the content system of the encryption database and enriches the information coverage of the database. On this basis, for feature segments determined to be non-sensitive information, by deeply analyzing the specific reasons and characteristics for their non-sensitive status, the accuracy of information identification can be optimized in reverse. At the same time, targeted rectification and optimization are implemented to address deviations in the information identification process and the problem of lagging updates to encryption rules in different regions. This not only effectively ensures the stable operation of the encryption database but also significantly improves its overall security, further enhancing the accuracy of the multi-regional data compliance adaptation engine intelligent system. Attached Figure Description

[0035] Figure 1 This is a structural block diagram of the intelligent system for multi-regional data compliance adaptation engine of the present invention;

[0036] Figure 2 This is a logic diagram for determining whether a text paragraph is a characteristic paragraph in an embodiment of the present invention;

[0037] Figure 3 This is a logic diagram for determining whether a feature segment is sensitive information in an embodiment of the present invention;

[0038] Figure 4 This is a logic diagram illustrating why the feature segment corresponding to the encrypted non-sensitive information is considered non-sensitive information in an embodiment of the present invention. Detailed Implementation

[0039] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0040] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0041] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0042] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0043] Please see Figure 1 The diagram shown is a structural block diagram of the multi-regional data compliance adaptation engine intelligent system of the present invention. An embodiment of the present invention provides a multi-regional data compliance adaptation engine intelligent system, comprising:

[0044] The information collection module is used to collect geographic information and user information. Geographic information includes user location information and geographic encryption information, while user information includes image information and text information.

[0045] The information extraction module, which is connected to the information acquisition module, includes an image classification unit for classifying image information into regular and non-regular images based on an encrypted database, and a text extraction unit for scanning text information or text in non-regular images to generate corresponding feature paragraphs.

[0046] The feature analysis module is connected to the information acquisition module and the information extraction module respectively. It includes an image filtering unit for filtering out images to be encrypted from regular images based on user location information, and a paragraph analysis unit for comparing feature paragraphs with preset paragraphs to generate paragraph similarity to determine sensitive information.

[0047] The information processing module is connected to the information acquisition module and the feature analysis module respectively. It is used to combine non-sensitive information and regional encrypted information to generate sensitive characterization values ​​to determine whether non-sensitive information should be encrypted, or to update the encrypted database.

[0048] The encryption module is connected to the information acquisition module, the feature analysis module, and the information processing module, and includes a first encryption unit for encrypting images and sensitive information based on regional encryption information, a second encryption unit for encrypting non-sensitive information determined to be encrypted, and a third encryption unit for determining non-sensitive information as non-sensitive information and encrypting it based on the boundary value analysis of the non-sensitive information.

[0049] Specifically, the image classification unit is used to divide image information into regular images and non-regular images, where,

[0050] Used to scan image information to obtain image location and image size;

[0051] It is used to compare image information with preset images in an encrypted database sequentially based on image location and image size, and determine whether the image information is a regular image or an unconventional image based on the comparison results.

[0052] Understandably, the image information scanning process does not scan the content of the image, but only obtains the image position and size. The scanned image position and size are compared with preset images in the encrypted database. If the image position and size are consistent with the image position and size in the preset images in the encrypted database, the image information is determined to be a regular image; otherwise, the image information is an unconventional image.

[0053] Specifically, this invention collects regional and user information, retrieves corresponding encrypted regional information based on user location information, and classifies user image information into regular and non-regular images based on the encrypted database corresponding to the encrypted regional information. This involves categorizing user image information, selecting officially unified and valid ID documents for regular images, and other non-official and non-regular images. This allows for accurate and targeted identification and processing of image information. Because the relative positions of information in standardized and unified official ID documents are fixed, the system can directly avoid scanning or blocking the areas requiring encryption by not recognizing the information. This improves both the efficiency and security of subsequent encryption, effectively enhancing the efficiency of image information processing and improving the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0054] Please see Figure 2 As shown, this is a logic diagram for determining whether a text segment is a feature segment according to an embodiment of the present invention. The text extraction unit is used to scan text information or text in unconventional images to generate several text segments, wherein...

[0055] Used to generate several digital feature values ​​based on an encrypted database;

[0056] Used to scan each text segment and generate corresponding numerical representation values;

[0057] It is used to determine whether a corresponding text paragraph is a feature paragraph based on the comparison results of numerical feature values ​​and numerical representation values.

[0058] It is understandable that the text extraction unit scans text information or text in unconventional images to generate several text segments and performs preprocessing such as noise reduction, tilt correction and contrast enhancement to make the scanned document clearer. This is existing technology for those skilled in the art and will not be described in detail here.

[0059] Understandably, several digital feature values ​​are generated based on the encrypted database. The digital sequence to be encrypted is extracted from the encrypted database, and the number of digits and the proportion of characters occupied by the digits in the digital sequence are counted. The digital feature value = ax number of digits + bx proportion of characters occupied by digits, where a and b are weighting values. Since the proportion of characters occupied by digits better reflects the characteristics of the digital sequence, a is generally taken as 0.3 and b as 0.7. Only the numerical values ​​are taken when calculating the above formula.

[0060] Understandably, scanning each text segment reveals a sequence of numbers within that segment. The number of numbers in each segment and the percentage of characters represented by those numbers are then calculated. The numerical representation value is calculated as: c x number of numbers + d x percentage of characters represented by those numbers. Here, c and d are weighted values. Since the percentage of characters represented by those numbers better reflects the characteristics of the text segment sequence, c is typically set to 0.3 and d to 0.7. Only the numerical values ​​are used in the above calculation.

[0061] Specifically, the comparison between the numerical feature value and the numerical representation value determines whether the corresponding text paragraph is a feature paragraph. If the numerical feature value and the numerical representation value are the same, then the corresponding text paragraph is determined to be a feature paragraph.

[0062] If the numerical feature value is different from the numerical representation value, then the corresponding text paragraph is determined not to be a feature paragraph.

[0063] In one specific embodiment, the encrypted database contains three sets of numeric sequences to be encrypted. The first set contains 11 digits with a character ratio of 0.84; the second set contains 10 digits with a character ratio of 1; and the third set contains 6 digits with a character ratio of 1. Let 'a' be 0.3 and 'b' be 0.7. The numeric feature value for the first set is 3.9, for the second set is 1, and for the third set is 2.5. Each text segment is then filtered to extract three sets of numeric sequences. The first set of text segment numeric sequences... The first group of text segments has 6 digits and a digit ratio of 0.5. The second group of text segments has 10 digits and a digit ratio of 1. The third group of text segments has 7 digits and a digit ratio of 1. Let c be 0.3 and d be 0.7. The numerical representation value of the first group of text segments is 2.2, the numerical representation value of the second group of text segments is 1, and the numerical representation value of the third group of text segments is 2.8. Since the numerical feature value of the second group of text segments is the same as the numerical representation value of the third group of text segments, the second group of text segments is determined to be the feature segment.

[0064] Specifically, the image filtering unit is used to obtain user location information and combine it with an encrypted database to determine several preset encrypted images and their corresponding blank matching degree, and to filter out images to be encrypted from regular images based on the blank matching degree.

[0065] It is understandable that the whitespace fit is the minimum distance (in mm) between the edge of the largest image and the nearest text in a regular image, and only the numerical value is taken when calculating.

[0066] Specifically, the blank space matching degree of a regular image is compared with the blank space matching degree of a preset encrypted image in the encrypted database. If they are the same, the regular image is determined to be an image to be encrypted; if they are different, the regular image is determined not to be an image to be encrypted.

[0067] Specifically, this invention determines whether text information and text in unconventional images are encrypted by combining numerical features of text information and text in unconventional images. Since much identity information is composed of numbers, improving the sensitivity to numbers in text plays an important role in identifying information that needs to be encrypted. At the same time, it can perform preliminary screening of text information to be encrypted. The image screening unit combines the acquired user location information to calculate the blank matching degree of regular images in the encrypted database. By screening the user's regular images based on the blank matching degree, it can quickly filter out images that need to be encrypted from regular images, further improving the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0068] Please seeFigure 3 As shown, this is a logic diagram for determining whether a feature paragraph is sensitive information according to an embodiment of the present invention. The paragraph analysis unit includes:

[0069] The character distance subunit, which is connected to the text extraction unit, is used to determine several sensitive words in the feature paragraph based on the encrypted database and to calculate the character distance between adjacent sensitive words in the feature paragraph.

[0070] The sensitive information subunit, which is connected to the character distance subunit, is used to select several preset paragraphs that include the same number of sensitive words, and to calculate several paragraph similarities by sequentially comparing the character distance with the sensitive character distances corresponding to each preset paragraph to determine whether the feature paragraph is sensitive information.

[0071] Understandably, based on the encrypted database, several sensitive words in the characteristic paragraphs are identified. The character distances between adjacent sensitive words in the characteristic paragraphs are counted, and each character distance is compiled into a character distance list. Several preset paragraphs containing the same sensitive words are selected, and the sensitive character distances between adjacent sensitive words in the preset paragraphs are counted. Each sensitive character distance is compiled into a sensitive character distance list. The similarity between the character distance list and each sensitive character distance list is calculated. The paragraph similarity MP between the character distance list YB = (YB1, YB2, ..., YBj, ..., YBm) and the sensitive character distance list EB = (EB1, EB2, ..., EBj, ..., EBm) is calculated; where j = 1, 2, ..., m; paragraph similarity: MP = ∑mj = 1YBj * EBj / (sqrt(∑mj = 1(YBj)2) * sqrt(∑mj = 1(EBj)2)).

[0072] Specifically, the character distance is sequentially calculated and compared with the distance of sensitive characters corresponding to each preset paragraph to generate several paragraph similarities. These similarities are then compared with preset similarities, and the comparison results are used to determine whether the feature paragraph contains sensitive information.

[0073] If the paragraph similarity is greater than or equal to the preset similarity, the feature paragraph is determined to be sensitive information;

[0074] If the paragraph similarity is less than the preset similarity, the feature paragraph is determined to be non-sensitive information;

[0075] In one specific embodiment, a preset similarity of 0.9 is set. If the paragraph similarity of 0.94 is greater than the preset similarity, the feature paragraph is determined to be sensitive information.

[0076] If the paragraph similarity is 0.72, which is less than the preset similarity, the feature paragraph is determined to be non-sensitive information.

[0077] It is understandable that the preset similarity is negatively correlated with the number of sensitive words contained in the feature paragraph. This is because the more sensitive words there are, the greater the character spacing, and the larger the variation in character spacing values. Therefore, the preset similarity is negatively correlated with the number of sensitive words contained in the feature paragraph. Preferably, the preset similarity value ranges from 0.8 to 0.9.

[0078] Specifically, the sensitivity indicator value is determined by the number of sensitive words and the encryption level of those words, where,

[0079] The encryption level of sensitive words is determined based on the regional encryption information, including the first encryption level, the second encryption level, and the third encryption level. The encryption level of the first encryption level is higher than that of the second encryption level, and the encryption level of the second encryption level is higher than that of the third encryption level.

[0080] It is understandable that the sensitivity index = α x number of sensitive words + β, where β is the weighted value corresponding to different encryption levels of sensitive words, α is the weighted value of the number of sensitive words, and α + β = 1. Since the encryption level of sensitive words has a significant impact on the sensitivity of sensitive information, different encryption levels of sensitive words correspond to different weighted values. Generally, the first encryption level β is 0.8, the second encryption level β is 0.7, and the third encryption level β is 0.6.

[0081] Specifically, the information processing module compares the sensitive characterization value with the preset sensitive value, and determines whether non-sensitive information should be encrypted based on the comparison result. The preset sensitive value is negatively correlated with the encryption level of sensitive words.

[0082] Understandably, given the same number of sensitive words, the higher the encryption level, the greater the impact of the number of sensitive words on the sensitivity indicator value. Therefore, the preset sensitivity value is negatively correlated with the encryption level. Preferably, the preset sensitivity value is 1.2 when the encryption level is the first level; 1.3 when the encryption level is the second level; and 1.4 when the encryption level is the third level.

[0083] Specifically, if the sensitive characterization value is greater than or equal to the preset sensitive value, then non-sensitive information is determined to be encrypted;

[0084] If the sensitivity value is less than the preset sensitivity value, then non-sensitive information is determined to be encrypted.

[0085] In one specific embodiment, a preset sensitivity value of 1.2 is set. If the sensitivity value of 1.6 is greater than the preset sensitivity value, then non-sensitive information is determined to be encrypted.

[0086] If the sensitivity value is 1.0, which is less than the preset sensitivity value, then non-sensitive information is determined to be encrypted.

[0087] Specifically, this invention uses an encrypted database to statistically analyze sensitive words related to encryption in the user's region, filters sensitive words in characteristic paragraphs, counts the number of characters in the text to determine the relative positions of each sensitive word in the characteristic paragraph, determines whether the characteristic paragraph is sensitive information based on the encryption properties of the preset paragraph with the closest character distance, and determines whether non-sensitive information should be encrypted by combining the number of sensitive words and the encryption level of the sensitive words. This improves the accuracy of determining whether information should be encrypted and further enhances the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0088] Specifically, the information processing module is used to update the encrypted database, wherein...

[0089] If non-sensitive information is determined to be encrypted, then the non-sensitive information will be updated in the encrypted database.

[0090] Please see Figure 4 As shown, it is a logic diagram of the reason why the feature segment corresponding to the encrypted non-sensitive information is non-sensitive information in an embodiment of the present invention. The second encryption unit compares the boundary value of the encrypted non-sensitive information with the preset boundary value. According to the comparison result, it determines that the reason why the feature segment corresponding to the encrypted non-sensitive information is non-sensitive information is a scanning problem of the information extraction module or an update of the regional encryption information.

[0091] Specifically, the boundary value for non-sensitive information is the percentage of blank areas within the total area of ​​the non-sensitive information. Updates to geo-encrypted information require precise selection of specific text during the conversion of non-sensitive information into preset sensitive information, leading to an increase in the percentage of blank areas and thus a larger preset boundary value. Conversely, issues such as scanning angle deviations during the information extraction module's scanning process can reduce the percentage of blank areas.

[0092] Specifically, if the boundary value of non-sensitive information is less than or equal to the preset boundary value, the reason for determining that the feature segment is non-sensitive information is that the regional encryption information has been updated.

[0093] If the boundary value of non-sensitive information is greater than the preset boundary value, the reason for determining that the feature paragraph is non-sensitive information is a scanning problem of the information extraction module;

[0094] In one specific embodiment, a preset boundary value of 70% is set. If the boundary value of non-sensitive information is 52%, which is less than the preset boundary value, the reason for determining that the feature segment is non-sensitive information is that the regional encryption information has been updated.

[0095] If the boundary value of non-sensitive information is 78%, which is greater than the preset boundary value, then the reason why the feature paragraph is determined to be non-sensitive information is a scanning problem of the information extraction module.

[0096] Among them, the preset boundary value is positively correlated with the total number of characters in non-sensitive information.

[0097] It is understandable that the higher the total number of characters in non-sensitive information, the greater the probability that it contains information that needs to be encrypted. Therefore, the preset boundary value is positively correlated with the total number of characters in non-sensitive information. Preferably, the preset boundary value ranges from 60% to 80%.

[0098] Specifically, the information collection module matches user location information with encrypted geographic information and calls the encrypted database corresponding to the encrypted geographic information.

[0099] Specifically, when this invention determines that non-sensitive information needs encryption, it updates the non-sensitive information to the encryption database, thus improving the encryption database. For feature segments corresponding to non-sensitive information, it analyzes the reasons why they are non-sensitive and rectifys problems in the information identification process and the failure to update encryption rules in different regions in a timely manner. During information processing, when a certain type of non-sensitive information is determined to require encryption, it is included in the encryption database for management. This operation first directly improves the content system of the encryption database and enriches the information coverage of the database. On this basis, for feature segments determined to be non-sensitive information, by deeply analyzing the specific reasons and characteristics for their non-sensitivity, the accuracy of information identification can be optimized in reverse. At the same time, targeted rectification and optimization are implemented to address deviations in the information identification process and the problem of lagging updates to encryption rules in different regions. This not only effectively ensures the stable operation of the encryption database but also significantly improves its overall security and further enhances the accuracy of the multi-regional data compliance adaptation engine intelligent system.

[0100] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-regional data compliance adaptation engine intelligent system, characterized in that, include: The information collection module is used to collect geographic information and user information. The geographic information includes user location information and geographic encryption information, and the user information includes image information and text information. The information extraction module, which is connected to the information acquisition module, includes an image classification unit for dividing the image information into regular images and non-regular images based on an encrypted database, and a text extraction unit for scanning the text information or the text in the non-regular images to generate corresponding feature paragraphs. The feature analysis module is connected to the information acquisition module and the information extraction module respectively, and includes an image filtering unit for filtering out the images to be encrypted from the regular images based on the user location information, and a paragraph analysis unit for comparing the feature paragraphs with preset paragraphs to generate paragraph similarity to determine sensitive information. An information processing module, which is connected to the information acquisition module and the feature analysis module respectively, is used to combine non-sensitive information and the regional encryption information to generate a sensitive characterization value to determine whether the non-sensitive information is encrypted, or to update the encryption database. The encryption module, which is connected to the information acquisition module, the feature analysis module, and the information processing module, includes a first encryption unit for encrypting images to be encrypted and sensitive information based on the regional encryption information, a second encryption unit for encrypting non-sensitive information determined to be encrypted, and a third encryption unit for determining non-sensitive information as non-sensitive information and encrypting it based on the boundary value analysis of the non-sensitive information.

2. The multi-regional data compliance adaptation engine intelligent system according to claim 1, characterized in that, The image classification unit is used to divide the image information into regular images and unregular images, wherein... Used to scan the image information to obtain the image location and image size; This is used to compare the image information sequentially with preset images in the encrypted database based on the image position and the image size, and determine whether the image information is a regular image or an unconventional image based on the comparison result.

3. The multi-regional data compliance adaptation engine intelligent system according to claim 2, characterized in that, The text extraction unit is used to scan text information or text in the unconventional image to generate several text paragraphs, wherein... Used to generate several digital feature values ​​based on the encrypted database; Used to scan each of the aforementioned text segments and generate corresponding numerical representation values; This is used to determine whether the corresponding text paragraph is a feature paragraph based on the comparison result of the numerical feature value and the numerical representation value.

4. The multi-regional data compliance adaptation engine intelligent system according to claim 2, characterized in that, The image filtering unit is used to obtain the user's location information and combine it with the encrypted database to determine a number of preset encrypted images and their corresponding blank matching degree, and to filter out the images to be encrypted from the regular images based on the blank matching degree.

5. The multi-regional data compliance adaptation engine intelligent system according to claim 3, characterized in that, The paragraph analysis unit includes: A character distance subunit, which is connected to the text extraction unit, is used to determine several sensitive words in the feature paragraph based on the encrypted database and to calculate the character distance between adjacent sensitive words in the feature paragraph. A sensitive information subunit, connected to the character distance subunit, is used to select several preset paragraphs that include the same number of sensitive words, and to calculate several paragraph similarities by sequentially comparing the character distance with the sensitive character distances corresponding to each preset paragraph to determine whether the feature paragraph is sensitive information.

6. The multi-regional data compliance adaptation engine intelligent system according to claim 5, characterized in that, The sensitivity indicator value is determined by the number of sensitive words and the encryption level of the sensitive words, wherein, Based on the geographical encryption information, the encryption level of the sensitive words is determined to include a first encryption level, a second encryption level, and a third encryption level.

7. The multi-regional data compliance adaptation engine intelligent system according to claim 6, characterized in that, The information processing module compares the sensitive characterization value with a preset sensitive value, and determines whether the non-sensitive information should be encrypted based on the comparison result. The preset sensitive value is negatively correlated with the sensitive word encryption level.

8. The multi-regional data compliance adaptation engine intelligent system according to claim 7, characterized in that, The information processing module is used to update the encrypted database, wherein, If it is determined that the non-sensitive information is encrypted, then it is determined that the non-sensitive information will be updated in the encrypted database.

9. The multi-regional data compliance adaptation engine intelligent system according to claim 8, characterized in that, The second encryption unit compares the boundary value of the non-sensitive information to be encrypted with a preset boundary value. Based on the comparison result, it determines that the reason why the feature segment corresponding to the non-sensitive information is non-sensitive information is either a scanning problem of the information extraction module or an update of the regional encryption information. The preset boundary value is positively correlated with the total number of characters in the non-sensitive information.

10. The multi-regional data compliance adaptation engine intelligent system according to claim 9, characterized in that, The information collection module matches the user location information with the encrypted regional information and calls the encrypted database corresponding to the encrypted regional information.

Citation Information

Patent Citations

  • Method, system and device for generating and decrypting secure electronic file and medium

    CN116108502A

  • Security compliance processing system and method for sensitive data

    CN110866281A

  • Data encryption compliance detection method and device

    CN115017519A

  • Network database user information encryption system and method

    CN115618398A

  • Sensitive data compliance control system and control method thereof

    CN117725280A