Student data desensitization method, device and system and storage medium
By correcting, segmenting and analyzing student information images, generating structured tagged text and masking it, the problems of privacy leakage and inefficiency in the supply of student file data are solved, and efficient and accurate data desensitization is achieved.
Patent Information
- Application Number
- CN202510706655.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-03
AI Technical Summary
The existing method of providing student file data is full exposure, resulting in the leakage of irrelevant students' privacy, low processing efficiency and high error rate.
By performing correction analysis, segmentation, coordinate analysis and table parameter analysis on the student information image, structured marked text is generated, and masking analysis is performed to obtain the desensitization result.
It improves the efficiency and accuracy of data desensitization, avoids the leakage of irrelevant sensitive information, and provides reliable technical guarantees for the safe sharing of educational data.
Smart Images

Figure CN120747495A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of data desensitization technology, and specifically to a student data desensitization method, device, system and storage medium. Background Art
[0002] Today, student rosters, as structured documents containing core data such as student status, academic records, and rewards and punishments, have become a crucial data carrier for schools, educational administration departments, and third-party service agencies to conduct their business. Current digital archive service systems generally use full-table exports or global sharing modes to retrieve information, leading to a typical dilemma: when a specific scenario only requires querying a single student's file, the system will still return complete student data, including sensitive fields such as the ID numbers, home addresses, and contact numbers of other students in the same class / grade. This "full exposure" data supply method carries two hidden dangers: first, information users may leak the privacy of unrelated students due to operational errors or abuse of authority; second, traditional solutions rely on manual screening of sensitive fields, which suffers from low processing efficiency, inconsistent rule enforcement, and delayed response in high-concurrency scenarios. These problems seriously restrict the compliant circulation and value release of educational data resources. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a student data desensitization method, device, system and storage medium in response to the shortcomings of the existing technology.
[0004] The present invention solves the above technical problems with the following technical solutions: A student data desensitization method comprises the following steps:
[0005] Importing a plurality of original student information images and a list of student names corresponding to each of the original student information images;
[0006] Performing correction analysis on each of the original student information images to obtain a corrected student information image corresponding to each of the original student information images;
[0007] Segmenting each of the corrected student information images to obtain a text feature image corresponding to each of the original student information images and a table structure image corresponding to each of the original student information images;
[0008] Performing coordinate analysis on the text feature images corresponding to the original student information images according to the student name lists to obtain a target coordinate vector set corresponding to the original student information images;
[0009] Performing table parameter analysis on each of the table structure images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images, to obtain a row height parameter sequence corresponding to each of the original student information images and a column width parameter sequence corresponding to each of the original student information images;
[0010] Performing registration analysis on each target coordinate vector set, the student name list corresponding to each original student information image, and the row height parameter sequence corresponding to each original student information image, respectively, to obtain a structured marked text corresponding to each original student information image;
[0011] A masking analysis is performed on each of the structured markup texts, the corrected student information images corresponding to each of the original student information images, and the column width parameter sequence corresponding to each of the original student information images, and the analysis results are used as the student data desensitization results.
[0012] Another technical solution of the present invention to solve the above technical problem is as follows: a student data desensitization device, comprising:
[0013] An import module, configured to import a plurality of original student information images and a list of student names corresponding to each of the original student information images;
[0014] a correction and analysis module, configured to perform correction and analysis on each of the original student information images to obtain a corrected student information image corresponding to each of the original student information images;
[0015] a segmentation module for segmenting each of the corrected student information images to obtain a text feature image corresponding to each of the original student information images and a table structure image corresponding to each of the original student information images;
[0016] A coordinate analysis module, configured to perform coordinate analysis on the text feature images corresponding to the original student information images according to the student name lists, to obtain a target coordinate vector set corresponding to each original student information image;
[0017] a table parameter analysis module, configured to perform table parameter analysis on each of the table structure images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images, to obtain a row height parameter sequence corresponding to each of the original student information images and a column width parameter sequence corresponding to each of the original student information images;
[0018] a registration analysis module, configured to perform registration analysis on each target coordinate vector set, the student name list corresponding to each original student information image, and the row height parameter sequence corresponding to each original student information image, to obtain a structured marked text corresponding to each original student information image;
[0019] The desensitization result acquisition module is used to perform masking analysis on each of the original student information images according to each of the structured marked texts, the corrected student information images corresponding to each of the original student information images, and the column width parameter sequence corresponding to each of the original student information images, and use the analysis results as the student data desensitization results.
[0020] Based on the above-mentioned student data desensitization method, the present invention also provides a student data desensitization system.
[0021] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: a student data desensitizing system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, the student data desensitizing method described above is implemented.
[0022] Based on the above-mentioned student data desensitization method, the present invention also provides a computer-readable storage medium.
[0023] Another technical solution of the present invention to solve the above technical problems is as follows: a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the student data desensitization method as described above is implemented.
[0024] The beneficial effects of the present invention are: a corrected student information image is obtained by correction analysis of the original student information image, a text feature image and a table structure image are obtained by segmentation of the corrected student information image, a target coordinate vector set is obtained by coordinate analysis of the text feature image according to the student name list, a row height parameter sequence and a column width parameter sequence are obtained by table parameter analysis of the table structure image according to the student name list and the corrected student information image, a structured marked text is obtained by alignment analysis of the target coordinate vector set, the student name list and the row height parameter sequence, a masking analysis of the original student information image is performed according to the structured marked text, the corrected student information image and the column width parameter sequence, and the analysis result is used as the student data desensitization result, which greatly improves the desensitization efficiency and accuracy, is compatible with digital service infrastructure in different scenarios, avoids the leakage of irrelevant sensitive information, effectively solves the problems of low efficiency, high error rate and privacy leakage risk in traditional manual processing, and provides a reliable technical guarantee for the safe sharing of education data. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic diagram of the process of desensitizing student data provided by an embodiment of the present invention;
[0026] Figure 2 This is a module block diagram of the student data desensitization device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0028] Figure 1 A flowchart of a student data desensitization method provided in an embodiment of the present invention.
[0029] like Figure 1 As shown, a student data desensitization method includes the following steps:
[0030] Importing a plurality of original student information images and a list of student names corresponding to each of the original student information images;
[0031] Performing correction analysis on each of the original student information images to obtain a corrected student information image corresponding to each of the original student information images;
[0032] Segmenting each of the corrected student information images to obtain a text feature image corresponding to each of the original student information images and a table structure image corresponding to each of the original student information images;
[0033] Performing coordinate analysis on the text feature images corresponding to the original student information images according to the student name lists to obtain a target coordinate vector set corresponding to the original student information images;
[0034] Performing table parameter analysis on each of the table structure images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images, to obtain a row height parameter sequence corresponding to each of the original student information images and a column width parameter sequence corresponding to each of the original student information images;
[0035] Performing registration analysis on each target coordinate vector set, the student name list corresponding to each original student information image, and the row height parameter sequence corresponding to each original student information image, respectively, to obtain a structured marked text corresponding to each original student information image;
[0036] A masking analysis is performed on each of the structured markup texts, the corrected student information images corresponding to each of the original student information images, and the column width parameter sequence corresponding to each of the original student information images, and the analysis results are used as the student data desensitization results.
[0037] It should be understood that when receiving the PDF document of the file roster (i.e., the original student information image) and the preset ordered list of student names (i.e., the student name list), the arrangement order of the list needs to establish a strict mapping relationship with the actual arrangement order of student information in the file roster.
[0038] In the above embodiment, the corrected student information image is obtained by correction analysis of the original student information image, the corrected student information image is segmented to obtain a text feature image and a table structure image, the coordinate analysis of the text feature image based on the student name list is performed to obtain a target coordinate vector set, the table parameter analysis of the table structure image based on the student name list and the corrected student information image is performed to obtain a row height parameter sequence and a column width parameter sequence, the registration analysis of the target coordinate vector set, the student name list and the row height parameter sequence is performed to obtain a structured marked text, the masking analysis of the original student information image is performed based on the structured marked text, the corrected student information image and the column width parameter sequence, and the analysis result is used as the student data desensitization result, which greatly improves the desensitization efficiency and accuracy, is compatible with digital service infrastructure in different scenarios, avoids the leakage of irrelevant sensitive information, effectively solves the problems of low efficiency, high error rate and privacy leakage risk in traditional manual processing, and provides a reliable technical guarantee for the safe sharing of education data.
[0039] Optionally, as an embodiment of the present invention, the process of performing correction analysis on each of the original student information images to obtain a corrected student information image corresponding to each of the original student information images includes:
[0040] Performing format conversion on each of the original student information images to obtain converted student information images corresponding to each of the original student information images;
[0041] performing binarization processing on each of the converted student information images respectively to obtain a binarized student information image corresponding to each of the original student information images;
[0042] Performing edge detection on each of the binarized student information images using a dual-threshold Canny edge detection algorithm to obtain a table outline image corresponding to each of the original student information images;
[0043] Using the probabilistic Hough transform algorithm to detect each of the table outline images, and obtain a set of straight line vectors corresponding to each of the original student information images;
[0044] Extracting the X-axis vectors corresponding to the original student information images from the table outline images respectively;
[0045] Calculating the angle between each straight line vector in each straight line vector set and the X-axis vector corresponding to each original student information image to obtain a plurality of target angles corresponding to each original student information image, and respectively combining the plurality of target angles corresponding to each original student information image to obtain a target angle set corresponding to each original student information image;
[0046] Extracting the median of each target angle set respectively, and using the extracted results as the tilt angle, thereby obtaining the tilt angle corresponding to each original student information image;
[0047] Determine whether the absolute value of the inclination angle is greater than or equal to the preset inclination angle. If not, use the binarized student information image as the corrected student information image; if so, use the affine transformation algorithm to correct the binarized student information image corresponding to the inclination angle, and use the correction result as the corrected student information image, thereby obtaining the corrected student information image corresponding to each of the original student information images.
[0048] Preferably, the preset inclination angle may be 1°.
[0049] It should be understood that the PDF document (ie, the original student information image) is converted into the original image in JPG format. orig (i.e. the converted student information image), rotation correction is performed to generate the corrected image I cor (i.e. the corrected student information image).
[0050] Specifically, for I orig (i.e. the converted student information image) is binarized to obtain image I gray (i.e., student information image after binarization); gray (i.e. the student information image after binarization) using the dual threshold Canny edge detection algorithm to enhance I gray (i.e. the student information image after binarization) and generates the document table outline gray_2 Image (i.e., table outline image); gray_2 (i.e., table outline image) uses the probabilistic Hough transform algorithm to obtain the line vector set L and calculate each line vector l in L (i.e., line vector set) i With image I gray_2The angle θ (i.e., the target angle) of the X-axis (i.e., the X-axis vector) is obtained to obtain the angle set θ L (i.e., target angle set); for θ L (i.e., the target angle set) is obtained by using the median method gray_2 The dominant tilt angle θ Main (i.e. tilt angle); when |θ Main |≥1°, for image I gray (i.e. the student information image after binarization) uses the affine transformation algorithm to perform rotation correction to generate image I cor (i.e. the corrected student information image).
[0051] In the above embodiment, each original student information image is corrected and analyzed to obtain a corrected student information image, which greatly improves the desensitization efficiency and accuracy, is compatible with digital service infrastructure in different scenarios, and avoids the leakage of irrelevant sensitive information.
[0052] Optionally, as an embodiment of the present invention, the corrected student information image includes a plurality of corrected student information coordinates;
[0053] The process of segmenting each corrected student information image to obtain a text feature image corresponding to each original student information image and a table structure image corresponding to each original student information image includes:
[0054] Import the maximum grayscale of the image, and segment the maximum grayscale of the image and each corrected student information coordinate using the first formula and the second formula respectively to obtain multiple text feature coordinates corresponding to each original student information image and multiple table structure coordinates corresponding to each original student information image. The first formula is:
[0055]
[0056] The second formula is:
[0057]
[0058] in, is the text feature coordinate corresponding to the jth corrected student information coordinate of the i-th original student information image, L is the maximum grayscale of the image, is the jth corrected student information coordinate corresponding to the i-th original student information image, is the table structure coordinate corresponding to the jth corrected student information coordinate of the i-th original student information image;
[0059] The text feature images corresponding to each original student information image are obtained through the multiple text feature coordinates corresponding to each original student information image, and the table structure images corresponding to each original student information image are obtained through the multiple table structure coordinates corresponding to each original student information image.
[0060] It should be understood that cor Generate text feature image I using dual-channel adaptive threshold segmentation algorithm text (i.e. text feature coordinates) and table structure image I table (i.e. table structure coordinates).
[0061] Specifically, for I cor The image (i.e. the corrected student information image) uses the Otsu algorithm to calculate the text area threshold T text ∈[0.7L,0.9L] and table line threshold T table ∈[0.3L,0.6L] (L is the maximum grayscale); generate a binary image I that satisfies text (i.e. text feature coordinates) and I table (i.e. table structure coordinates).
[0062] Specifically, for I cor The image (i.e. the corrected student information image) uses the Otsu algorithm to calculate the text area threshold T text and table line threshold T table , satisfying T text ∈[0.7L,0.9L], T table ∈[0.3L,0.6L] (L is the maximum grayscale (i.e. the maximum grayscale of the image)), and the binary images I are generated by the following formulas. text (i.e. text feature coordinates) and I table (i.e. table structure coordinates):
[0063]
[0064] In the above embodiment, the segmentation of the corrected student information image obtains a text feature image and a table structure image, which greatly improves the desensitization efficiency and accuracy, is compatible with digital service infrastructure in different scenarios, and avoids the leakage of irrelevant sensitive information.
[0065] Optionally, as an embodiment of the present invention, the process of performing coordinate analysis on the text feature images corresponding to the respective original student information images according to the respective student name lists to obtain a target coordinate vector set corresponding to each of the original student information images includes:
[0066] Cutting off the text feature images corresponding to the respective original student information images according to a preset width to obtain interest region images corresponding to the respective original student information images;
[0067] Extracting the length of the student list corresponding to each of the original student information images from each of the student name lists;
[0068] Import the first prime number set corresponding to each of the original student information images, calculate the length of each of the student lists and the first prime number set corresponding to each of the original student information images by the third formula, and obtain the target number of blocks corresponding to each of the original student information images. The third formula is:
[0069]
[0070] in, is the target block number corresponding to the i-th original student information image, is the wth prime number in the first prime number set corresponding to the i-th original student information image, P1 i is the first prime number set corresponding to the i-th original student information image, N i is the length of the student list corresponding to the i-th original student information image;
[0071] Dividing the interest region images corresponding to the respective original student information images into a plurality of recognition units corresponding to the respective original student information images at equal intervals according to the respective target block numbers;
[0072] Using an optical character recognition algorithm to detect each of the recognition units, respectively, to obtain coordinates of a plurality of student name content frames corresponding to each of the original student information images;
[0073] Extracting the interest region image height corresponding to each of the original student information images from each of the interest region images respectively;
[0074] The height of each of the image regions of interest, the number of target blocks corresponding to each of the original student information images, and the coordinates of multiple student name content boxes corresponding to each of the original student information images are calculated using the fourth formula to obtain multiple global coordinates corresponding to each of the original student information images. The fourth formula is:
[0075]
[0076] in, is the global coordinate corresponding to the coordinates of the kth student name content box in the i-th original student information image, is the coordinate of the kth student name content box corresponding to the i-th original student information image, m is the number of target blocks, is the image height of the region of interest corresponding to the i-th original student information image, is the target block number corresponding to the i-th original student information image;
[0077] A plurality of global coordinates corresponding to each of the original student information images and a list of student names corresponding to each of the original student information images are respectively collected to obtain a target coordinate vector set corresponding to each of the original student information images.
[0078] Preferably, the preset width is 50% of the width of the text feature image.
[0079] It should be understood that based on I text (i.e., text feature image) executes the improved OCR recognition algorithm to obtain the coordinate vector set K of the student name (i.e., the target coordinate vector set), as shown in the following formula:
[0080] K={(name i , x1, y1, x2, y2)|i∈[1, N]}.
[0081] Specifically, the longitudinal section I text (i.e., text feature image) and take the left half of the region of interest to generate image I roi (i.e., image of the region of interest); calculate the optimal number of blocks (i.e., target number of blocks) N based on the length of the student list N prime =min{p1∈P1|p1≥N}(P1 is a prime number set (i.e., the first prime number set), p1 is a prime number in the prime number set (i.e., the first prime number set)); roi (i.e., the image of the region of interest) is divided into equal intervals in the horizontal direction to generate N prime +1 recognition unit, and perform OCR detection on each recognition unit separately. If the OCR output of the mth unit contains any name in the initial name list, a global coordinate mapping is established for the content box f(x′1, y′1, x′2, y′2) containing the name in the recognition unit, and finally a coordinate vector set K (i.e., the target coordinate vector set) is obtained.
[0082] Specifically, when the OCR output of the mth unit contains any name in the initial name list, the following coordinate mapping relationship is established for the content frame coordinates f(x′1, y′1, x′2, y′2) containing the name in the recognized unit:
[0083] k(x1, y1, x2, y2)=Fn(x′1, y′1, x′2, y′2)
[0084] Among them Fm The coordinates f(x′1, y′1, x′2, y′2) in the mth recognition unit are converted to image I roi Matrix transformation of global coordinate k (x1, y1, x2, y2). The specific matrix transformation process is as follows:
[0085]
[0086] in, For image I roi Height, N prime For I roi The optimal number of blocks (i.e. the target number of blocks).
[0087] Finally, N prime All student name coordinate vectors k obtained in the +1 recognition unit are combined with the student names represented by their coordinate vectors to form a coordinate vector set (i.e., target coordinate vector set) K = {(name i , x1, y1, x2, y2)|i∈[1, N]}.
[0088] In the above embodiment, the coordinate analysis of the text feature image is performed based on the student name list to obtain a target coordinate vector set, which speeds up the overall detection speed, avoids the leakage of irrelevant sensitive information, and effectively solves the problems of low efficiency, high error rate and privacy leakage risk in traditional manual processing, providing reliable technical guarantees for the safe sharing of educational data.
[0089] Optionally, as an embodiment of the present invention, the process of performing table parameter analysis on each of the table structure images based on each of the student name lists and the corrected student information images corresponding to each of the original student information images to obtain a row height parameter sequence corresponding to each of the original student information images and a column width parameter sequence corresponding to each of the original student information images includes:
[0090] Performing expansion processing on each of the table structure images respectively to obtain an original table image corresponding to each of the original student information images;
[0091] Using a Hough transform algorithm, each of the original table images is divided into an original horizontal table line image corresponding to each of the original student information images and an original vertical table line image corresponding to each of the original student information images;
[0092] Redrawing each of the original horizontal table line images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images to obtain a redrawn horizontal table line image corresponding to each of the original student information images;
[0093] Redrawing each of the original vertical table line images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images to obtain a redrawn vertical table line image corresponding to each of the original student information images;
[0094] Using a probabilistic Hough line detection algorithm, extracting horizontal and vertical table line coordinate sets corresponding to each of the original student information images from each of the redrawn horizontal table line images and each of the redrawn vertical table line images;
[0095] The cluster analysis algorithm is used to calculate the parameter distribution of each horizontal and vertical table line coordinate set, and a row height parameter sequence corresponding to each original student information image and a column width parameter sequence corresponding to each original student information image are obtained.
[0096] It should be understood that the data processing process of redrawing each of the original horizontal table line images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images is the same as the data processing process of redrawing each of the original vertical table line images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images, and only the processed data is different.
[0097] It should be understood that table (i.e., table structure image) implements diffusion redrawing algorithm based on prime radius to generate accurate horizontal table line image I′ h-line (ie, the original horizontal table line image) and the vertical table line image I′ v-line (i.e. the original vertical table line image), and obtain the row height parameter sequence H={h i} and column width parameter sequence W = {w i}.
[0098] Specifically, a 3×15 pixel rectangular kernel is used as the structural element. table (i.e., the table structure image) performs morphological corrosion and expansion operations, and then generates an image containing only horizontal table lines I through Hough transform. h-line (ie, the original horizontal table line image) and the vertical table line image I v-line (i.e. original vertical table line image); respectively for I h-line (i.e. original horizontal table line image) and I v-line The image (i.e. the original vertical table line image) performs a diffusion redrawing algorithm based on a prime radius; for the redrawn I′ h-line (i.e. horizontal table line image after redrawing) and I′ v-lineThe image (ie, the vertical table line image after redrawing) executes the table parsing algorithm to obtain the table row height parameter sequence H and column width parameter sequence W.
[0099] Specifically, for I1′ h-line (i.e. horizontal table line image after redrawing) and I1′ v-line The image (i.e. the vertical table line image after redrawing) uses the probabilistic Hough line detection algorithm to extract the coordinate sets of the horizontal and vertical table lines (i.e. the horizontal and vertical table line coordinate sets), and then cluster analysis is used to calculate the row height distribution H = {h i} (i.e. row height parameter sequence) and column width distribution W = {w i} (i.e., column width parameter sequence), where i∈[1,N].
[0100] In the above embodiment, table parameter analysis is performed on the table structure image based on the student name list and the corrected student information image to obtain a row height parameter sequence and a column width parameter sequence, which effectively solves the problems of low efficiency, high error rate and privacy leakage risk existing in traditional manual processing, and provides reliable technical guarantee for the safe sharing of educational data.
[0101] Optionally, as an embodiment of the present invention, the process of redrawing each of the original horizontal table line images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images to obtain the redrawn horizontal table line images corresponding to each of the original student information images includes:
[0102] Extracting the number of student list items corresponding to each of the original student information images from each of the student name lists;
[0103] Each of the original horizontal table line images is divided into equal sizes according to an M×M grid to obtain a plurality of processing units corresponding to each of the original student information images, wherein: Among them, N′ i is the number of student list items corresponding to the i-th original student information image;
[0104] extracting from each of the processing units a plurality of processing unit widths corresponding to each of the original student information images and a plurality of processing unit heights corresponding to each of the original student information images;
[0105] The fifth formula is used to calculate the pixel coordinates of all original horizontal table lines in each processing unit, the widths of multiple processing units corresponding to each original student information image, and the heights of multiple processing units corresponding to each original student information image, to obtain multiple grayscale means corresponding to each original student information image. The fifth formula is:
[0106]
[0107] in, is the grayscale mean of the ath processing unit corresponding to the i-th original student information image, is the processing unit width of the ath processing unit corresponding to the i-th original student information image, is the processing unit height of the ath processing unit corresponding to the i-th original student information image, is the ath processing unit corresponding to the i-th original student information image, is the pixel coordinate of the oth original horizontal table line of the ath processing unit corresponding to the i-th original student information image;
[0108] Using the maximum inter-class variance algorithm, threshold calculation is performed on each of the corrected student information images to obtain a table line threshold corresponding to each of the original student information images;
[0109] If the grayscale mean satisfies the first judgment condition, all original horizontal table line pixel coordinates corresponding to the grayscale mean are used as target pixel points, thereby obtaining multiple target pixel points corresponding to each of the original student information images. The first judgment condition is:
[0110]
[0111] in, is the grayscale mean of the ath processing unit corresponding to the i-th original student information image, is the table line threshold corresponding to the i-th original student information image, and δ is the preset noise tolerance threshold;
[0112] If the target pixel point is greater than or equal to the table line threshold corresponding to each of the original student information images, raster scanning is performed on the target pixel point to obtain a plurality of raster points corresponding to each of the original student information images;
[0113] Obtaining, from each of the processing units, a processing unit left boundary corresponding to each of the processing units, a processing unit upper boundary corresponding to each of the processing units, a processing unit right boundary corresponding to each of the processing units, and a processing unit lower boundary corresponding to each of the processing units;
[0114] The sixth formula is used to calculate each of the grating points, the left boundary of the processing unit corresponding to each of the processing units, the upper boundary of the processing unit corresponding to each of the processing units, the right boundary of the processing unit corresponding to each of the processing units, and the lower boundary of the processing unit corresponding to each of the processing units, to obtain a plurality of boundary distances corresponding to each of the original student information images. The sixth formula is:
[0115]
[0116] in, is the boundary distance of the ath processing unit corresponding to the i-th original student information image, is the x-axis coordinate of the raster point of the a-th processing unit corresponding to the i-th original student information image, is the left boundary of the processing unit of the ath processing unit corresponding to the i-th original student information image, is the y-axis coordinate of the raster point of the a-th processing unit corresponding to the i-th original student information image, is the right boundary of the processing unit of the ath processing unit corresponding to the i-th original student information image, is the upper boundary of the processing unit of the ath processing unit corresponding to the i-th original student information image, is the lower boundary of the processing unit of the ath processing unit corresponding to the i-th original student information image;
[0117] Import the second prime number set corresponding to each of the original student information images, and calculate each of the boundary distances and the second prime number set corresponding to each of the original student information images using the seventh formula to obtain a plurality of prime number radii corresponding to each of the original student information images. The seventh formula is:
[0118]
[0119] in, is the prime radius of the ath processing unit corresponding to the i-th original student information image, is the w′th prime number in the second prime number set corresponding to the i-th original student information image, P2 i is the second prime number set corresponding to the i-th original student information image, is the boundary distance of the ath processing unit corresponding to the i-th original student information image;
[0120] The eighth formula is used to calculate each of the grating points and a plurality of prime number radii corresponding to each of the original student information images, thereby obtaining a plurality of circular search domains corresponding to each of the original student information images. The eighth formula is:
[0121]
[0122] in, is the circular search domain of the ath processing unit corresponding to the i-th original student information image, is the prime radius of the ath processing unit corresponding to the i-th original student information image, is the raster point of the ath processing unit corresponding to the i-th original student information image, is the ath processing unit corresponding to the i-th original student information image, is the adjacent raster point of the ath processing unit corresponding to the i-th original student information image;
[0123] If the adjacent grating points satisfy the second judgment condition, the ninth formula is used to calculate each of the grating points and the multiple adjacent grating points corresponding to each of the original student information images, thereby obtaining the nearest neighbor point corresponding to each of the grating points. The second judgment condition is:
[0124]
[0125] in, is the adjacent raster point of the ath processing unit corresponding to the i-th original student information image, is the table line threshold corresponding to the i-th original student information image, is the raster point of the ath processing unit corresponding to the i-th original student information image;
[0126] The ninth formula is:
[0127]
[0128] in, is the nearest neighbor point of the ath processing unit corresponding to the i-th original student information image, is the raster point of the ath processing unit corresponding to the i-th original student information image, is the adjacent raster point of the ath processing unit corresponding to the i-th original student information image;
[0129] Drawing line segments for each of the nearest neighbor points and each of the raster points corresponding to each of the original student information images to obtain a plurality of target line segments corresponding to each of the original student information images, and respectively gathering the plurality of target line segments corresponding to each of the original student information images to obtain a line segment set corresponding to each of the original student information images;
[0130] Performing path filling on each of the line segment sets using the Bresenham algorithm to obtain filled horizontal table line images corresponding to each of the original student information images;
[0131] Using a 3×3 mean algorithm to filter each of the filled horizontal table line images, to obtain a filtered horizontal table line image corresponding to each of the original student information images;
[0132] The pixel values of each filtered horizontal table line image are modified according to the preset table line pixel values to obtain a redrawn horizontal table line image corresponding to each original student information image.
[0133] Specifically, the steps for redrawing an image are as follows:
[0134] 1. Perform M×M grid block operation on the image (i.e. the original horizontal table line image) N is the number of student list items), divide the image (i.e. the original horizontal table line image) into equal-sized processing units B ij , where i, j∈[1,M] represents the row and column indices.
[0135] 2. For each block B ij (i.e., processing unit) performs the following steps:
[0136] a) Calculate the grayscale mean of the block (i.e. grayscale mean):
[0137] b) If and only if μ ij >T table +δ (δ is the noise tolerance threshold, the default is δ=0.05L), the block is activated and the following steps are entered.
[0138] 3. Traverse the pixel point pixel(x, y) (i.e. target pixel point) in the active block in raster scan order. pixel (x, y) ≥ T table If yes, execute steps a), b), c), d); otherwise, continue to traverse the next pixel (i.e., the target pixel):
[0139] a) Calculate the boundary distance parameter (i.e., boundary distance) of the pixel point (i.e., raster point): d = min (xx min ,yy min , x max -x,y max -y), where x min / y min is the left / top boundary of the block, x max / y max for the right / bottom border;
[0140] b) Determine the prime radius: Wherein P2 is a prime number set (i.e., the second prime number set), and p2 is a prime number in the prime number set (i.e., the second prime number set);
[0141] c) Construct a circular search domain with pixel (x, y) as the center: Ω(p, r) = {q∈B ij |||qp||2≤r};
[0142] d) Perform proximity detection:
[0143] i.if (p, r) satisfies I(q) ≥ T table When ∧q≠p, execute steps (1) and (2):
[0144] (1) Select the nearest neighbor point q * =argmin q ||qp||2;
[0145] (2) Draw the line segment L(p,q*) and update the current processing point to q*;
[0146] ii. Otherwise, continue to traverse the next pixel.
[0147] 4. For the line segment set L generated in step 3 k For each line segment (i.e., target line segment) in the (i.e., line segment set), perform the following steps:
[0148] a) Use Bresenham algorithm to fill pixel paths;
[0149] b) Apply 3×3 mean filtering to eliminate the aliasing effect;
[0150] c) Update the grayscale value of the pixel in the line segment to 0.
[0151] 5. Terminate the algorithm when all blocks have been traversed.
[0152] It should be understood that the Bresenham algorithm is a classic algorithm in computer graphics, primarily used for drawing straight lines on raster displays (devices such as screens and printers). Proposed by Jack Elton Bresenham in 1962, it efficiently determines the pixels through which a line passes, avoiding floating-point operations and significantly improving computational speed.
[0153] It should be understood that the 3×3 mean algorithm is a basic linear filtering method in image processing, commonly used in scenarios such as image smoothing and noise reduction. Its core concept is to use a 3×3 sliding window to move pixel by pixel across the image, replacing the center pixel value with the average of all pixels within the window, thereby achieving the effect of smoothing the image and reducing noise.
[0154] In the above embodiment, the original horizontal table line image is redrawn according to the student name list and the corrected student information image to obtain a redrawn horizontal table line image. Compared with the traditional method of constructing a table, it is more accurate, avoids the leakage of irrelevant sensitive information, and effectively solves the problems of low efficiency, high error rate and privacy leakage risk in traditional manual processing, providing reliable technical guarantees for the safe sharing of educational data.
[0155] Optionally, as an embodiment of the present invention, the target coordinate vector set includes a plurality of target coordinate vectors, the student name list includes a student name sublist, and the row height parameter sequence includes a plurality of row height parameters;
[0156] The process of performing registration analysis on each target coordinate vector set, the student name list corresponding to each original student information image, and the line height parameter sequence corresponding to each original student information image to obtain the structured marked text corresponding to each original student information image includes:
[0157] Determine whether the student name sublist is the same as the target coordinate vector at the same position,
[0158] If so, the target triplet is constructed from the student name sublist and the target coordinate vector at the same position by the tenth formula, which is:
[0159]
[0160] in, is the v1th target triplet corresponding to the i-th original student information image, is the v1th target coordinate vector corresponding to the i-th original student information image, is the v1th row height parameter corresponding to the i-th original student information image;
[0161] If not, then the target triplet is constructed by the eleventh formula from the student name sublist, the target triplet at the previous position, and the target coordinate vector at the same position. The eleventh formula is:
[0162]
[0163] in, is the v2th target triplet corresponding to the i-th original student information image, is the x1 coordinate of the v1-1th target triplet corresponding to the i-th original student information image or the x1 coordinate of the v2-1th target triplet corresponding to the i-th original student information image, is the v2th row height parameter corresponding to the i-th original student information image, is the y1 coordinate of the v1-1th target triplet corresponding to the i-th original student information image or the y1 coordinate of the v2-1th target triplet corresponding to the i-th original student information image, is the x2 coordinate of the v1-1th target triplet corresponding to the i-th original student information image or the x2 coordinate of the v2-1th target triplet corresponding to the i-th original student information image, is the y2 coordinate of the v1-1th target triplet corresponding to the i-th original student information image or the y2 coordinate of the v2-1th target triplet corresponding to the i-th original student information image;
[0164] The structured marked text corresponding to each of the original student information images is obtained through the multiple target triples corresponding to each of the original student information images.
[0165] It should be understood that a spatial registration algorithm is used to generate a structured markup (XML) file (ie, structured markup text) containing name-coordinate-line height triples.
[0166] Specifically, receive the student name coordinate vector set K (i.e., the target coordinate vector set); receive the row height sequence H (i.e., the row height parameter sequence) and the column width sequence W; use a spatial alignment algorithm on the coordinate vector set K (i.e., the target coordinate vector set), the row height sequence H (i.e., the row height parameter sequence) and the column width sequence W to generate triple data containing name-coordinate-row height (i.e., the target triple), and record it in an XML file.
[0167] Specifically, combining K (i.e., target coordinate vector set), H (i.e., row height parameter sequence), W, and the ordered list of student names U (i.e., student name list) of length N, constructs a name-coordinate-row height triple T = {(name i ,x1,y1,x2,y2,h i )|i∈[1,N]}, the specific steps are as follows:
[0168] a) Traverse list U (i.e., the list of student names) in sequence and perform the following steps:
[0169] i.if u i The name is equivalent to k i If the name in k i With h i Combine to generate the sub-item t of triple T i =(name i ,x1,y1,x2,y2,h i ), where (name i , x1, y1, x2, y2) from k i ;
[0170] where u i is a subitem of list U (i.e., a list of student names), k i is a sub-item of the name coordinate vector set K (i.e., the target coordinate vector set), h i It is a sub-item of the row height sequence H (i.e., row height parameter sequence);
[0171] ii.if u i The name is not equal to k i The names in t are reconstructed based on the following method i =(name i ,x1,y1,x2,y2,h i ):
[0172] 1)name i =u i ,
[0173] 2)
[0174] 3)
[0175] 4)h i =h i ,
[0176] Finally, a complete triplet T after compensation (ie, target triplet) is generated and recorded in a standard structured XML tag file (ie, structured tag text).
[0177] In the above embodiment, the target coordinate vector set, the student name list and the row height parameter sequence are aligned and analyzed to obtain structured marked text, which can automatically identify the file table structure and select the optimal masking strategy, avoiding the leakage of irrelevant sensitive information, and effectively solving the problems of low efficiency, high error rate and privacy leakage risk in traditional manual processing, providing reliable technical guarantees for the safe sharing of educational data.
[0178] Optionally, as an embodiment of the present invention, the process of performing masking analysis on each of the original student information images according to each of the structured markup texts, the corrected student information images corresponding to each of the original student information images, and the column width parameter sequences corresponding to each of the original student information images, and using the analysis results as the student data desensitization results includes:
[0179] Counting the number of parameters in each column width parameter sequence respectively to obtain the number of column width parameters corresponding to each original student information image;
[0180] Determine whether the number of column width parameters is 0,
[0181] If so, the width of the corrected student information image corresponding to the number of column width parameters is used as the masking width;
[0182] If not, all parameters in the column width parameter sequence corresponding to the number of column width parameters are added, and the result of the addition is used as the mask width;
[0183] performing masking processing on each of the original student information images according to each of the structured markup texts and the masking width corresponding to each of the original student information images, to obtain a first masking area corresponding to each of the original student information images;
[0184] Using a pixel mask algorithm to mask the remaining areas of each of the original student information images, respectively, to obtain second masked areas corresponding to each of the original student information images;
[0185] The desensitized images corresponding to each of the original student information images are obtained through the first masked areas corresponding to each of the original student information images and the second masked areas corresponding to each of the original student information images, and all of the desensitized images are taken as the student data desensitization results.
[0186] It should be understood that the remaining area in the original student information image may be an area other than the first shielded area.
[0187] It should be understood that the masking algorithm is used to orig (i.e. the original student information image) performs regional selective pixel coverage and outputs a desensitized image I that meets the standard result .
[0188] Specifically, when the column width sequence W (i.e., the column width parameter sequence) contains sub-items, the mask width is calculated as W mask =W table , where W table is the sub-item w in the column width sequence W i When the column width sequence W (i.e., the column width parameter sequence) does not contain sub-items, the masking width is the image width (i.e., the width of the corrected student information image); pixel masking is performed on the non-target area (i.e., the remaining area in the original student information image) in combination with the XML tag file (i.e., structured tag text).
[0189] It should be understood that, in combination with the shielding width W mask XML markup files (structured markup text) orig(i.e. the original student information image) is selectively covered with pixels. Based on the target student's name, the triple data in the XML file (i.e. the structured markup text) is searched, the final masked area (i.e. the first masked area) is calculated, and the non-target area is masked by the pixel mask algorithm, and the final desensitized image file I is output. result .
[0190] In the above embodiment, the original student information image is masked and analyzed based on the structured marked text, the corrected student information image, and the column width parameter sequence, and the analysis result is used as the student data desensitization result, which effectively solves the problems of low efficiency, high error rate and privacy leakage risk in traditional manual processing, and provides reliable technical guarantee for the safe sharing of educational data.
[0191] Alternatively, as another embodiment of the present invention, the present invention belongs to the field of digital archive services. This method achieves precise information desensitization through an automated processing flow and primarily includes the following technical solutions: receiving a PDF of an archive roster and an ordered list of student names; generating a text feature image and a table structure image through tilt correction and dual-channel threshold segmentation; acquiring name coordinate vectors using prime number block optimization OCR recognition; accurately extracting table row height and column width parameters using a prime number radius diffusion redrawing algorithm; generating a structured XML file through spatial registration; and finally implementing a dynamic masking strategy to output the desensitized image.
[0192] The advantages of the present invention are reflected in three aspects: First, it uses dual-channel image processing technology to improve positioning accuracy by separating and analyzing text and tables; second, it reconstructs accurate table row height and column width parameters based on an improved OCR recognition algorithm combined with a prime radius diffusion redrawing algorithm; generates a structured XML file through a spatial registration algorithm, and finally implements pixel coverage of non-target areas based on a dynamic masking strategy; third, it establishes a dynamic compensation mechanism to infer coordinates based on the global average interval distance when local recognition fails, greatly improving the desensitization accuracy. The system architecture includes three modules: a digital archive service system, an index library, and a database. It supports real-time / offline dual-mode operation, can handle high concurrent requests, and effectively solves the problems of low efficiency, high error rate, and privacy leakage risks in traditional manual processing, providing reliable technical guarantees for the secure sharing of educational data.
[0193] Optionally, as another embodiment of the present invention, the present invention includes the following steps:
[0194] S1. Receive a PDF document of the file roster and a preset ordered list of student names. The order of the list must be strictly mapped to the order of the actual student information in the file roster.
[0195] S2. Convert PDF documents to original images in JPG format I orig, using the tilt detection algorithm to perform rotation correction to generate the corrected image I cor , then I cor Generate text feature image I using dual-channel adaptive threshold segmentation algorithm text and table structure image I table ;
[0196] S3, based on I text Execute the improved OCR recognition algorithm to obtain the coordinate vector set K of the student name
[0197] K={(name i ,x1,y1,x2,y2)|i∈[1,N]}
[0198] S4, to I table Implement a diffusion redrawing algorithm based on prime radius to generate an accurate horizontal table line image I' h-line and vertical table line image I′ v-line , and obtain the row height parameter sequence H={h i} and column width parameter sequence W = {w i};
[0199] S5. Apply a spatial registration algorithm to the output of steps S3-S4 to generate a structured markup (XML) file containing name-coordinate-row height triples;
[0200] S6, using masking algorithm to orig Implement region-selective pixel coverage and output a desensitized image that meets the standard I result .
[0201] Optionally, as another embodiment of the present invention, the present invention includes a digital archive service system, an archive information index library and an archive database.
[0202] Digital archive service system: provides services including user login, archive information query, archive information download, etc.
[0203] Archive information index database: including file numbers, archive roster name lists;
[0204] File database: including file number and file scan file path address.
[0205] Alternatively, as another embodiment of the present invention, the present invention can provide online digital archive services around the clock, avoiding the efficiency bottlenecks and fixed working hours of manual desensitization. Compared with existing methods, the innovation of the present invention lies in:
[0206] (1) A diffusion redrawing algorithm based on prime radius is proposed, which is more accurate than the traditional Hough transform method in constructing tables;
[0207] (2) A spatial registration algorithm was proposed to automatically identify the archival table structure and select the optimal masking strategy;
[0208] Compared with existing methods, its beneficial effects are:
[0209] (1) Provide 24 / 7 online digital archive services;
[0210] (2) Desensitization services can be provided by using pre-desensitization or real-time desensitization methods, and can be compatible with digital service infrastructure in different scenarios;
[0211] (3) Avoiding the leakage of irrelevant sensitive information;
[0212] (4) Adapt to the needs of high-concurrency scenarios;
[0213] (5) Greatly improved desensitization efficiency and accuracy.
[0214] Figure 2 A module block diagram of a student data desensitizing device provided in an embodiment of the present invention.
[0215] Alternatively, as another embodiment of the present invention, Figure 2 As shown, a student data desensitization device includes:
[0216] An import module, configured to import a plurality of original student information images and a list of student names corresponding to each of the original student information images;
[0217] a correction and analysis module, configured to perform correction and analysis on each of the original student information images to obtain a corrected student information image corresponding to each of the original student information images;
[0218] a segmentation module for segmenting each of the corrected student information images to obtain a text feature image corresponding to each of the original student information images and a table structure image corresponding to each of the original student information images;
[0219] A coordinate analysis module, configured to perform coordinate analysis on the text feature images corresponding to the original student information images according to the student name lists, to obtain a target coordinate vector set corresponding to each original student information image;
[0220] a table parameter analysis module, configured to perform table parameter analysis on each of the table structure images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images, to obtain a row height parameter sequence corresponding to each of the original student information images and a column width parameter sequence corresponding to each of the original student information images;
[0221] a registration analysis module, configured to perform registration analysis on each target coordinate vector set, the student name list corresponding to each original student information image, and the row height parameter sequence corresponding to each original student information image, to obtain a structured marked text corresponding to each original student information image;
[0222] The desensitization result acquisition module is used to perform masking analysis on each of the original student information images according to each of the structured marked texts, the corrected student information images corresponding to each of the original student information images, and the column width parameter sequence corresponding to each of the original student information images, and use the analysis results as the student data desensitization results.
[0223] Alternatively, another embodiment of the present invention provides a student data desensitization system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the student data desensitization method described above is implemented. The system may be a computer or other system.
[0224] Optionally, another embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the student data desensitization method as described above is implemented.
[0225] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0226] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0227] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or ignoring or not implementing certain features.
[0228] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the objectives of the embodiments of the present invention.
[0229] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0230] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0231] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A student data desensitization method, characterized in that: The steps include: Importing a plurality of original student information images and a list of student names corresponding to each of the original student information images; Performing correction analysis on each of the original student information images to obtain a corrected student information image corresponding to each of the original student information images; Segmenting each of the corrected student information images to obtain a text feature image corresponding to each of the original student information images and a table structure image corresponding to each of the original student information images; Performing coordinate analysis on the text feature images corresponding to the original student information images according to the student name lists to obtain a target coordinate vector set corresponding to the original student information images; Performing table parameter analysis on each of the table structure images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images, to obtain a row height parameter sequence corresponding to each of the original student information images and a column width parameter sequence corresponding to each of the original student information images; Performing registration analysis on each target coordinate vector set, the student name list corresponding to each original student information image, and the row height parameter sequence corresponding to each original student information image, respectively, to obtain a structured marked text corresponding to each original student information image; A masking analysis is performed on each of the structured markup texts, the corrected student information images corresponding to each of the original student information images, and the column width parameter sequence corresponding to each of the original student information images, and the analysis results are used as the student data desensitization results.
2. The student data desensitization method according to claim 1, characterized in that: The process of respectively performing correction analysis on each of the original student information images to obtain a corrected student information image corresponding to each of the original student information images includes: Performing format conversion on each of the original student information images to obtain converted student information images corresponding to each of the original student information images; performing binarization processing on each of the converted student information images respectively to obtain a binarized student information image corresponding to each of the original student information images; Performing edge detection on each of the binarized student information images using a dual-threshold Canny edge detection algorithm to obtain a table outline image corresponding to each of the original student information images; Using the probabilistic Hough transform algorithm to detect each of the table outline images, and obtain a set of straight line vectors corresponding to each of the original student information images; Extracting the X-axis vectors corresponding to the original student information images from the table outline images respectively; Calculating the angle between each straight line vector in each straight line vector set and the X-axis vector corresponding to each original student information image to obtain a plurality of target angles corresponding to each original student information image, and respectively combining the plurality of target angles corresponding to each original student information image to obtain a target angle set corresponding to each original student information image; Extracting the median of each target angle set respectively, and using the extracted results as the tilt angle, thereby obtaining the tilt angle corresponding to each original student information image; Determine whether the absolute value of the inclination angle is greater than or equal to the preset inclination angle. If not, use the binarized student information image as the corrected student information image; if so, use the affine transformation algorithm to correct the binarized student information image corresponding to the inclination angle, and use the correction result as the corrected student information image, thereby obtaining the corrected student information image corresponding to each of the original student information images.
3. The student data desensitization method according to claim 1, characterized in that: The corrected student information image includes a plurality of corrected student information coordinates; The process of segmenting each corrected student information image to obtain a text feature image corresponding to each original student information image and a table structure image corresponding to each original student information image includes: Import the maximum grayscale of the image, and segment the maximum grayscale of the image and each corrected student information coordinate using the first formula and the second formula respectively to obtain multiple text feature coordinates corresponding to each original student information image and multiple table structure coordinates corresponding to each original student information image. The first formula is: The second formula is: in, is the text feature coordinate corresponding to the jth corrected student information coordinate of the i-th original student information image, L is the maximum grayscale of the image, is the jth corrected student information coordinate corresponding to the i-th original student information image, is the table structure coordinate corresponding to the jth corrected student information coordinate of the i-th original student information image; The text feature images corresponding to each original student information image are obtained through the multiple text feature coordinates corresponding to each original student information image, and the table structure images corresponding to each original student information image are obtained through the multiple table structure coordinates corresponding to each original student information image.
4. The student data desensitization method according to claim 1, characterized in that: The process of performing coordinate analysis on the text feature images corresponding to the original student information images according to the student name lists to obtain a target coordinate vector set corresponding to each original student information image includes: Cutting off the text feature images corresponding to the respective original student information images according to a preset width to obtain interest region images corresponding to the respective original student information images; Extracting the length of the student list corresponding to each of the original student information images from each of the student name lists; Import the first prime number set corresponding to each of the original student information images, calculate the length of each of the student lists and the first prime number set corresponding to each of the original student information images by the third formula, and obtain the target number of blocks corresponding to each of the original student information images. The third formula is: in, is the target block number corresponding to the i-th original student information image, is the wth prime number in the first prime number set corresponding to the i-th original student information image, P1 i is the first prime number set corresponding to the i-th original student information image, N i is the length of the student list corresponding to the i-th original student information image; Dividing the interest region images corresponding to the respective original student information images into a plurality of recognition units corresponding to the respective original student information images at equal intervals according to the respective target block numbers; Using an optical character recognition algorithm to detect each of the recognition units, respectively, to obtain coordinates of a plurality of student name content frames corresponding to each of the original student information images; Extracting the interest region image height corresponding to each of the original student information images from each of the interest region images respectively; The height of each of the image regions of interest, the number of target blocks corresponding to each of the original student information images, and the coordinates of multiple student name content boxes corresponding to each of the original student information images are calculated using the fourth formula to obtain multiple global coordinates corresponding to each of the original student information images. The fourth formula is: in, is the global coordinate corresponding to the coordinates of the kth student name content box in the i-th original student information image, is the coordinate of the kth student name content box corresponding to the i-th original student information image, m is the number of target blocks, is the image height of the region of interest corresponding to the i-th original student information image, is the target block number corresponding to the i-th original student information image; A plurality of global coordinates corresponding to each of the original student information images and a list of student names corresponding to each of the original student information images are respectively collected to obtain a target coordinate vector set corresponding to each of the original student information images.
5. The student data desensitization method according to claim 1, characterized in that: The process of performing table parameter analysis on each of the table structure images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images to obtain a row height parameter sequence corresponding to each of the original student information images and a column width parameter sequence corresponding to each of the original student information images includes: Performing expansion processing on each of the table structure images respectively to obtain an original table image corresponding to each of the original student information images; Using a Hough transform algorithm, each of the original table images is divided into an original horizontal table line image corresponding to each of the original student information images and an original vertical table line image corresponding to each of the original student information images; Redrawing each of the original horizontal table line images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images to obtain a redrawn horizontal table line image corresponding to each of the original student information images; Redrawing each of the original vertical table line images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images to obtain a redrawn vertical table line image corresponding to each of the original student information images; Using a probabilistic Hough line detection algorithm, extracting horizontal and vertical table line coordinate sets corresponding to each of the original student information images from each of the redrawn horizontal table line images and each of the redrawn vertical table line images; The cluster analysis algorithm is used to calculate the parameter distribution of each horizontal and vertical table line coordinate set, and a row height parameter sequence corresponding to each original student information image and a column width parameter sequence corresponding to each original student information image are obtained.
6. The student data desensitization method according to claim 5, characterized in that: The process of redrawing each of the original horizontal table line images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images to obtain the redrawn horizontal table line images corresponding to each of the original student information images includes: Extracting the number of student list items corresponding to each of the original student information images from each of the student name lists; Each of the original horizontal table line images is divided into equal sizes according to an M×M grid to obtain a plurality of processing units corresponding to each of the original student information images, wherein: Among them, N′ i is the number of student list items corresponding to the i-th original student information image; extracting from each of the processing units a plurality of processing unit widths corresponding to each of the original student information images and a plurality of processing unit heights corresponding to each of the original student information images; The fifth formula is used to calculate the pixel coordinates of all original horizontal table lines in each processing unit, the widths of multiple processing units corresponding to each original student information image, and the heights of multiple processing units corresponding to each original student information image, to obtain multiple grayscale means corresponding to each original student information image. The fifth formula is: in, is the grayscale mean of the ath processing unit corresponding to the i-th original student information image, is the processing unit width of the ath processing unit corresponding to the i-th original student information image, is the processing unit height of the ath processing unit corresponding to the i-th original student information image, is the ath processing unit corresponding to the i-th original student information image, is the pixel coordinate of the oth original horizontal table line of the ath processing unit corresponding to the i-th original student information image; Using the maximum inter-class variance algorithm, threshold calculation is performed on each of the corrected student information images to obtain a table line threshold corresponding to each of the original student information images; If the grayscale mean satisfies the first judgment condition, all original horizontal table line pixel coordinates corresponding to the grayscale mean are used as target pixel points, thereby obtaining multiple target pixel points corresponding to each of the original student information images. The first judgment condition is: in, is the grayscale mean of the ath processing unit corresponding to the i-th original student information image, is the table line threshold corresponding to the i-th original student information image, and δ is the preset noise tolerance threshold; If the target pixel point is greater than or equal to the table line threshold corresponding to each of the original student information images, raster scanning is performed on the target pixel point to obtain a plurality of raster points corresponding to each of the original student information images; Obtaining, from each of the processing units, a processing unit left boundary corresponding to each of the processing units, a processing unit upper boundary corresponding to each of the processing units, a processing unit right boundary corresponding to each of the processing units, and a processing unit lower boundary corresponding to each of the processing units; The sixth formula is used to calculate each of the grating points, the left boundary of the processing unit corresponding to each of the processing units, the upper boundary of the processing unit corresponding to each of the processing units, the right boundary of the processing unit corresponding to each of the processing units, and the lower boundary of the processing unit corresponding to each of the processing units, to obtain a plurality of boundary distances corresponding to each of the original student information images. The sixth formula is: in, is the boundary distance of the ath processing unit corresponding to the i-th original student information image, is the x-axis coordinate of the raster point of the a-th processing unit corresponding to the i-th original student information image, is the left boundary of the processing unit of the ath processing unit corresponding to the i-th original student information image, is the y-axis coordinate of the raster point of the a-th processing unit corresponding to the i-th original student information image, is the right boundary of the processing unit of the ath processing unit corresponding to the i-th original student information image, is the upper boundary of the processing unit of the ath processing unit corresponding to the i-th original student information image, is the lower boundary of the processing unit of the ath processing unit corresponding to the i-th original student information image; Import the second prime number set corresponding to each of the original student information images, and calculate each of the boundary distances and the second prime number set corresponding to each of the original student information images using the seventh formula to obtain a plurality of prime number radii corresponding to each of the original student information images. The seventh formula is: in, is the prime radius of the ath processing unit corresponding to the i-th original student information image, is the w′th prime number in the second prime number set corresponding to the i-th original student information image, P2 i is the second prime number set corresponding to the i-th original student information image, is the boundary distance of the ath processing unit corresponding to the i-th original student information image; The eighth formula is used to calculate each of the grating points and a plurality of prime number radii corresponding to each of the original student information images, thereby obtaining a plurality of circular search domains corresponding to each of the original student information images. The eighth formula is: in, is the circular search domain of the ath processing unit corresponding to the i-th original student information image, is the prime radius of the ath processing unit corresponding to the i-th original student information image, is the raster point of the ath processing unit corresponding to the i-th original student information image, is the ath processing unit corresponding to the i-th original student information image, is the adjacent raster point of the ath processing unit corresponding to the i-th original student information image; If the adjacent grating points satisfy the second judgment condition, the ninth formula is used to calculate each of the grating points and the multiple adjacent grating points corresponding to each of the original student information images, thereby obtaining the nearest neighbor point corresponding to each of the grating points. The second judgment condition is: in, is the adjacent raster point of the ath processing unit corresponding to the i-th original student information image, is the table line threshold corresponding to the i-th original student information image, is the raster point of the ath processing unit corresponding to the i-th original student information image; The ninth formula is: in, is the nearest neighbor point of the ath processing unit corresponding to the i-th original student information image, is the raster point of the ath processing unit corresponding to the i-th original student information image, is the adjacent raster point of the ath processing unit corresponding to the i-th original student information image; Drawing line segments for each of the nearest neighbor points and each of the raster points corresponding to each of the original student information images to obtain a plurality of target line segments corresponding to each of the original student information images, and respectively gathering the plurality of target line segments corresponding to each of the original student information images to obtain a line segment set corresponding to each of the original student information images; Performing path filling on each of the line segment sets using the Bresenham algorithm to obtain filled horizontal table line images corresponding to each of the original student information images; Using a 3×3 mean algorithm to filter each of the filled horizontal table line images, to obtain a filtered horizontal table line image corresponding to each of the original student information images; The pixel values of each filtered horizontal table line image are modified according to the preset table line pixel values to obtain a redrawn horizontal table line image corresponding to each original student information image.
7. The student data desensitization method according to claim 1, characterized in that: The target coordinate vector set includes a plurality of target coordinate vectors, the student name list includes a student name sublist, and the row height parameter sequence includes a plurality of row height parameters; The process of performing registration analysis on each target coordinate vector set, the student name list corresponding to each original student information image, and the line height parameter sequence corresponding to each original student information image to obtain the structured marked text corresponding to each original student information image includes: Determine whether the student name sublist is the same as the target coordinate vector at the same position, If so, the target triplet is constructed from the student name sublist and the target coordinate vector at the same position by the tenth formula, which is: in, is the v1th target triplet corresponding to the i-th original student information image, is the v1th target coordinate vector corresponding to the i-th original student information image, is the v1th row height parameter corresponding to the i-th original student information image; If not, then the target triplet is constructed by the eleventh formula from the student name sublist, the target triplet at the previous position, and the target coordinate vector at the same position. The eleventh formula is: in, is the v2th target triplet corresponding to the i-th original student information image, is the x1 coordinate of the v1-1th target triplet corresponding to the i-th original student information image or the x1 coordinate of the v2-1th target triplet corresponding to the i-th original student information image, is the v2th row height parameter corresponding to the i-th original student information image, is the y1 coordinate of the v1-1th target triplet corresponding to the i-th original student information image or the y1 coordinate of the v2-1th target triplet corresponding to the i-th original student information image, is the x2 coordinate of the v1-1th target triplet corresponding to the i-th original student information image or the x2 coordinate of the v2-1th target triplet corresponding to the i-th original student information image, is the y2 coordinate of the v1-1th target triplet corresponding to the i-th original student information image or the y2 coordinate of the v2-1th target triplet corresponding to the i-th original student information image; The structured marked text corresponding to each of the original student information images is obtained through the multiple target triples corresponding to each of the original student information images.
8. The student data desensitization method according to any one of claims 1 to 7, characterized in that: The process of performing masking analysis on each of the original student information images according to each of the structured markup texts, the corrected student information images corresponding to each of the original student information images, and the column width parameter sequences corresponding to each of the original student information images, and using the analysis results as student data desensitization results includes: Counting the number of parameters in each column width parameter sequence respectively to obtain the number of column width parameters corresponding to each original student information image; Determine whether the number of column width parameters is 0, If yes, the width of the corrected student information image corresponding to the number of column width parameters is used as the mask width; if no, all parameters in the column width parameter sequence corresponding to the number of column width parameters are added together, and the result of the addition is used as the mask width; performing masking processing on each of the original student information images according to each of the structured markup texts and the masking width corresponding to each of the original student information images, to obtain a first masking area corresponding to each of the original student information images; Using a pixel mask algorithm to mask the remaining areas of each of the original student information images, respectively, to obtain second masked areas corresponding to each of the original student information images; The desensitized images corresponding to each of the original student information images are obtained through the first masked areas corresponding to each of the original student information images and the second masked areas corresponding to each of the original student information images, and all of the desensitized images are taken as the student data desensitization results.
9. A student data desensitization device, characterized in that: include: An import module, configured to import a plurality of original student information images and a list of student names corresponding to each of the original student information images; a correction and analysis module, configured to perform correction and analysis on each of the original student information images to obtain a corrected student information image corresponding to each of the original student information images; a segmentation module for segmenting each of the corrected student information images to obtain a text feature image corresponding to each of the original student information images and a table structure image corresponding to each of the original student information images; A coordinate analysis module, configured to perform coordinate analysis on the text feature images corresponding to the original student information images according to the student name lists, to obtain a target coordinate vector set corresponding to each original student information image; a table parameter analysis module, configured to perform table parameter analysis on each of the table structure images according to each of the student name lists and the corrected student information images corresponding to each of the original student information images, to obtain a row height parameter sequence corresponding to each of the original student information images and a column width parameter sequence corresponding to each of the original student information images; a registration analysis module, configured to perform registration analysis on each target coordinate vector set, the student name list corresponding to each original student information image, and the row height parameter sequence corresponding to each original student information image, to obtain a structured marked text corresponding to each original student information image; The desensitization result acquisition module is used to perform masking analysis on each of the original student information images according to each of the structured marked texts, the corrected student information images corresponding to each of the original student information images, and the column width parameter sequence corresponding to each of the original student information images, and use the analysis results as the student data desensitization results.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the student data desensitization method as described in any one of claims 1 to 8 is implemented.