Artificial intelligence-based method for digitizing paper textbooks

By extracting sample images of key areas of paper textbooks of different subject types and calculating sensitive feature representation parameters and sorting sequences, the problem of slow review of textbooks of different subject types is solved, and efficient and accurate paper textbook publication review is achieved.

CN120279558BActive Publication Date: 2025-10-17GUANGDONG PUBLICATION GRP DIGITAL PUBLICATION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510766924.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-17
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing technology does not take into account the data representation of different key areas of paper textbooks of different subject types in the complex dimensions of image recognition, resulting in slow review and publication of massive paper textbooks and low efficiency.

Method used

Extract sample images of different key areas of paper textbooks of different subject types, calculate sensitive feature representation parameters, and quickly identify abnormal areas of paper textbooks through sensitive feature sorting sequence and similarity analysis to optimize the review process.

Benefits of technology

It improves the efficiency and accuracy of pre-publication review of paper textbooks, reduces noise data, and ensures the accuracy and efficiency of the review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279558B_ABST
    Figure CN120279558B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image recognition, and more particularly to a paper teaching material publishing digitization method based on artificial intelligence, which extracts sample image information of different key areas of various subject paper teaching materials to establish a database, determines characteristic influence factors of various key areas of different subject paper teaching materials, calculates sensitive characteristic representation parameters of various key areas of various subject paper teaching materials, quickly determines a sensitive characteristic sorting sequence of various key areas of various subject paper teaching materials, accurately obtains whether there is an abnormality in each key area of the uploaded paper teaching material by acquiring actual images of the paper teaching material uploaded by a user end and sample images of corresponding key areas, improves the auditing and publishing precision of paper teaching materials, and overcomes the problem of low auditing and publishing efficiency and low precision of corresponding subject paper teaching materials caused by different sensitive characteristics of different key areas of different subject types of paper teaching materials in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and particularly relates to a paper teaching material publishing digitization method based on artificial intelligence. BACKGROUND

[0002] The rise of artificial intelligence technology has brought new breakthroughs in the recognition of paper teaching materials. Deep learning algorithms, such as convolutional neural networks (CNN), have achieved remarkable results in image recognition. Through a large amount of labeled data training, the model can automatically learn the features in the image, thereby achieving efficient recognition of paper teaching materials. In the teaching system, artificial intelligence image recognition technology is widely used in the digitization of teaching material content, the integration of teaching resources, and intelligent tutoring, etc. For example, by recognizing the knowledge points in the teaching materials, personalized learning suggestions and tutoring are provided for students.

[0003] With the development of digital technology, the recognition technology of paper teaching materials is also constantly improving. In the early stage, it was mainly based on simple optical character recognition (OCR) technology, which scanned and recognized the text on paper teaching materials and converted them into editable text format. Although this method can process text information, it has limited processing capacity for complex content such as images and charts in teaching materials. Later, image processing-based recognition technology appeared, which can comprehensively process the layout, images, and text of teaching materials, improving the accuracy and completeness of recognition. For example, through image segmentation technology, different elements in the page are separated and recognized and analyzed respectively.

[0004] Chinese patent publication number CN113569540B discloses a test paper generation method and device based on social science teaching materials. The method includes identifying paper social science teaching materials to obtain corresponding electronic documents; finding multiple text segments (composed of one or more continuous sentences in the electronic document) that meet the text characteristics in the electronic document; using a semantic recognition model to identify text segments containing professional terms in the multiple text segments; generating at least one question according to each text segment containing professional terms; extracting multiple knowledge points from the directory of the electronic document and each question; combining the knowledge points and the location information of the text segments used to generate the questions to construct a knowledge graph through knowledge fusion; according to the specified test paper generation parameters, searching the knowledge graph to obtain a set of test paper generation questions that meet the test paper generation parameters, and combining the multiple questions in the set of test paper generation questions into a test paper. This scheme automatically generates questions from teaching materials, effectively improving the efficiency of automatic test paper generation.

[0005] Chinese patent publication number: CN114220305A discloses a teaching system based on artificial intelligence image recognition technology, which includes an image input module, an AI recognition module, a teaching knowledge point input module, a modification and input module, a label and index module, a touch panel terminal, a data interaction module, and a knowledge point display module. The application applies AI technology and image recognition technology to modern teaching systems. Students organize knowledge points in related teaching materials. For related difficult points or unknown knowledge points in the teaching materials, image recognition technology and AI technology are used to identify and calculate difficult points or unknown knowledge points on paper teaching materials to obtain corresponding answers from the database without the need for users to spend time organizing and manually entering.

[0006] However, the prior art still has the following problems:

[0007] In the existing paper teaching material recognition method and artificial intelligence image recognition teaching system, the data representation of different key areas of different subject type paper teaching materials in image recognition complex dimension is not considered. In actual situations, some subject type paper teaching materials have more obvious characteristics in some key areas, and the area characteristics of some subject type paper teaching materials are more similar to the corresponding key areas of other paper teaching materials. If the same analysis method is used to audit whether to publish each type of paper teaching material, the publishing speed is slow and the efficiency is low when facing a large amount of paper teaching material audit and publishing data information. SUMMARY

[0008] Therefore, the present application provides a paper teaching material publishing digitization method based on artificial intelligence to overcome the problem that the prior art does not consider the data representation of different key areas of different subject type paper teaching materials in image recognition complex dimension, and the same analysis method is used to audit whether to publish each type of paper teaching material, resulting in slow publishing speed and low efficiency when facing a large amount of paper teaching material audit and publishing data information.

[0009] To achieve the above purpose, the present application provides a paper teaching material publishing digitization method based on artificial intelligence, which comprises:

[0010] Extracting sample images of different key areas of different subject type paper teaching materials to determine the characteristic influence factors of each key area of different subject type paper teaching materials;

[0011] According to the characteristic influence factors of each key area of different subject type paper teaching materials, the sensitive feature representation parameters of each subject type paper teaching material are calculated to determine the sensitive feature sorting sequence of each key area of a single subject type paper teaching material;

[0012] The sample image of the paper teaching material uploaded by the user terminal and the corresponding teaching material subject type are acquired, the sensitive feature representation parameters of each key area of the paper teaching material of the corresponding subject type are acquired according to the actual teaching material subject type, the abnormal tendency label of the paper teaching material is divided, the actual image of the paper teaching material is analyzed according to the determination result, including,

[0013] If it is the first abnormal tendency label, the sensitive feature sorting sequence of each key area of the paper teaching material of the corresponding subject type is acquired, the comparison order of the local sample image of the key area of the paper teaching material to be published is determined according to the sensitive feature sorting sequence, the corresponding key area of the paper teaching material of the corresponding subject type is compared with the corresponding key area of the actual paper teaching material according to the comparison order, and the similarity is determined to determine whether there is an abnormality.

[0014] If it is the second abnormal tendency label, the complete sample image of the paper teaching material is compared with the sample image of the corresponding complete paper teaching material to determine the similarity, so as to determine whether there is an abnormality.

[0015] Further, the process of determining the feature influence factor of each key area of the paper teaching material of different subject types includes,

[0016] The sample image of each key area of a single subject type is compared with the local sample image of the paper teaching material of the corresponding key area of each subject type.

[0017] The color saturation is solved, and the edge complexity of each key area of the paper teaching material of a single subject type is acquired.

[0018] Further, the edge complexity of each key area of the paper teaching material of a single subject type includes,

[0019] The edge contour line of each key area of the paper teaching material of a single subject type is calibrated, and the edge texture density is extracted.

[0020] The coincidence degree of the edge contour line in the paper teaching material and the reference contour line is calculated, and the ratio of the coincidence degree and the reference contour coincidence threshold value is taken as the first edge complexity influence factor.

[0021] The ratio of the edge texture density in the paper teaching material and the reference texture density threshold value is taken as the second edge complexity influence factor.

[0022] The sum of the first edge complexity influence factor and the second edge complexity influence factor is determined as the edge complexity.

[0023] Further, the process of calculating the sensitive feature representation parameters of the paper teaching material of a single subject type includes,

[0024] The feature influence factor of the single-subject type paper textbook is obtained, including color saturation mean value and edge complexity;

[0025] A ratio of a predetermined color saturation threshold value to the color saturation mean value is determined as a first feature influence factor;

[0026] A ratio of the edge complexity to a predetermined edge complexity threshold value is determined as a second feature influence factor;

[0027] A sum of the first feature influence factor and the second feature influence factor is determined as the sensitive feature representation parameter.

[0028] Further, the process of determining the sensitive feature sorting sequence of each key area of the single-subject type paper textbook includes,

[0029] The sensitive feature representation parameter corresponding to each key area of the single-subject type paper textbook is determined;

[0030] The sensitive feature representation parameters are arranged in descending order to obtain the sensitive feature sorting sequence.

[0031] Further, the process of dividing the abnormal tendency label of the paper textbook includes,

[0032] The significant sensitive feature representation parameter of each key area of the corresponding subject type paper textbook is extracted;

[0033] If the significant sensitive feature representation parameter is greater than a predetermined sensitive feature representation parameter threshold value, it is determined as a first abnormal tendency label;

[0034] If the significant sensitive feature representation parameter is less than or equal to the predetermined sensitive feature representation parameter threshold value, it is determined as a second abnormal tendency label.

[0035] Further, the comparison order of the actual image of the paper textbook to be published is determined according to the sensitive feature sorting sequence, including,

[0036] According to the sensitive feature sorting sequence, the corresponding key area is determined in sequence, and the key area is compared with the actual image;

[0037] Each sensitive feature representation parameter corresponds to a key area number one by one.

[0038] Further, according to the comparison order, the key area of the corresponding subject type paper textbook is called in sequence and compared with the corresponding key area of the actual paper textbook to determine the similarity, so as to determine whether there is an abnormality, including,

[0039] If the similarity corresponding to each key area actual image is less than or equal to a predetermined key area similarity threshold value, it is determined that there is an abnormality.

[0040] Further, the complete sample image of the paper-based teaching material is acquired and compared with the sample image corresponding to the complete paper-based teaching material to determine the similarity, so as to determine whether there is an anomaly, including,

[0041] The complete sample image of the paper-based teaching material uploaded by the user terminal is acquired and compared with the complete sample image of the paper-based teaching material corresponding to the subject type, and the similarity is determined.

[0042] If the similarity between the complete sample image of the paper-based teaching material and the sample image corresponding to the complete paper-based teaching material is less than or equal to a predetermined complete sample image similarity threshold, it is determined that there is an anomaly.

[0043] Further, the sample image of the paper-based teaching material uploaded by the user terminal needs to include a complete sample image of the paper-based teaching material and a local sample image of each key area of the paper-based teaching material.

[0044] Compared with the prior art, the present application establishes a database by extracting sample image information of different key areas of each subject type paper-based teaching material, and can calculate sensitive feature representation parameters of each key area of each subject type paper-based teaching material by determining feature influence factors of each key area of each subject type paper-based teaching material, so as to quickly determine the sensitive feature sorting sequence of each key area of each subject type paper-based teaching material, thereby improving the auditing efficiency before the publication of the paper-based teaching material. At the same time, by acquiring the actual image of the paper-based teaching material uploaded by the user terminal and the sample image corresponding to the key area, the significant sensitive feature representation parameters of each key area of the uploaded paper-based teaching material can be accurately acquired, thereby improving the auditing and publishing accuracy of the paper-based teaching material. The problem of low auditing and publishing efficiency and low accuracy of the corresponding subject paper-based teaching material caused by different sensitive features of different key areas of different subject type paper-based teaching materials in the prior art is overcome. By determining the comparison order of the actual image of each key area of the paper-based teaching material, and sequentially calling the actual image of the corresponding key area and the sample image of the corresponding key area for comparison according to the comparison order, whether the paper-based teaching material uploaded by the user terminal is abnormal and whether to enable publishing can be accurately determined.

[0045] Especially, the application compares sample images of different key areas of single subject paper teaching materials with corresponding key area image information of other subject paper teaching materials, comprehensively considers the differences in paper teaching material publishing review of different key areas, and in actual situations, some subject paper teaching materials have more obvious sensitive features in some key areas, and some sensitive features of key areas of some subject type paper teaching materials are more consistent with corresponding key areas of other subject types. Therefore, the sensitive feature representation of the sample images of each key area is different, and in some cases, the sample images of the key areas with obvious sensitive feature representation can identify whether there is an anomaly. Therefore, the application considers to determine the sensitive feature representation parameters of each key area, provides data support for selecting sensitive feature representation parameters when subsequent actual image of uploaded paper teaching material is reviewed and published, and then adaptively selects the actual review and publishing analysis, reduces the review processing amount under the premise of ensuring accuracy, and improves the publishing review efficiency of paper teaching materials.

[0046] Especially, the application obtains the color saturation and edge complexity of each key area of single subject type paper teaching materials, which provides an important feature dimension for whether there is an anomaly in paper teaching material publishing review. The edge complexity reflects the structural features of each key area of paper teaching materials, and different subject type attribute factors may cause different edge complexities of paper teaching materials of different subject types. The edge complexity and color saturation are combined to represent the sensitive feature influence situation and data representation of the key area.

[0047] Especially, the application analyzes the actual image in the case that the significant sensitive feature representation parameter is greater than the predetermined sensitive feature representation parameter threshold. In the above case, the key area of the paper teaching material of the corresponding subject type has a more prominent sensitive feature, and the data representation is strong. Therefore, the comparison order of the actual image of each key area of the paper teaching material of the subject type is determined according to the sensitive feature sorting sequence, and the actual image of the corresponding key area is compared with the local sample image of the corresponding key area according to the comparison order. This can reduce the noise data introduced by full area sample image comparison, and preferentially analyze the actual image of the area with strong image representation under the premise of ensuring reliability, improve the publishing review efficiency of paper teaching materials, and ensure the accuracy of paper teaching material publishing review. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 The step flowchart of the paper teaching material publishing digitization method based on artificial intelligence of the embodiment of the application;

[0049] Figure 2 The step flowchart of obtaining the edge complexity of each key area of single subject type paper teaching materials of the embodiment of the application;

[0050] Figure 3 A flow chart of steps for determining feature influence factors of each key area of paper teaching materials of different subject types for the embodiments of the present application is shown in the figure;

[0051] Figure 4 A logic determination diagram for dividing the abnormal tendency labels of paper teaching materials for the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION

[0052] In order to make the objects and advantages of the present application clearer, the present application will be further described below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0053] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the protection scope of the present application.

[0054] Please refer to Figure 1 The figure is a step flow chart of paper teaching material publishing digitization based on artificial intelligence for the embodiments of the present application. The present application provides a paper teaching material publishing digitization method based on artificial intelligence, which comprises:

[0055] Step S1, sample images of different key areas of paper teaching materials of different subject types are extracted to determine feature influence factors of each key area of paper teaching materials of different subject types;

[0056] Step S2, sensitive feature representation parameters of paper teaching materials of each subject type are calculated according to the feature influence factors of each key area of paper teaching materials of different subject types to determine the sensitive feature sorting sequence of each key area of a single subject type paper teaching material;

[0057] Step S3, sample images of paper teaching materials uploaded by a user terminal and corresponding teaching subject types are obtained, sensitive feature representation parameters of each key area of paper teaching materials of the corresponding subject type are obtained according to the actual teaching subject type, and the abnormal tendency label of the paper teaching material is divided. According to the determination result, the actual image of the paper teaching material is analyzed, including,

[0058] If it is the first abnormal tendency label, the sensitive feature sorting sequence of each key area of paper teaching materials of the corresponding subject type is obtained, the comparison order of the local sample image of the key area of the paper teaching material to be published is determined according to the sensitive feature sorting sequence, and the corresponding key area of the paper teaching material of the corresponding subject type is called in turn according to the comparison order to compare with the corresponding key area of the actual paper teaching material to determine the similarity and determine whether there is an abnormality.

[0059] If it is the second abnormal tendency label, a complete sample image of the paper teaching material is obtained and compared with a sample image corresponding to the complete paper teaching material to determine a similarity, so as to determine whether there is an abnormality.

[0060] Specifically, the subject types of the paper teaching materials are generally divided into Chinese, mathematics, biology, politics, physics, chemistry, biology, and history according to subjects, or can be divided into first grade, second grade, third grade, and so on, n (n is a positive integer), or other forms can be used, and in a preferred implementation, the subject division is used, which will not be repeated here.

[0061] Specifically, the division of the key regions of the paper teaching materials is not limited, and in the implementation, the key regions can be divided into formula regions, curve graph regions, table regions, and illustration regions, or other forms can be used, which will not be repeated here.

[0062] Specifically, the present application establishes a database by extracting sample image information of different key regions of each subject type of paper teaching material, and can calculate sensitive feature representation parameters of each key region of each subject type of paper teaching material by determining feature influence factors of each key region of different subject types of paper teaching material, so as to quickly determine the sensitive feature sorting sequence of each key region of each subject type of paper teaching material, improve the auditing efficiency of paper teaching materials before publication, at the same time, the significant sensitive feature representation parameters of each key region of the uploaded paper teaching material can be accurately obtained by obtaining the actual image of the paper teaching material uploaded by the user end and the sample image corresponding to the key region, and the auditing and publishing accuracy of the paper teaching material is improved, and the problem that the sensitive features of different key regions of different subject types of paper teaching materials are different, resulting in low auditing and publishing efficiency and low accuracy of corresponding subject paper teaching materials is overcome. By determining the comparison order of the actual image of each key region of the paper teaching material, and sequentially calling the corresponding key region actual image and the corresponding key region sample image for comparison, it can be accurately determined whether the paper teaching material uploaded by the user end is abnormal, and whether to enable publishing.

[0063] Specifically, the determination method of the similarity is not limited, the corresponding data information processing algorithm or model can be imported into the logic component to realize the corresponding function, and the data information processing algorithm is not limited, for example, the similarity index calculation method can be used, or a pre-trained deep learning model can be used to extract the feature representation of the data information, and then the cosine similarity between the features is calculated as the similarity, of course, other methods can also be used, which will not be repeated here.

[0064] Please refer to Figure 2 As shown in the figure, it is a step flow chart for determining the feature influence factors of each key region of different subject types of paper teaching materials, in step S1, the process of determining the feature influence factors of each key region of different subject types of paper teaching materials includes,

[0065] comparing the sample image of each key area of the single subject type with the sample image of the corresponding key area of the paper teaching material of each subject type;

[0066] solving the color saturation, and obtaining the edge complexity of each key area of the paper teaching material of the single subject type.

[0067] Specifically, for the color saturation, the OpenCV or Pillow library in Python can be used to calculate the color saturation between the sample image of each key area and the sample image of the corresponding key area of each subject, or other forms can be used, which will not be repeated here.

[0068] Please refer to Figure 3 The step flow chart for obtaining the edge complexity of each key area of the paper teaching material of the single subject type is shown in the figure, and in step S1, the process of obtaining the edge complexity of each key area of the paper teaching material of the single subject type includes,

[0069] calibrating the edge contour line of each key area of the paper teaching material of the single subject type and extracting the edge texture density;

[0070] calculating the coincidence degree of the edge contour line in the paper teaching material and the reference contour line, and taking the ratio of the coincidence degree and the reference contour coincidence threshold value as the first edge complexity influence factor;

[0071] calculating the ratio of the edge texture density in the paper teaching material and the reference texture density threshold value as the second edge complexity influence factor;

[0072] determining the sum of the first edge complexity influence factor and the second edge complexity influence factor as the edge complexity;

[0073] The coincidence degree is the ratio of the total area of the overlapping part of the edge contour in the paper teaching material and the reference contour to the area of the reference contour.

[0074] Specifically, the reference contour coincidence threshold value and the reference texture density threshold value are obtained by pre-setting, wherein the average value of the ratio of the total area of the overlapping part of the edge contour and the basic contour in the historical period of three months of the system to the area of the reference contour is multiplied by the precision coefficient to obtain the reference contour coincidence threshold value, and the precision coefficient is between the interval [0.92, 0.98], and the average value of the edge texture density in the historical period of three months of the system is multiplied by the deviation coefficient to obtain the reference texture density threshold value, and the deviation coefficient is between the interval [0.95, 0.99].

[0075] The present application provides an important feature dimension for the paper textbook publishing review whether there is an anomaly by obtaining the edge profile coincidence degree and the edge texture density of each key area of a single subject type paper textbook. The edge complexity reflects the structural characteristics of each key area of the paper textbook. Different subject types may have different edge complexity due to different attribute factors. The edge profile coincidence degree and the edge texture density are combined to represent the complex feature influence of the key area.

[0076] Specifically, the process of calculating the sensitive feature representation parameter of a single subject type paper textbook includes,

[0077] The feature influence factor of a single subject type paper textbook includes the color saturation mean value and the edge complexity.

[0078] The ratio of the predetermined color saturation threshold value to the color saturation mean value is determined as the first feature influence factor.

[0079] The ratio of the edge complexity to the predetermined edge complexity threshold value is determined as the second feature influence factor.

[0080] The sum of the first feature influence factor and the second feature influence factor is determined as the sensitive feature representation parameter.

[0081] In implementation, the predetermined color saturation threshold value is pre-set, wherein the color saturation mean value of different key areas of each subject type paper textbook is pre-determined, the color saturation threshold value is set as the product of the color saturation mean value and a precision coefficient, and the precision coefficient is selected within the interval [0.90, 0.95].

[0082] In implementation, the predetermined edge complexity threshold value is pre-set, wherein the edge complexity of different key areas of each subject type paper textbook is pre-determined, the edge complexity mean value is calculated, and the edge complexity threshold value is set as the product of the edge complexity mean value and a complexity offset coefficient, and the complexity offset coefficient is selected between the interval [1.10, 1.15].

[0083] The application compares sample images of different key areas of single subject paper teaching materials with corresponding key area image information of other subject paper teaching materials, comprehensively considers the differences in paper teaching material publishing review of different key areas, and in actual situations, some subject paper teaching materials have more obvious sensitive features in some key areas, and some sensitive features of key areas of some subject type paper teaching materials are more consistent with corresponding key areas of other subject types, therefore, the sensitive feature representation of sample images of each key area is different, in some cases, the sample image of the key area with obvious sensitive feature representation can identify whether there is an anomaly, therefore, the application considers to determine the sensitive feature representation parameter of each key area, provides data support for selecting the sensitive feature representation parameter mode when subsequent paper teaching material actual image uploading is reviewed and published, and then adaptively selects the actual review and publishing analysis, reduces the review processing amount under the premise of ensuring accuracy, and improves the paper teaching material publishing review efficiency.

[0084] Specifically, the process of determining the sensitive feature sorting sequence of each key area of a single subject type paper teaching material includes,

[0085] Determine the sensitive feature representation parameter corresponding to each key area of a single subject type paper teaching material.

[0086] The sensitive feature representation parameters are arranged from large to small to obtain the sensitive feature sorting sequence.

[0087] Please refer to Figure 4 The application is an embodiment of the application, which is a logical judgment schematic diagram for dividing the abnormal tendency label of paper teaching materials, and the process of dividing the abnormal tendency label of paper teaching materials includes,

[0088] Extract the significant sensitive feature representation parameter of each key area of the corresponding subject type paper teaching material.

[0089] If the significant sensitive feature representation parameter is greater than the predetermined sensitive feature representation parameter threshold, the first abnormal tendency label is determined.

[0090] If the significant sensitive feature representation parameter is less than or equal to the predetermined sensitive feature representation parameter threshold, the second abnormal tendency label is determined.

[0091] In implementation, the significant sensitive feature representation parameter is the maximum sensitive feature representation parameter, and the predetermined sensitive feature representation parameter threshold is in the interval [2.15, 2.45].

[0092] Specifically, the comparison order of the actual image of the paper teaching material to be published is determined according to the sensitive feature sorting sequence, including,

[0093] According to the sensitive feature ranking sequence, the corresponding key region is determined in sequence, and the key region is compared with the actual image.

[0094] The sensitive feature characteristic parameters correspond to the key region sequence one by one.

[0095] In the implementation, optionally,

[0096] The key regions can be assigned sequence numbers. Taking four key regions as an example, 1-formula region, 2-graph region, 3-table region, and 4-illustration region; formula region, graph region, table region, and illustration region

[0097] For example, the sensitive feature characteristic parameter of the formula region is 2.5, the sensitive feature characteristic parameter of the graph region is 2.65, the sensitive feature characteristic parameter of the table region is 2.55, and the sensitive feature characteristic parameter of the illustration region is 2.45.

[0098] The corresponding key region ranking sequence is 2.65, 2.55, 2.5, and 2.45.

[0099] The corresponding key region ranking is 2, 3, 1, and 4.

[0100] Specifically, according to the comparison sequence, the corresponding subject type paper teaching material key region and the actual paper teaching material corresponding key region are compared in sequence to determine the similarity, so as to determine whether there is an anomaly, including,

[0101] If the similarity corresponding to any key region actual image is greater than the predetermined key region similarity threshold, it is determined that there is no anomaly, and the paper teaching material is published.

[0102] If the similarity corresponding to each key region actual image is less than or equal to the predetermined key region similarity threshold, it is determined that there is an anomaly, and the paper teaching material is not published, and the paper teaching material is corrected and re-uploaded for analysis.

[0103] In the implementation, the similarity threshold of the key region is obtained by pre-setting, a plurality of sample image information of the same key region of the same subject type paper teaching material is obtained in advance, the average similarity between the sample image information is determined, and the ratio of the average similarity to the offset coefficient is determined as the similarity threshold of the key region. The local image offset coefficient is selected between the interval [0.85, 0.95].

[0104] Specifically, the complete sample image of the paper teaching material is obtained, and the sample image corresponding to the complete paper teaching material is compared to determine the similarity, so as to determine whether there is an anomaly, including,

[0105] The complete sample image of the paper teaching material uploaded by the user terminal is compared with the complete sample image of the corresponding subject type paper teaching material to determine the similarity, so as to determine whether to enable the paper teaching material publishing;

[0106] If the similarity between the current complete sample image of the paper teaching material and the sample image of the corresponding complete paper teaching material is less than or equal to the predetermined complete sample image similarity threshold, it is determined that there is an abnormality, and the paper teaching material publishing is not enabled, and the paper teaching material is corrected and then re-uploaded and analyzed;

[0107] If the similarity between the current complete sample image of the paper teaching material and the sample image of the corresponding complete paper teaching material is greater than the predetermined complete sample image similarity threshold, it is determined that there is no abnormality, and the paper teaching material publishing is enabled.

[0108] Specifically, the sample similarity threshold of the complete paper teaching material sample image is obtained by prior setting, a plurality of complete sample image information of all key regions of the paper teaching material of the same subject is obtained in advance, the average similarity between the complete sample image information is determined, and the ratio of the average similarity and an offset coefficient is determined as the key region similarity threshold. The complete image offset coefficient is selected in the interval [1.15, 1.25].

[0109] Specifically, the sample image of the paper teaching material uploaded by the user terminal needs to include a complete sample image of the paper teaching material and a local sample image of each key region of the paper teaching material.

[0110] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical solutions after the changes or replacements will fall within the protection scope of the present application.

Claims

1. A method for digitalizing paper textbook publishing based on artificial intelligence, characterized in that: include: Extract sample images of different key areas of paper textbooks of different subject types to determine the characteristic influencing factors of each key area of ​​paper textbooks of different subject types; Calculate the sensitive feature characterization parameters of each subject type of paper textbook based on the feature impact factors of each key area of ​​the paper textbook of different subject types, so as to determine the sensitive feature ranking sequence of each key area of ​​the paper textbook of a single subject type; Obtain sample images of paper textbooks uploaded by the user and the corresponding textbook subject types, obtain sensitive feature representation parameters of key areas of the paper textbooks of the corresponding subject types based on the actual textbook subject types, and classify abnormal tendency labels of the paper textbooks. Analyze the actual images of the paper textbooks based on the judgment results. include, If it is the first abnormal tendency label, obtain the sensitive feature sorting sequence of each key area of ​​the paper textbook of the corresponding subject type, determine the comparison order of the local sample images of the key area of ​​the paper textbook to be published based on the sensitive feature sorting sequence, and sequentially call the key area of ​​the paper textbook of the corresponding subject type and the corresponding key area of ​​the actual paper textbook according to the comparison order to determine the similarity and whether there is any abnormality; If the similarity of the actual images of each key area is less than or equal to the predetermined key area similarity threshold, it is determined that there is an anomaly, and the paper textbook is activated for correction and re-uploaded for analysis; If it is the second abnormal tendency label, then obtain the complete sample image of the paper textbook and compare it with the sample image of the corresponding complete paper textbook to determine the similarity to determine whether there is an abnormality; The process of determining the characteristic influencing factors of key areas of paper textbooks of different subject types includes: Compare the sample images of each key area of ​​a single subject type with the partial sample images of the corresponding key areas of paper textbooks of other subject types; Calculate color saturation and obtain edge complexity of key areas of paper textbooks for a single subject type; Among them, obtaining the edge complexity of each key area of ​​the paper textbook of a single subject type includes: Calibrate the edge contours of key areas of paper textbooks for a single subject type and extract edge texture density; Calculating the degree of overlap between the edge contour line in the paper textbook and the reference contour line, and taking the ratio of the degree of overlap to the reference contour overlap threshold as the first edge complexity influencing factor; The ratio of the edge texture density in the paper textbook to the benchmark texture density threshold is calculated as the second edge complexity influencing factor; determining a sum of the first edge complexity influencing factor and the second edge complexity influencing factor as the edge complexity; The characteristic influencing factors of paper textbooks of a single subject type include color saturation mean and edge complexity; Among them, the key areas include the formula area, the graph area, the table area, and the illustration area.

2. The method for digitalizing paper textbook publication based on artificial intelligence according to claim 1, characterized in that: The process of calculating the sensitive feature representation parameters of a paper textbook of a single subject type includes: Determining a ratio of a predetermined color saturation threshold to a color saturation mean as a first characteristic influencing factor; Determining a ratio of the edge complexity to a predetermined edge complexity threshold as a second feature influencing factor; The sum of the first feature impact factor and the second feature impact factor is determined as the sensitive feature characterization parameter.

3. The method for digitalizing paper textbook publication based on artificial intelligence according to claim 1, characterized in that: The process of determining the ranking sequence of sensitive features of key areas of paper textbooks of a single subject type includes: Determine the sensitive feature representation parameters corresponding to each key area of ​​paper textbooks of a single subject type; Arrange the sensitive feature characterization parameters in descending order to obtain the sensitive feature sorting sequence.

4. The method for digitalizing paper textbook publication based on artificial intelligence according to claim 1, characterized in that: The process of classifying abnormal tendency labels of paper textbooks includes: Extract significant sensitive feature representation parameters of key areas of paper textbooks of corresponding subject types; If the significant sensitive feature characterization parameter is greater than a predetermined sensitive feature characterization parameter threshold, it is determined to be a first abnormal tendency label; If the significant sensitive feature characterization parameter is less than or equal to the predetermined sensitive feature characterization parameter threshold, it is determined to be a second abnormal tendency label.

5. The method for digitalizing paper textbook publication based on artificial intelligence according to claim 3, characterized in that: Determining a comparison order for actual images of paper teaching materials to be published based on the sensitive feature sorting sequence includes: Determine corresponding key areas in sequence according to the sensitive feature sorting sequence, and compare the key areas with the actual image; Among them, each sensitive feature characterization parameter corresponds one-to-one to the key area serial number.

6. The method for digitalizing paper textbook publication based on artificial intelligence according to claim 5, characterized in that: Obtain a complete sample image of the paper textbook and compare it with the sample image of the corresponding complete paper textbook to determine the similarity to determine whether there is any anomaly, including, Obtain the complete sample image of the paper textbook uploaded by the user and compare it with the complete sample image of the paper textbook of the corresponding subject type to determine the similarity; If the similarity between the current paper textbook complete sample image and the sample image of the corresponding complete paper textbook is less than or equal to a predetermined complete sample image similarity threshold, it is determined that an abnormality exists.

7. The method for digitalizing paper textbook publication based on artificial intelligence according to claim 6, characterized in that: The sample images of the paper textbooks uploaded by the user terminal must include complete sample images of the paper textbooks and local sample images of key areas of the paper textbooks.

Citation Information

Patent Citations

  • Test paper generation method and device based on social science textbooks

    CN113569540B

  • Teaching system based on artificial intelligence image recognition technology

    CN114220305A

  • Method for quantitatively evaluating perceived quality of digital printed lines and texts

    CN103076334A

  • Agricultural product tracing method and system based on data analysis

    CN119762090A