Method and system for automatic title labeling and classification
By photographing and text data on the questions, identifying English and Chinese texts and annotating subject types, the problem of difficult question positioning in the question bank is solved, and efficient and accurate automatic classification of questions is achieved.
Patent Information
- Application Number
- CN202011048811.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-29
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2040-09-29
AI Technical Summary
The lack of deep processing of the existing test question bank makes it impossible to quickly and accurately locate the required questions, affecting the practicality and reliability of use.
Automatic classification of questions is achieved by taking images, converting text data, identifying English and Chinese texts and labeling subject types.
It improves the efficiency and accuracy of the title and classification, reduces the calculation workload, and improves the convenience of subsequent analysis and processing.
Smart Images

Figure CN111985193B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent education, and in particular to a method and system for automatically labeling and classifying questions. Background Art
[0002] At present, in order to improve the efficiency and accuracy of the test paper formation process, the different types of questions in the test paper are usually formed by using a corresponding test question bank. However, the existing test question bank is formed based on the questions used in historical homework, tests, and exams. In order to save the time of forming the test question bank, the existing technology simply classifies the historical questions into simple assessment knowledge points and directly stores them in the test question bank without performing corresponding deep processing on the questions. Although this can effectively increase the question data volume and the knowledge point coverage of the questions in the test question bank, the lack of corresponding labeling and classification deep processing makes it impossible to quickly and accurately locate the required questions from the test question bank, which seriously affects the practicality and reliability of the test question bank. It can be seen that the existing technology urgently needs a processing method that can accurately and effectively automatically label and classify different types of questions. Summary of the Invention
[0003] In response to the defects of the existing technology, the present invention provides a method and system for automatic topic labeling and classification, which obtains image information about each target topic by photographing several target topics respectively, and performs text data conversion processing on the image information to obtain topic text data samples about several target topics, and performs text language information recognition processing on the topic text data samples to obtain English text information and Chinese text information corresponding to the topic text data samples, and then performs word type recognition processing on the English text information and Chinese text information to obtain subject type labeling information corresponding to the topic text data samples, and finally, according to the subject type labeling information, the several target topics are matched and divided into different topic sets. In the combination, automatic classification of several target topics is achieved; it can be seen that the method and system for automatic tagging and classification of topics photographs the target topics and converts the photographed images into corresponding topic text data, and identifies the English text and Chinese text respectively contained in the topic text data samples, and then obtains the corresponding subject type according to the vocabulary types respectively contained in the English text and the Chinese text, and performs adaptive tagging, and finally automatically classifies the target topics into the corresponding topic set according to the tagging results, so that a large number of target topics of different types can be automatically labeled and classified in a targeted and efficient manner, thereby improving the efficiency of deep processing of topics and facilitating subsequent analysis and processing of topics.
[0004] The present invention provides a method for automatically labeling and classifying topics, which is characterized by comprising the following steps:
[0005] Step S1: photographing a plurality of target topics respectively to obtain image information about each of the target topics, and performing text data conversion processing on the image information to obtain topic text data samples about the plurality of target topics;
[0006] Step S2, performing text language information recognition processing on the question text data sample, thereby obtaining English text information and Chinese text information corresponding to the question text data sample;
[0007] Step S3, performing word type recognition processing on the English text information and the Chinese text information, thereby obtaining subject type labeling information corresponding to the question text data sample;
[0008] Step S4, matching and classifying the target topics into different topic sets according to the subject type annotation information, thereby achieving automatic classification of the target topics;
[0009] Furthermore, in step S1, a plurality of target topics are photographed respectively to obtain image information about each target topic, and the image information is converted into text data to obtain topic text data samples about the plurality of target topics. Specifically, the following steps are performed:
[0010] Step S101, scanning and photographing each of the target topics to obtain a two-dimensional image of each of the target topics;
[0011] Step S102, performing pixel binarization processing and background noise reduction filtering processing on the two-dimensional image, thereby converting the two-dimensional image into a grayscale image;
[0012] Step S103: extracting corresponding question text character outline information from the grayscale image, and converting the grayscale image into question text data corresponding to the target question based on the question text character outline information, thereby forming a question text data sample from the question text data corresponding to all target questions;
[0013] Furthermore, in step S2, the text language information recognition process is performed on the question text data sample to obtain English text information and Chinese text information corresponding to the question text data sample, specifically including:
[0014] According to the following formula (1), the text language information recognition processing is performed on the subject text data sample to obtain the English text information A corresponding to the subject text data sample. n and Chinese text message B m ,
[0015]
[0016] In the above formula (1), Title(A n , B m ) represents the semantic approximation value of the title text composed of the semantic approximation value of the English text and the semantic approximation value of the Chinese text included in the title text data sample, π represents the pi, arctan represents the inverse tangent function operator, A n Indicates the text semantic approximation value corresponding to the nth English text in the title, B m represents the text semantic approximation value corresponding to the mth Chinese text, N represents the total number of English text data contained in the English text information, and its maximum value is 40, n is an integer between 1 and 40, M represents the total number of Chinese text data contained in the Chinese text information, and its maximum value is 20, m is an integer between 1 and 20, j represents any Chinese text character in the question text data sample, which is split into eight intervals according to the rice grid, and each interval is marked in a counterclockwise order in the right horizontal axis direction, and the value of j can only be 1, 2, 3, 4, 5, 6, 7, 8, l i represents the horizontal length of the jth interval of any Chinese text character, h j represents the vertical length corresponding to the j-th interval of any Chinese text character, Indicates the horizontal stroke space vector corresponding to any one of the Chinese text characters, represents the vertical stroke space vector corresponding to any one of the Chinese characters, f(a) represents the character area value corresponding to any one of the English characters in the question text data sample, Represents the recognition result of the English text characters of the subject text data sample;
[0017] as well as,
[0018] In step S3, word type recognition processing is performed on the English text information and the Chinese text information to obtain subject type annotation information corresponding to the question text data sample, specifically including:
[0019] According to the following formula (2), the English text information and the Chinese text information are processed for word type recognition to obtain the subject type labeling information corresponding to the question text data sample.
[0020]
[0021] In the above formula (2), Match(q, d) represents the subject type label matching value corresponding to the question text data sample, Q represents the total number of subjects contained in the question text data sample, D represents the total number of English words and Chinese words contained in the question text data sample, q represents any positive integer between [1, Q], and d represents any positive integer between [1, D].
[0022] Furthermore, in step S4, the target topics are matched and divided into different topic sets according to the subject type annotation information, thereby achieving automatic classification of the target topics. Specifically, the method includes:
[0023] Step S401: According to the following formula (3) and the subject type annotation information, the matching value Disp(o) between each target topic and the corresponding keyword in the preset classification keyword library is determined. r ,t r ),
[0024]
[0025] In the above formula (3), o r Indicates the number of characters corresponding to the rth keyword, t r Indicates the character bit length corresponding to the rth keyword, where r is any positive integer greater than or equal to 1;
[0026] Step S402: When the matching value Disp( r ,t r ) is equal to 1, indicating that the current target topic matches the current keyword, and the current target topic is divided into the topic set corresponding to the current keyword, thereby realizing automatic classification of the target topic.
[0027] The present invention also provides a system for automatic question labeling and classification, which is characterized by comprising a target question shooting module, a question text data sample acquisition module, a question English text / Chinese text information acquisition module, a subject type labeling information acquisition module and a question automatic classification module; wherein,
[0028] The target topic shooting module is used to shoot a plurality of target topics respectively, so as to obtain image information about each of the target topics;
[0029] The topic text data sample acquisition module is used to perform text data conversion processing on the image information, thereby obtaining topic text data samples about a number of the target topics;
[0030] The question English text / Chinese text information acquisition module is used to perform text language information recognition processing on the question text data sample, thereby obtaining English text information and Chinese text information corresponding to the question text data sample;
[0031] The subject type annotation information acquisition module is used to perform word type recognition processing on the English text information and the Chinese text information, so as to obtain the subject type annotation information corresponding to the question text data sample;
[0032] The automatic topic classification module is used to match and classify the target topics into different topic sets according to the subject type annotation information, thereby realizing automatic classification of the target topics;
[0033] Furthermore, the target topic shooting module shoots a plurality of target topics respectively to obtain image information about each target topic, specifically including:
[0034] Scanning and photographing each of the target topics to obtain a two-dimensional image of each of the target topics;
[0035] as well as,
[0036] The topic text data sample acquisition module performs text data conversion processing on the image information to obtain topic text data samples about several target topics, specifically including:
[0037] Performing pixel binarization processing and background noise reduction filtering processing on the two-dimensional image, thereby converting the two-dimensional image into a grayscale image;
[0038] Then, extracting corresponding question text character outline information from the grayscale image, and converting the grayscale image into question text data corresponding to the target question based on the question text character outline information, thereby forming a question text data sample from the question text data corresponding to all target questions;
[0039] Furthermore, the question English text / Chinese text information acquisition module performs text language information recognition processing on the question text data sample to obtain English text information and Chinese text information corresponding to the question text data sample, specifically including:
[0040] According to the following formula (1), the text language information recognition processing is performed on the subject text data sample to obtain the English text information A corresponding to the subject text data sample. n and Chinese text message B m ,
[0041]
[0042] In the above formula (1), Title(A n , B m ) represents the semantic approximation value of the title text composed of the semantic approximation value of the English text and the semantic approximation value of the Chinese text included in the title text data sample, π represents the pi, arctan represents the inverse tangent function operator, A n Indicates the text semantic approximation value corresponding to the nth English text in the title, B m represents the text semantic approximation value corresponding to the mth Chinese text, N represents the total number of English text data contained in the English text information, and its maximum value is 40, n is an integer between 1 and 40, M represents the total number of Chinese text data contained in the Chinese text information, and its maximum value is 20, m is an integer between 1 and 20, j represents any Chinese text character in the question text data sample, which is split into eight intervals according to the rice grid, and each interval is marked in a counterclockwise order in the right horizontal axis direction, and the value of j can only be 1, 2, 3, 4, 5, 6, 7, 8, l i represents the horizontal length of the jth interval of any Chinese text character, h j represents the vertical length corresponding to the j-th interval of any Chinese text character, Indicates the horizontal stroke space vector corresponding to any one of the Chinese text characters, represents the vertical stroke space vector corresponding to any one of the Chinese characters, f(a) represents the character area value corresponding to any one of the English characters in the question text data sample, Represents the recognition result of the English text characters of the subject text data sample;
[0043] as well as,
[0044] The subject type annotation information acquisition module performs word type recognition processing on the English text information and the Chinese text information to obtain the subject type annotation information corresponding to the question text data sample, specifically including:
[0045] According to the following formula (2), the English text information and the Chinese text information are processed for word type recognition to obtain the subject type labeling information corresponding to the question text data sample.
[0046]
[0047] In the above formula (2), Match(q, d) represents the subject type label matching value corresponding to the question text data sample, Q represents the total number of subjects contained in the question text data sample, D represents the total number of English words and Chinese words contained in the question text data sample, q represents any positive integer between [1, Q], and d represents any positive integer between [1, D].
[0048] Furthermore, the automatic topic classification module matches and divides the target topics into different topic sets according to the subject type annotation information, thereby achieving automatic classification of the target topics. Specifically, the automatic classification includes:
[0049] According to the following formula (3) and the subject type annotation information, the matching value Disp(o) between each target topic and the corresponding keyword in the preset classification keyword library is determined. r ,t r ),
[0050]
[0051] In the above formula (3), o r Indicates the number of characters corresponding to the rth keyword, t r Indicates the character bit length corresponding to the rth keyword, where r is any positive integer greater than or equal to 1;
[0052] And when the matching value Disp(o r ,t r ) is equal to 1, indicating that the current target topic matches the current keyword, and the current target topic is divided into the topic set corresponding to the current keyword, thereby realizing automatic classification of the target topic.
[0053] Compared with the existing technology, the method and system for automatic topic labeling and classification takes photos of several target topics respectively to obtain image information about each target topic, and performs text data conversion processing on the image information to obtain topic text data samples about several target topics, and performs text language information recognition processing on the topic text data samples to obtain English text information and Chinese text information corresponding to the topic text data samples, and then performs word type recognition processing on the English text information and Chinese text information to obtain subject type labeling information corresponding to the topic text data samples, and finally, according to the subject type labeling information, matches and divides several target topics into different topic sets, from And realize the automatic classification of several target topics; it can be seen that the method and system for automatic labeling and classification of topics shoots the target topics and converts the photographed images into corresponding topic text data, and identifies the English text and Chinese text contained in the topic text data samples respectively, and then obtains the corresponding subject type according to the vocabulary types contained in the English text and the Chinese text respectively, and performs adaptive labeling, and finally automatically classifies the target topics into the corresponding topic set according to the labeling results, so that a large number of target topics of different types can be automatically labeled and classified in a targeted and efficient manner, thereby improving the efficiency of deep processing of topics and facilitating subsequent analysis and processing of topics.
[0054] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0055] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0057] Figure 1 A flowchart of the method for automatically labeling and classifying topics provided by the present invention.
[0058] Figure 2 This is a structural diagram of the system for automatically labeling and classifying topics provided by the present invention. DETAILED DESCRIPTION
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0060] See Figure 1 , which is a flow chart of a method for automatically labeling and classifying topics provided by an embodiment of the present invention. The method for automatically labeling and classifying topics includes the following steps:
[0061] Step S1: photographing a plurality of target topics respectively to obtain image information about each target topic, and converting the image information into text data to obtain topic text data samples about the plurality of target topics;
[0062] Step S2: performing text language information recognition processing on the question text data sample to obtain English text information and Chinese text information corresponding to the question text data sample;
[0063] Step S3, performing word type recognition processing on the English text information and the Chinese text information, thereby obtaining subject type labeling information corresponding to the question text data sample;
[0064] Step S4, according to the subject type annotation information, the target topics are matched and divided into different topic sets, thereby achieving automatic classification of the target topics.
[0065] The beneficial effects of the above technical solution are: the method of automatic question labeling and classification can convert target questions with different entity forms into corresponding text forms through image shooting and image conversion, and can also effectively and accurately label and classify question text data samples through different processing processes of English / Chinese text information recognition, word type recognition and subject type labeling, thereby effectively reducing the computational workload of labeling and classifying target questions, and also improving the speed and accuracy of labeling and classifying target questions.
[0066] Preferably, in step S1, a plurality of target topics are photographed respectively to obtain image information about each target topic, and the image information is converted into text data to obtain topic text data samples about the plurality of target topics. Specifically, the following steps are performed:
[0067] Step S101, scanning and photographing each target topic to obtain a two-dimensional image of each target topic;
[0068] Step S102, performing pixel binarization processing and background noise reduction filtering processing on the two-dimensional image, thereby converting the two-dimensional image into a grayscale image;
[0069] Step S103: extract the corresponding question text character outline information from the grayscale image, and convert the grayscale image into question text data corresponding to the target question based on the question text character outline information, so that the question text data corresponding to all target questions form a question text data sample.
[0070] The beneficial effects of the above technical solution are: by performing pixel binarization and background noise reduction filtering on the image obtained by shooting the target title, the noise information in the image that is irrelevant to the target title can be effectively extracted, and the corresponding title text data can be obtained based on the character contour information of the title text contained in the image. The above image-text conversion method can be applied to images of different resolution levels, and can also effectively reduce the error rate of text recognition conversion and improve the reliability of text recognition conversion.
[0071] Preferably, in step S2, performing text language information recognition processing on the question text data sample to obtain English text information and Chinese text information corresponding to the question text data sample specifically includes:
[0072] According to the following formula (1), the text language information of the question text data sample is identified, thereby obtaining the English text information A corresponding to the question text data sample. n and Chinese text message B m ,
[0073]
[0074] In the above formula (1), Title(A n , B m ) represents the semantic approximation value of the title text composed of the semantic approximation value of the English text and the semantic approximation value of the Chinese text included in the title text data sample, π represents the pi, arctan represents the inverse tangent function operator, A n Indicates the text semantic approximation value corresponding to the nth English text in the title, B mIndicates the text semantic approximation value corresponding to the mth Chinese text, N indicates the total amount of English text data contained in the English text information, and its maximum value is 40, n is an integer between 1 and 40, M indicates the total amount of Chinese text data contained in the Chinese text information, and its maximum value is 20, m is an integer between 1 and 20, j indicates that any Chinese text character in the question text data sample is split into eight intervals according to the rice grid, and each interval is marked in counterclockwise order in the right horizontal axis direction, and the value of j can only be 1, 2, 3, 4, 5, 6, 7, 8, l i Indicates the horizontal length of the jth interval of any Chinese text character, h j Indicates the vertical length of the jth interval of any Chinese text character. Indicates the horizontal stroke space vector corresponding to any Chinese text character, represents the vertical stroke space vector corresponding to any Chinese text character, f(a) represents the character area value corresponding to any English text character in the text data sample of the question, Represents the recognition result of the English text characters of the text data sample of the question, wherein the text semantic approximation value corresponding to the nth English text can be determined in the following manner: when the nth English text is determined to be similar to the letter A or a, the corresponding text semantic approximation value is 1; when the nth English text is determined to be similar to the letter B or b, the corresponding text semantic approximation value is 2; and so on, when the nth English text is determined to be similar to the letter Z or z, the corresponding text semantic approximation value is 26; the text semantic approximation value corresponding to the mth Chinese text can be determined in the following manner: according to the input mode of the Wubi input method, all stroke types contained in the mth Chinese text are determined, and a corresponding stroke score is preset for each stroke type, and then the stroke scores corresponding to all stroke types contained in the mth Chinese text are accumulated, so that the accumulated result is used as the text semantic approximation value of the mth Chinese text;
[0075] as well as,
[0076] In step S3, word type recognition processing is performed on the English text information and the Chinese text information to obtain the subject type annotation information corresponding to the question text data sample, which specifically includes:
[0077] According to the following formula (2), the English text information and the Chinese text information are processed for word type recognition to obtain the subject type labeling information corresponding to the question text data sample.
[0078]
[0079] In the above formula (2), Match(q, d) represents the subject type label matching value corresponding to the question text data sample, Q represents the total number of subjects contained in the question text data sample, D represents the total number of English words and Chinese words contained in the question text data sample, q represents any positive integer between [1, Q], and d represents any positive integer between [1, D]. For example, when the text semantic similarity value of the nth English text in the English text information is 16, the text semantic similarity value of the n+1th English text is 8, and the text semantic similarity value of the n+2th English text is 25, then the English text information composed of the nth English text, the n+1th English text and the n+2th English text is identified as "PHY", that is, the corresponding subject type is "physics".
[0080] The beneficial effects of the above technical solution are: by synchronously distinguishing and identifying the English text information and Chinese text information in the question text data sample, the efficiency of identifying information in different languages in the question text data sample can be improved, and it can also accurately mark the questions in different languages in the question text data sample according to actual needs, so as to carry out targeted differentiation and processing of questions in different languages.
[0081] Preferably, in step S4, according to the subject type annotation information, the target topics are matched and divided into different topic sets, thereby achieving automatic classification of the target topics. Specifically, the process includes:
[0082] Step S401: According to the following formula (3) and the subject type annotation information, determine the matching value Disp(o) between each target topic and the corresponding keyword in the preset classification keyword library. r ,t r ),
[0083]
[0084] In the above formula (3), o r Indicates the number of characters corresponding to the rth keyword, t r Indicates the character bit length corresponding to the rth keyword, where r is any positive integer greater than or equal to 1;
[0085] Step S402: When the matching value Disp( r ,t r ) is equal to 1, indicating that the current target topic matches the current keyword, and the current target topic is divided into the topic set corresponding to the current keyword, thereby realizing automatic classification of the target topic.
[0086] The beneficial effects of the above technical solution are: by determining the matching value between each target topic and the preset keywords, it is possible to use the preset keywords as the classification standard for the target topics, and only when there is a complete match between the target topic and the preset keywords, the target topic will be automatically classified into the corresponding topic set, thereby improving the classification automation and classification accuracy of the target topic.
[0087] See Figure 2 , is a structural diagram of a system for automatic question annotation and classification provided by an embodiment of the present invention. The system for automatic question annotation and classification includes a target question shooting module, a question text data sample acquisition module, a question English text / Chinese text information acquisition module, a subject type annotation information acquisition module, and a question automatic classification module; wherein,
[0088] The target topic shooting module is used to shoot a plurality of target topics respectively, so as to obtain image information about each target topic;
[0089] The topic text data sample acquisition module is used to perform text data conversion processing on the image information, thereby obtaining topic text data samples related to a number of the target topics;
[0090] The question English text / Chinese text information acquisition module is used to perform text language information recognition processing on the question text data sample, so as to obtain the English text information and Chinese text information corresponding to the question text data sample;
[0091] The subject type annotation information acquisition module is used to perform word type recognition processing on the English text information and the Chinese text information, so as to obtain the subject type annotation information corresponding to the question text data sample;
[0092] The topic automatic classification module is used to match and classify a number of target topics into different topic sets according to the subject type annotation information, thereby realizing automatic classification of a number of target topics.
[0093] The beneficial effects of the above technical solution are: the system for automatic topic labeling and classification can convert target topics with different entity forms into corresponding text forms through image shooting and image conversion, and can also perform effective and accurate labeling and classification deep processing on topic text data samples through different processing processes of English / Chinese text information recognition, word type recognition and subject type labeling, thereby effectively reducing the computational workload of labeling and classifying target topics, and also improving the speed and accuracy of labeling and classifying target topics.
[0094] Preferably, the target topic shooting module shoots several target topics respectively, thereby obtaining image information about each target topic, specifically including:
[0095] Scanning and photographing each target subject to obtain a two-dimensional image of each target subject;
[0096] as well as,
[0097] The topic text data sample acquisition module performs text data conversion processing on the image information, thereby obtaining topic text data samples about several target topics, specifically including:
[0098] Performing pixel binarization processing and background noise reduction filtering processing on the two-dimensional image, thereby converting the two-dimensional image into a grayscale image;
[0099] Then, the corresponding question text character contour information is extracted from the grayscale image, and based on the question text character contour information, the grayscale image is converted into question text data corresponding to the target question, so that the question text data corresponding to all target questions form a question text data sample.
[0100] The beneficial effects of the above technical solution are: by performing pixel binarization and background noise reduction filtering on the image obtained by shooting the target title, the noise information in the image that is irrelevant to the target title can be effectively extracted, and the corresponding title text data can be obtained based on the character contour information of the title text contained in the image. The above image-text conversion method can be applied to images of different resolution levels, and can also effectively reduce the error rate of text recognition conversion and improve the reliability of text recognition conversion.
[0101] Preferably, the question English text / Chinese text information acquisition module performs text language information recognition processing on the question text data sample to obtain English text information and Chinese text information corresponding to the question text data sample, specifically including:
[0102] According to the following formula (1), the text language information of the question text data sample is identified, thereby obtaining the English text information A corresponding to the question text data sample. n and Chinese text message B m ,
[0103]
[0104] In the above formula (1), Title(A n , B m) represents the semantic approximation value of the title text composed of the semantic approximation value of the English text and the semantic approximation value of the Chinese text included in the title text data sample, π represents the pi, arctan represents the inverse tangent function operator, A n Indicates the text semantic approximation value corresponding to the nth English text in the title, B m Indicates the text semantic approximation value corresponding to the mth Chinese text, N indicates the total amount of English text data contained in the English text information, and its maximum value is 40, n is an integer between 1 and 40, M indicates the total amount of Chinese text data contained in the Chinese text information, and its maximum value is 20, m is an integer between 1 and 20, j indicates that any Chinese text character in the question text data sample is split into eight intervals according to the rice grid, and each interval is marked in counterclockwise order in the right horizontal axis direction, and the value of j can only be 1, 2, 3, 4, 5, 6, 7, 8, l i Indicates the horizontal length of the jth interval of any Chinese text character, h j Indicates the vertical length of the jth interval of any Chinese text character. Indicates the horizontal stroke space vector corresponding to any Chinese text character, represents the vertical stroke space vector corresponding to any Chinese text character, f(a) represents the character area value corresponding to any English text character in the text data sample of the question, Represents the recognition result of the English text characters of the text data sample of the question, wherein the text semantic approximation value corresponding to the nth English text can be determined in the following manner: when the nth English text is determined to be similar to the letter A or a, the corresponding text semantic approximation value is 1; when the nth English text is determined to be similar to the letter B or b, the corresponding text semantic approximation value is 2; and so on, when the nth English text is determined to be similar to the letter Z or z, the corresponding text semantic approximation value is 26; the text semantic approximation value corresponding to the mth Chinese text can be determined in the following manner: according to the input mode of the Wubi input method, all stroke types contained in the mth Chinese text are determined, and a corresponding stroke score is preset for each stroke type, and then the stroke scores corresponding to all stroke types contained in the mth Chinese text are accumulated, so that the accumulated result is used as the text semantic approximation value of the mth Chinese text;
[0105] as well as,
[0106] The subject type annotation information acquisition module performs word type recognition processing on the English text information and the Chinese text information to obtain the subject type annotation information corresponding to the question text data sample, specifically including:
[0107] According to the following formula (2), the English text information and the Chinese text information are processed for word type recognition to obtain the subject type labeling information corresponding to the question text data sample.
[0108]
[0109] In the above formula (2), Match(q, d) represents the subject type label matching value corresponding to the question text data sample, Q represents the total number of subjects contained in the question text data sample, D represents the total number of English words and Chinese words contained in the question text data sample, q represents any positive integer between [1, Q], and d represents any positive integer between [1, D]. For example, when the text semantic similarity value of the nth English text in the English text information is 16, the text semantic similarity value of the n+1th English text is 8, and the text semantic similarity value of the n+2th English text is 25, then the English text information composed of the nth English text, the n+1th English text and the n+2th English text is identified as "PHY", that is, the corresponding subject type is "physics".
[0110] The beneficial effects of the above technical solution are: by synchronously distinguishing and identifying the English text information and Chinese text information in the question text data sample, the efficiency of identifying information in different languages in the question text data sample can be improved, and it can also accurately mark the questions in different languages in the question text data sample according to actual needs, so as to carry out targeted differentiation and processing of questions in different languages.
[0111] Preferably, the automatic topic classification module matches and divides the target topics into different topic sets according to the subject type annotation information, thereby achieving automatic classification of the target topics. Specifically, the automatic classification includes:
[0112] According to the following formula (3) and the subject type annotation information, the matching value Disp(o) between each target topic and the corresponding keyword in the preset classification keyword library is determined. r ,t r ),
[0113]
[0114] In the above formula (3), o r Indicates the number of characters corresponding to the rth keyword, t r Indicates the character bit length corresponding to the rth keyword, where r is any positive integer greater than or equal to 1;
[0115] And when the matching value Disp(o r ,t r) is equal to 1, indicating that the current target topic matches the current keyword, and the current target topic is divided into the topic set corresponding to the current keyword, thereby realizing automatic classification of the target topic.
[0116] The beneficial effects of the above technical solution are: by determining the matching value between each target topic and the preset keywords, it is possible to use the preset keywords as the classification standard for the target topics, and only when there is a complete match between the target topic and the preset keywords, the target topic will be automatically classified into the corresponding topic set, thereby improving the classification automation and classification accuracy of the target topic.
[0117] From the contents of the above embodiments, it can be seen that the method and system for automatic topic labeling and classification obtains image information about each target topic by photographing several target topics respectively, and performs text data conversion processing on the image information to obtain topic text data samples about several target topics, and performs text language information recognition processing on the topic text data samples to obtain English text information and Chinese text information corresponding to the topic text data samples, and then performs word type recognition processing on the English text information and Chinese text information to obtain subject type labeling information corresponding to the topic text data samples, and finally, according to the subject type labeling information, matches and divides several target topics into different topic sets. , thereby realizing the automatic classification of several target topics; it can be seen that the method and system for automatic topic labeling and classification shoots the target topic and converts the photographed image into the corresponding topic text data, and identifies the English text and Chinese text respectively contained in the topic text data samples, and then obtains the corresponding subject type according to the vocabulary types contained in the English text and the Chinese text, and performs adaptive labeling, and finally automatically classifies the target topic into the corresponding topic set according to the labeling results. In this way, a large number of target topics of different types can be automatically labeled and classified in a targeted and efficient manner, thereby improving the efficiency of deep processing of the topics and facilitating subsequent analysis and processing of the topics.
[0118] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. The method for automatically labeling and classifying questions is characterized by: It includes the following steps: Step S1: photographing a plurality of target topics respectively to obtain image information about each of the target topics, and performing text data conversion processing on the image information to obtain topic text data samples about the plurality of target topics; Step S2, performing text language information recognition processing on the question text data sample, thereby obtaining English text information and Chinese text information corresponding to the question text data sample; Step S3, performing word type recognition processing on the English text information and the Chinese text information, thereby obtaining subject type labeling information corresponding to the question text data sample; Step S4, matching and classifying the target topics into different topic sets according to the subject type annotation information, thereby achieving automatic classification of the target topics; Wherein, in the step S2, performing text language information recognition processing on the question text data sample to obtain English text information and Chinese text information corresponding to the question text data sample specifically includes: According to the following formula (1), the text language information recognition process is performed on the question text data sample, thereby obtaining the English text semantic approximation value and the Chinese text semantic approximation value included in the question text data sample: In the above formula (1), Title(A n , B m ) represents the semantic approximation value of the title text composed of the semantic approximation value of the English text and the semantic approximation value of the Chinese text included in the title text data sample, π represents the pi, arctan represents the inverse tangent function operator, A n Indicates the text semantic approximation value corresponding to the nth English text in the title, B m represents the text semantic approximation value corresponding to the mth Chinese text, N represents the total number of English text data contained in the English text information, and its maximum value is 40, n is an integer between 1 and 40, M represents the total number of Chinese text data contained in the Chinese text information, and its maximum value is 20, m is an integer between 1 and 20, j represents any Chinese text character in the question text data sample, which is split into eight intervals according to the rice grid, and each interval is marked in a counterclockwise order in the right horizontal axis direction, and the value of j can only be 1, 2, 3, 4, 5, 6, 7, 8, l j represents the horizontal length of the jth interval of any Chinese text character, h j represents the vertical length corresponding to the j-th interval of any Chinese text character, Indicates the horizontal stroke space vector corresponding to any one of the Chinese text characters, represents the vertical stroke space vector corresponding to any one of the Chinese text characters; f(a) represents the character area value corresponding to any one of the English text characters in the question text data sample, represents the recognition result of the English text characters of the subject text data sample, Represents the recognition result of the Chinese text characters of the subject text data sample; as well as, In step S3, word type recognition processing is performed on the English text information and the Chinese text information to obtain subject type annotation information corresponding to the question text data sample, specifically including: According to the following formula (2), the English text information and the Chinese text information are processed for word type recognition to obtain the subject type labeling information corresponding to the question text data sample: In the above formula (2), Match(q, d) represents the subject type label matching value corresponding to the question text data sample, Q represents the total number of subjects contained in the question text data sample, D represents the total number of English words and Chinese words contained in the question text data sample, q represents any positive integer between [1, Q], and d represents any positive integer between [1, D].
2. The method for automatically labeling and classifying questions according to claim 1, wherein: In step S1, a plurality of target topics are photographed respectively to obtain image information about each target topic, and the image information is converted into text data to obtain topic text data samples about the plurality of target topics. Specifically, the following steps are performed: Step S101, scanning and photographing each of the target topics to obtain a two-dimensional image of each of the target topics; Step S102, performing pixel binarization processing and background noise reduction filtering processing on the two-dimensional image, thereby converting the two-dimensional image into a grayscale image; Step S103: extract the corresponding question text character contour information from the grayscale image, and convert the grayscale image into question text data corresponding to the target question based on the question text character contour information, so that the question text data corresponding to all target questions form a question text data sample.
3. The method and system for automatically labeling and classifying questions according to claim 1, characterized in that: In step S4, according to the subject type annotation information, the target topics are matched and divided into different topic sets, thereby achieving automatic classification of the target topics. Specifically, the following steps are performed: Step S401: According to the following formula (3) and the subject type annotation information, the matching value Disp(o) between each target topic and the corresponding keyword in the preset classification keyword library is determined. r ,t r ), In the above formula (3), o r Indicates the number of characters corresponding to the rth keyword, t r Indicates the character bit length corresponding to the rth keyword, where r is any positive integer greater than or equal to 1; Step S402: When the matching value Disp( r ,t r ) is equal to 1, indicating that the current target topic matches the current keyword, and the current target topic is divided into the topic set corresponding to the current keyword, thereby realizing automatic classification of the target topic.
4. The system for automatic question labeling and classification is characterized by: It includes a target topic shooting module, a topic text data sample acquisition module, a topic English text / Chinese text information acquisition module, a subject type annotation information acquisition module and a topic automatic classification module; among which, The target topic shooting module is used to shoot a plurality of target topics respectively, so as to obtain image information about each of the target topics; The topic text data sample acquisition module is used to perform text data conversion processing on the image information, thereby obtaining topic text data samples about a number of the target topics; The question English text / Chinese text information acquisition module is used to perform text language information recognition processing on the question text data sample, thereby obtaining English text information and Chinese text information corresponding to the question text data sample; The subject type annotation information acquisition module is used to perform word type recognition processing on the English text information and the Chinese text information, so as to obtain the subject type annotation information corresponding to the question text data sample; The automatic topic classification module is used to match and classify the target topics into different topic sets according to the subject type annotation information, thereby realizing automatic classification of the target topics; The English / Chinese text information acquisition module performs text language information recognition processing on the question text data sample to obtain the English text information and Chinese text information corresponding to the question text data sample, specifically including: According to the following formula (1), the text language information recognition processing is performed on the subject text data sample to obtain the English text information A corresponding to the subject text data sample. n and Chinese text message B m , In the above formula (1), Title(A n , B m ) represents the semantic approximation value of the title text composed of the semantic approximation value of the English text and the semantic approximation value of the Chinese text included in the title text data sample, π represents the pi, arctan represents the inverse tangent function operator, A n Indicates the text semantic approximation value corresponding to the nth English text in the title, B m represents the text semantic approximation value corresponding to the mth Chinese text, N represents the total number of English text data contained in the English text information, and its maximum value is 40, n is an integer between 1 and 40, M represents the total number of Chinese text data contained in the Chinese text information, and its maximum value is 20, m is an integer between 1 and 20, j represents any Chinese text character in the question text data sample, which is split into eight intervals according to the rice grid, and each interval is marked in a counterclockwise order in the right horizontal axis direction, and the value of j can only be 1, 2, 3, 4, 5, 6, 7, 8, l j represents the horizontal length of the jth interval of any Chinese text character, h j represents the vertical length corresponding to the j-th interval of any Chinese text character, Indicates the horizontal stroke space vector corresponding to any one of the Chinese text characters, represents the vertical stroke space vector corresponding to any one of the Chinese characters, f(a) represents the character area value corresponding to any one of the English characters in the question text data sample, Represents the recognition result of the English text characters of the subject text data sample; is the symbol of partial derivative function; as well as, The subject type annotation information acquisition module performs word type recognition processing on the English text information and the Chinese text information to obtain the subject type annotation information corresponding to the question text data sample, specifically including: According to the following formula (2), the English text information and the Chinese text information are processed for word type recognition to obtain the subject type labeling information corresponding to the question text data sample. In the above formula (2), Match(q, d) represents the subject type label matching value corresponding to the question text data sample, Q represents the total number of subjects contained in the question text data sample, D represents the total number of English words and Chinese words contained in the question text data sample, q represents any positive integer between [1, Q], and d represents any positive integer between [1, D].
5. The system for automatically labeling and classifying questions according to claim 4, characterized in that: The target topic shooting module shoots a plurality of target topics respectively, thereby obtaining image information about each target topic, specifically including: Scanning and photographing each of the target topics to obtain a two-dimensional image of each of the target topics; as well as, The question text data sample acquisition module performs text data conversion processing on the image information to obtain question text data samples about the target questions, specifically including: performing pixel binarization processing and background noise reduction filtering processing on the two-dimensional image to convert the two-dimensional image into a grayscale image; Then, the corresponding question text character contour information is extracted from the grayscale image, and based on the question text character contour information, the grayscale image is converted into question text data corresponding to the target question, so that the question text data corresponding to all target questions form a question text data sample.
6. The system for automatically labeling and classifying questions according to claim 4, characterized in that: The automatic topic classification module matches and divides the target topics into different topic sets according to the subject type annotation information, thereby achieving automatic classification of the target topics. Specifically, the module includes: According to the following formula (3) and the subject type annotation information, the matching value Disp(o) between each target topic and the corresponding keyword in the preset classification keyword library is determined. r ,t r ), In the above formula (3), o r Indicates the number of characters corresponding to the rth keyword, t r Indicates the character bit length corresponding to the rth keyword, where r is any positive integer greater than or equal to 1; And when the matching value Disp(o r ,t r ) is equal to 1, indicating that the current target topic matches the current keyword, and the current target topic is divided into the topic set corresponding to the current keyword, thereby realizing automatic classification of the target topic.
Citation Information
Patent Citations
Topic classification method and system
CN106776724A
Multidisciplinary test paper content detection and recognition system and method based on deep learning
CN110210413A