Fishery literature subject classification method and system and electronic equipment

By constructing narrative mask images and subject relationship clustering images and superimposing them, the classification types of fishery literature are accurately determined, and the problem of inconsistency in classification errors and standards in the prior art is solved, and a high-accurate literature classification is achieved.

CN120144771APending Publication Date: 2025-06-13CHINESE ACAD OF FISHERY SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510276036.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When processing fishery literature, it is difficult for the prior art to effectively capture semantic correlation and identify unique attributes of the literature, resulting in low classification error and accuracy, increasing the need for human classification, and there are inconsistencies in standards caused by subjective evaluation.

Method used

A discipline classification method for fishery literature is proposed. By constructing narrative mask images and discipline relationship cluster images, and superimposing the two to obtain residual images, thereby accurately determining the classification type of target classification literature.

Benefits of technology

Deeply capture and understand the semantic correlations of sentences in the target classification literature, effectively distinguish and identify the unique attributes of different literatures, reduce classification errors, improve classification accuracy, reduce human classification needs, and solve the problem of inconsistent standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144771A_ABST
    Figure CN120144771A_ABST
Patent Text Reader

Abstract

The invention discloses a fishery literature subject classification method, a fishery literature subject classification system and electronic equipment. The fishery literature subject classification method comprises the following steps: obtaining thesaurus information of fishery literature and target classification literature; constructing a thesaurus mask image according to thesaurus information; constructing a subject relationship clustering image according to the target classification literature; overlapping the descriptor mask image positive film to the subject relationship clustering image to obtain a residual image; and determining a classification type corresponding to the target classification literature according to the residual image. The method can deeply capture and understand semantic association of statements in target classification literatures, effectively distinguish and identify unique attributes of different literatures, reduce classification errors caused by approximate topics, improve classification accuracy, reduce manpower for classifying the literatures in the fishery professional field, and improve classification efficiency. The problem of standard inconsistency caused by subjective evaluation in a human participation process when the documents in the fishery professional field are classified is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image technology, and more particularly to a method, system, and electronic device for subject classification of fishery literature. Background Art

[0002] In the current fishery system, fishery literature has a wide coverage, involves multiple aspects, is complex and intersecting, and there are close connections between fishery science and other disciplines such as biology, environmental science, economics, etc., which increases the difficulty of literature classification.

[0003] To solve the above problems, in the prior art, clustering methods are usually used to classify literature in the fishery professional field, such as using clustering algorithms such as k - clustering, k - means, and k - center. However, such algorithms have limitations in processing high - dimensional data. For example, problems such as semantic association problems in fishery literature and the large number of professional terms and strong theme intersection in fishery literature will lead to poor traditional clustering effects, resulting in classification errors for approximate themes in the automatic classification process, and further leading to a large amount of manpower spent on manually classifying fishery professional field literature, and increasing the problem of inconsistent standards caused by subjective evaluation in the process of manual participation. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art.

[0005] For this reason, an object of the present invention is to provide a method for subject classification of fishery literature, which can deeply capture and understand the semantic associations of sentences in the target - classified literature, effectively distinguish and identify the unique attributes of different literatures, reduce classification errors caused by approximate themes, improve the accuracy of classification, thereby reducing the manpower for classifying fishery professional field literature and solving the problem of inconsistent standards caused by subjective evaluation in the process of manually classifying fishery professional field literature.

[0006] For this reason, a second object of the present invention is to provide a system for subject classification of fishery literature.

[0007] For this reason, a third object of the present invention is to provide an electronic device.

[0008] For this reason, a fourth object of the present invention is to provide a computer - readable storage medium.

[0009] To achieve the above object, an embodiment of the first aspect of the present invention provides a method for subject classification of fishery literature, the method comprising the following steps: obtaining the thesaurus information of fishery literature and the target classification literature; constructing a thesaurus mask image according to the thesaurus information; constructing a subject relationship clustering image according to the target classification literature; overlaying the thesaurus mask image on the subject relationship clustering image to obtain a residual image; and determining the classification type corresponding to the target classification literature according to the residual image.

[0010] According to the method for subject classification of fishery literature of the embodiment of the present invention, by constructing a thesaurus mask image and a subject relationship clustering image and overlaying the two to obtain a residual image, it is convenient to accurately determine the classification type of the target classification literature according to the residual image, can deeply capture and understand the semantic association of the sentences in the target classification literature, effectively distinguish and identify the unique attributes of different literatures, reduce the classification error caused by approximate topics, improve the accuracy of classification, and thus can reduce the manpower for classifying the literature in the fishery professional field and solve the problem of inconsistent standards caused by subjective evaluation in the process of human participation in classifying the literature in the fishery professional field.

[0011] In addition, the method for subject classification of fishery literature according to the embodiment of the present invention may further have the following additional technical features: In some examples, constructing a thesaurus mask image according to the thesaurus information includes: annotating the subject relationships of the words in the thesaurus information; obtaining the part-of-speech approximation relationships of the annotated words; clustering the annotated words according to the part-of-speech approximation relationships; and constructing the thesaurus mask image according to the clustered words.

[0012] In some examples, constructing the thesaurus mask image according to the clustered words includes: obtaining the normalized value of the distance between the clustered words and the clustering center points to which their corresponding subjects belong; drawing the corresponding clustering images based on the clustered words and using the normalized value as the transparency index of the clustering images; and keeping the transparency index unchanged and converting the clustering images into grayscale images to obtain the thesaurus mask images.

[0013] In some examples, constructing a subject relationship clustering image according to the target classification literature includes: preprocessing the target classification literature, where the preprocessing includes word segmentation and stop word removal; and constructing a subject relationship clustering image according to the preprocessed target classification literature.

[0014] In some examples, constructing a disciplinary relationship clustering image based on the preprocessed target classification literature includes: extracting subject terms from the preprocessed target classification literature; obtaining the disciplinary themes corresponding to the extracted subject terms and the approximation between the extracted subject terms and the disciplinary themes; and constructing a disciplinary relationship clustering image according to the approximation.

[0015] In some examples, multiplying the gray space of the thesaurus mask image with the color space of the disciplinary relationship clustering image and multiplying the inverted value of the transparency space of the thesaurus mask image with all channel values of the disciplinary relationship clustering image to obtain the residual image includes: multiplying the thesaurus mask image onto the disciplinary relationship clustering image to obtain a residual image.

[0016] In some examples, determining the classification type corresponding to the target classification literature according to the residual image includes: determining the corresponding closed region image according to the residual image; determining the maximum area clustering corresponding to the residual image according to the closed region image; and determining the classification type corresponding to the target classification literature according to the maximum area clustering corresponding to the residual image.

[0017] In some examples, determining the classification type corresponding to the target classification literature according to the maximum area clustering corresponding to the residual image includes: obtaining the clustering center of the maximum area clustering corresponding to the residual image; and taking the classification type corresponding to the clustering center in the thesaurus mask image as the classification type corresponding to the target classification literature.

[0018] To achieve the above object, an embodiment of the second aspect of the present invention provides a disciplinary classification system for fishery literature, the disciplinary classification system for fishery literature including: an acquisition module, configured to acquire the thesaurus information of fishery literature and target classification literature; a first construction module, configured to construct a thesaurus mask image according to the thesaurus information; a second construction module, configured to construct a disciplinary relationship clustering image according to the target classification literature; a third construction module, configured to multiply the thesaurus mask image onto the disciplinary relationship clustering image to obtain a residual image; and a determination module, configured to determine the classification type corresponding to the target classification literature according to the residual image.

[0019] According to the subject classification system of fishery literature of the present invention, by constructing a thesaurus mask image and a subject relationship clustering image, and superimposing the two to obtain a residual image, it is convenient to accurately determine the classification type of the target classified literature according to the residual image, can deeply capture and understand the semantic associations of the statements in the target classified literature, effectively distinguish and identify the unique attributes of different literatures, reduce the classification errors caused by approximate topics, improve the accuracy of classification, and thus can reduce the manpower for classifying the literature in the fishery professional field, and solve the problem of inconsistent standards caused by subjective evaluation in the process of human participation when classifying the literature in the fishery professional field.

[0020] To achieve the above object, the third aspect embodiment of the present invention discloses an electronic device, which includes: the subject classification system of fishery literature described in the second aspect embodiment of the present invention; or, a processor, a memory, and a subject classification program of fishery literature stored on the memory and executable on the processor, and when the subject classification program of fishery literature is executed by the processor, it implements the subject classification method of fishery literature described in the first aspect embodiment of the present invention.

[0021] According to the electronic device of the embodiment of the present invention, by constructing a thesaurus mask image and a subject relationship clustering image, and superimposing the two to obtain a residual image, it is convenient to accurately determine the classification type of the target classified literature according to the residual image, can deeply capture and understand the semantic associations of the statements in the target classified literature, effectively distinguish and identify the unique attributes of different literatures, reduce the classification errors caused by approximate topics, improve the accuracy of classification, and thus can reduce the manpower for classifying the literature in the fishery professional field, and solve the problem of inconsistent standards caused by subjective evaluation in the process of human participation when classifying the literature in the fishery professional field.

[0022] To achieve the above object, the fourth aspect embodiment of the present invention discloses a computer-readable storage medium, on which a subject classification program of fishery literature is stored, and when the subject classification program of fishery literature is executed by a processor, it implements the subject classification method of fishery literature described in the first aspect embodiment of the present invention.

[0023] When the subject classification program of fishery literature stored on a computer-readable storage medium according to an embodiment of the present invention is executed by a processor, by constructing a thesaurus mask image and a subject relationship clustering image and superimposing the two to obtain a residual image, it is convenient to accurately determine the classification type of the target classification literature, deeply capture and understand the semantic associations of the sentences in the target classification literature, effectively distinguish and identify the unique attributes of different literatures, reduce classification errors caused by approximate topics, improve the accuracy of classification, and thus can reduce the manpower for classifying fishery professional literature and solve the problem of inconsistent standards caused by subjective evaluation in the process of human participation in classifying fishery professional literature.

[0024] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 is a schematic flowchart of a subject classification method for fishery literature according to an embodiment of the present invention; Figure 2 is a schematic structural diagram of a subject classification system for fishery literature according to an embodiment of the present invention.

[0026] Reference numerals: Subject classification system for fishery literature - 100; Acquisition module - 110; First construction module - 120; Second construction module - 130; Third construction module - 140; Determination module - 150. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] In order to be able to understand the features and technical content of the embodiments of the present invention in more detail, the implementation of the embodiments of the present invention will be described in detail below with reference to the drawings. The attached drawings are for reference and illustration only and are not used to limit the embodiments of the present invention. In the following technical description, for the sake of explanation, numerous details are provided to give a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner.

[0028] The following refers to Figure 1 - Figure 2 Describe a subject classification method and device for fishery literature according to an embodiment of the present invention.

[0029] Figure 1 is a schematic flowchart of a subject classification method for fishery literature according to an embodiment of the present invention, in combination with Figure 1As shown in the figure, the subject classification method of this fishery literature includes the following steps: Step S1: Obtain the thesaurus information of fishery literature and the target classification literature.

[0030] Specifically, in the process of subject classification of fishery literature, the thesaurus information of fishery literature and the target classification literature can be obtained. Among them, the thesaurus information is also called the subject thesaurus information, which contains terms and concepts in the fishery field, including synonyms, near-synonyms, antonyms, and relationships between terms. For example, it includes a large number of subject terms in categories such as fishery resources, fishery technology, fishery economy, and fishery management; the target classification literature is the specific literature that needs to be classified, which can come from various channels, including but not limited to academic journal papers, conference papers, research reports, government publications, etc.

[0031] Step S2: Construct a thesaurus mask image according to the thesaurus information.

[0032] Specifically, in the process of subject classification of fishery literature, the thesaurus information can be converted into an image form, that is, each vocabulary in the thesaurus information is represented as a specific area or feature in an image. For example, each vocabulary in the thesaurus can be converted into a vector form in a high-dimensional space, and the vector is projected into a two-dimensional space, and the position of each vocabulary is marked in a two-dimensional coordinate system. At the same time, different-sized, -colored, or -shaped points can be selected to represent different features represented by each vocabulary, so as to construct a thesaurus mask image.

[0033] Step S3: Construct a subject relationship clustering image according to the target classification literature.

[0034] Specifically, in the process of subject classification of fishery literature, a subject relationship clustering image can be constructed according to the target classification literature. For example, the target classification literature can be preprocessed to extract the keywords of the target classification literature, so that a subject relationship network can be constructed according to the co-occurrence relationship, citation relationship, or semantic similarity between the keywords, so as to select a suitable clustering algorithm for clustering according to the characteristics and data scale of the subject relationship network, and finally present the clustering result in the form of an image, so as to obtain the subject relationship clustering image. It can be understood that when presenting the clustering result in the form of an image, the clustering image can be constructed in the same two-dimensional space as the thesaurus mask image, that is, the X / Y coordinate units of the two images are the same.

[0035] Step S4: Overlay the thesaurus mask image on the subject relationship clustering image in a multiply mode to obtain a residual image.

[0036] Specifically, after obtaining the thesaurus mask image and the subject relationship clustering image, an image synthesis technique can be used to overlay the two, that is, to multiply the thesaurus mask image onto the subject relationship clustering image to obtain a residual image. For example, the thesaurus mask image and the subject relationship clustering image can be imported into an image production software respectively as two independent layers, and in the layer property panel, the mode can be set to multiply. In this mode, the two images will be overlaid to obtain a residual image.

[0037] Step S5: Determine the classification type corresponding to the target classified document according to the residual image.

[0038] Specifically, after obtaining the residual image, the classification type corresponding to the target classified document can be determined according to the residual image. For example, image segmentation technology can be used to identify different regions in the residual image, and calculate the area sizes corresponding to different regions, and sort according to the area sizes corresponding to different regions, so as to determine the classification type corresponding to the target classified document according to the sorted area sizes corresponding to different regions.

[0039] Thus, the above-mentioned subject classification method for fishery documents, by constructing a thesaurus mask image and a subject relationship clustering image, and overlaying the two to obtain a residual image, facilitates accurately determining the classification type of the target classified document according to the residual image, can deeply capture and understand the semantic associations of the sentences in the target classified document, effectively distinguish and identify the unique attributes of different documents, reduce classification errors caused by approximate topics, improve the accuracy of classification, and thus can reduce the manpower for classifying documents in the fishery professional field and solve the problem of inconsistent standards caused by subjective evaluation in the process of human participation when classifying documents in the fishery professional field.

[0040] In an embodiment of the present invention, constructing a thesaurus mask image according to thesaurus information includes: annotating the subject relationships of the vocabulary in the thesaurus information; obtaining the part-of-speech approximation relationships of the annotated vocabulary; clustering the annotated vocabulary according to the part-of-speech approximation relationships; and constructing a thesaurus mask image according to the clustered vocabulary.

[0041] Specifically, in the process of constructing a thesaurus mask image according to thesaurus information, first, each vocabulary in the thesaurus can be annotated with subject relationships through manual annotation or using an automatic annotation tool, that is, marking the specific subject field or topic category to which each vocabulary belongs. For example, terms containing "cultivation" and "enhancement" can be annotated and classified into aquaculture.

[0042] Furthermore, the approximate relationship of the part-of-speech of the labeled words can be obtained, including but not limited to using natural language processing techniques to calculate the approximate relationship between different parts of speech. For example, the vector representation between words can be calculated through a pre-trained language model, and the part-of-speech similarity of the labeled words can be evaluated based on the distance metric of vectors (such as cosine similarity) to obtain its approximate relationship.

[0043] Furthermore, different clustering algorithms such as K-means, hierarchical clustering, etc. can be adopted to cluster the labeled words according to the approximate relationship of the part-of-speech, that is, taking the labeled words and their approximate relationship of the part-of-speech as input, performing the clustering algorithm, and dividing the words into different clusters.

[0044] Furthermore, a thesaurus mask image is constructed according to the clustered words, that is, a clustering image is drawn according to the clustered words in a two-dimensional space, and the clustering image is converted into a thesaurus mask image.

[0045] In an embodiment of the present invention, constructing a thesaurus mask image according to the clustered words includes: obtaining a normalized value of the distance between the clustered words and the clustering center point to which their corresponding disciplines belong; Drawing the corresponding clustering image based on the clustered words, and using the normalized value as the transparency index of the clustering image; Keeping the transparency index unchanged, converting the clustering image into a grayscale image to obtain a thesaurus mask image.

[0046] Specifically, in the process of constructing a thesaurus mask image according to the clustered words, after clustering the labeled words according to the approximate relationship of the part-of-speech, each cluster will have a center point (also called the centroid). At this time, a normalized value of the distance between the clustered words and the clustering center point to which their corresponding disciplines belong can be obtained, that is, normalizing the calculated distance value to ensure that it is on a unified scale. For example, all distance values can be scaled to between 0 and 1, or relatively scaled according to a certain specific maximum distance value.

[0047] Furthermore, a corresponding clustering image can be drawn according to the number of clusters and the distribution of words, and the obtained normalized distance value can be used as the transparency index of the clustering image. For example, the closer the normalized distance value is to 0, the closer the word is to the clustering center point, and the transparency can be set to be higher (that is, more opaque); the closer the normalized distance value is to 1, the farther the word is from the clustering center point, and the transparency can be set to be lower (that is, more transparent).

[0048] Furthermore, on the premise of keeping the transparency index unchanged, the colored clustering image can be converted into a grayscale image, that is, removing the color information used to distinguish different clusters and replacing it with gray levels to obtain a thesaurus mask image.

[0049] In an embodiment of the present invention, constructing a disciplinary relationship clustering image according to target classified documents includes: preprocessing the target classified documents, and the preprocessing includes word segmentation and stop word removal; Constructing a disciplinary relationship clustering image according to the preprocessed target classified documents.

[0050] Specifically, in the process of constructing a disciplinary relationship clustering image according to target classified documents, first, existing Chinese word segmentation tools (such as Jieba word segmentation, etc.) can be used to perform word segmentation on the target classified documents. Then, a general stop word list can be used, or a custom stop word list can be constructed according to the domain characteristics of the target classified documents to remove the corresponding stop words in the text after word segmentation according to the stop word list, so as to complete the preprocessing of the target classified documents.

[0051] Furthermore, after preprocessing the target classified documents, a disciplinary relationship clustering image can be constructed according to the preprocessed target classified documents. For example, a disciplinary relationship network can be constructed according to the preprocessed target classified documents, and a suitable clustering algorithm can be selected for clustering according to the characteristics and data scale of the disciplinary relationship network. Finally, the clustering result is presented in the form of an image to obtain the disciplinary relationship clustering image.

[0052] In an embodiment of the present invention, constructing a disciplinary relationship clustering image according to the preprocessed target classified documents includes: extracting subject words from the preprocessed target classified documents; Obtaining the disciplinary subjects corresponding to the extracted subject words and the approximation degree between the extracted subject words and the disciplinary subjects; Constructing a disciplinary relationship clustering image according to the approximation degree.

[0053] Specifically, after preprocessing the target documents, algorithms such as K-clustering and support vector machines can be used to extract subject words from the preprocessed target classified documents, including but not limited to obtaining high-frequency words in the preprocessed target classified documents and screening out the keywords that can best represent the document theme from the high-frequency words.

[0054] Furthermore, the disciplinary subjects corresponding to the extracted subject words can be obtained, that is, mapping the extracted subject words to specific disciplinary subjects, including but not limited to matching the extracted subject words with the words in the disciplinary subject library to find the most suitable disciplinary subject. At the same time, the subject words and disciplinary subjects can both be converted into vector forms, and the cosine similarity between the two can be calculated.

[0055] Furthermore, a disciplinary relationship clustering image can be constructed according to the approximation degree, including but not limited to using different clustering algorithms, such as K-means, hierarchical clustering, etc., clustering according to the size relationship of the approximation degree, and constructing a disciplinary relationship clustering image according to the subject words after clustering.

[0056] In one embodiment of the present invention, the thesaurus mask image is multiplied with the subject relationship clustering image to obtain a residual image, including: multiplying the gray scale space of the thesaurus mask image with the color space of the subject relationship clustering image, and multiplying the inverted value of the transparency space of the thesaurus mask image with all channel values of the subject relationship clustering image to obtain a residual image.

[0057] Specifically, in the process of multiplying the thesaurus mask image with the subject relationship clustering image, the gray scale value of the thesaurus mask image can be mapped to the color space of the subject relationship clustering image, and the two are superimposed, that is, the gray scale space of the thesaurus mask image is multiplied with the color space of the subject relationship clustering image. For example, the color channel value corresponding to each pixel of the subject relationship clustering image is multiplied with the gray scale value at the corresponding position of the thesaurus mask image respectively; at the same time, in order to perform the multiply operation, the inverted value of the transparency space of the thesaurus mask image is multiplied with all channel values of the subject relationship clustering image respectively to obtain the final residual image.

[0058] In one embodiment of the present invention, determining the classification type corresponding to the target classification document according to the residual image includes: determining the corresponding closed region image according to the residual image; Determining the maximum area clustering corresponding to the residual image according to the closed region image; Determining the classification type corresponding to the target classification document according to the maximum area clustering corresponding to the residual image.

[0059] Specifically, after the thesaurus mask image and the subject relationship clustering image are superimposed, a residual image can be formed, where the residual image includes multiple closed region images, that is, multiple independent and visually distinguishable clustering images.

[0060] Further, the backpack algorithm can be used to calculate the area of the closed region image and sort the calculation results to obtain the closed region image with the largest area, that is, the maximum area clustering corresponding to the residual image.

[0061] Further, determining the classification type corresponding to the target classification document according to the maximum area clustering corresponding to the residual image includes, but is not limited to, corresponding the maximum area clustering with the thesaurus mask image to determine the classification type corresponding to the target classification document.

[0062] In one embodiment of the present invention, determining the classification type corresponding to the target classification document according to the maximum area clustering corresponding to the residual image includes: obtaining the clustering center of the maximum area clustering corresponding to the residual image; Taking the classification type corresponding to the clustering center in the thesaurus mask image as the classification type corresponding to the target classification document.

[0063] Specifically, after determining the maximum area cluster corresponding to the residual image, the cluster center corresponding to the maximum area cluster can be obtained, including but not limited to an average or representative point of all pixel positions within the cluster area.

[0064] Furthermore, the cluster center point obtained from the residual image can be mapped back to the original thesaurus mask image to obtain the classification type corresponding to the target classified document. It can be understood that since the thesaurus mask image is constructed based on the thesaurus, which contains keywords related to a specific subject area and their distribution, the position of the cluster center point on the thesaurus mask image can reflect the theme or classification to which the cluster belongs.

[0065] In summary, according to the subject classification method of fishery literature in the embodiments of the present invention, by constructing a thesaurus mask image and a subject relationship cluster image, and overlaying the two to obtain a residual image, it is convenient to accurately determine the classification type of the target classified document according to the residual image, deeply capture and understand the semantic associations of the sentences in the target classified document, effectively distinguish and identify the unique attributes of different documents, reduce classification errors caused by approximate themes, improve the accuracy of classification, and thus can reduce the manpower for classifying fishery professional literature and solve the problem of inconsistent standards caused by subjective evaluation in the process of manual participation when classifying fishery professional literature.

[0066] A further embodiment of the present invention proposes a subject classification system 100 for fishery literature, as Figure 2 shown. The subject classification system 100 for fishery literature includes: an acquisition module 110, a first construction module 120, a second construction module 130, a third construction module 140, and a determination module 150, where The acquisition module 110 is used to acquire the thesaurus information of fishery literature and the target classified document.

[0067] The first construction module 120 is used to construct a thesaurus mask image according to the thesaurus information.

[0068] The second construction module 130 is used to construct a subject relationship cluster image according to the target classified document.

[0069] The third construction module 140 is used to multiply the thesaurus mask image with the subject relationship cluster image in a multiply blend mode to obtain a residual image.

[0070] The determination module 150 is used to determine the classification type corresponding to the target classified document according to the residual image.

[0071] In some embodiments, when constructing a thesaurus mask image according to thesaurus information, the first construction module 120 is specifically configured to: label the discipline relationships of the words in the thesaurus information; obtain the part-of-speech approximation relationships of the labeled words; cluster the labeled words according to the part-of-speech approximation relationships; and construct a thesaurus mask image based on the clustered words.

[0072] In some embodiments, when constructing a thesaurus mask image based on the clustered words, the first construction module 120 is specifically configured to: obtain the normalized value of the distance between the clustered words and the cluster center points of their corresponding disciplines; draw the corresponding cluster images based on the clustered words, and use the normalized value as the transparency index of the cluster images; keep the transparency index unchanged and convert the cluster images into grayscale images to obtain the thesaurus mask image.

[0073] In some embodiments, when constructing a discipline relationship cluster image according to the target classified documents, the second construction module 130 is specifically configured to: preprocess the target classified documents, and the preprocessing includes word segmentation and stop word removal; and construct a discipline relationship cluster image according to the preprocessed target classified documents.

[0074] In some embodiments, when constructing a discipline relationship cluster image according to the preprocessed target classified documents, the second construction module 130 is specifically configured to: extract the subject words from the preprocessed target classified documents; obtain the discipline themes corresponding to the extracted subject words and the approximation degrees between the extracted subject words and the discipline themes; and construct a discipline relationship cluster image according to the approximation degrees.

[0075] In some embodiments, when multiplying the thesaurus mask image with the discipline relationship cluster image to obtain a residual image, the third construction module 140 is specifically configured to: perform multiplication on the grayscale space of the thesaurus mask image and the color space of the discipline relationship cluster image, and perform multiplication on the inverted value of the transparency space of the thesaurus mask image and all channel values of the discipline relationship cluster image to obtain a residual image.

[0076] In some embodiments, when determining the classification type corresponding to the target classified documents according to the residual image, the determination module 150 is specifically configured to: determine the corresponding closed region image according to the residual image; determine the maximum area cluster corresponding to the residual image according to the closed region image; and determine the classification type corresponding to the target classified documents according to the maximum area cluster corresponding to the residual image.

[0077] In some embodiments, when determining the classification type corresponding to the target classified documents according to the maximum area cluster corresponding to the residual image, the determination module 150 is specifically configured to: obtain the cluster center of the maximum area cluster corresponding to the residual image; Use the classification type corresponding to the cluster center in the thesaurus mask image as the classification type corresponding to the target classified documents.

[0078] According to the subject classification system 100 of fishery literature of the present invention, by constructing a thesaurus mask image and a subject relationship clustering image, and superimposing the two to obtain a residual image, it is convenient to accurately determine the classification type of the target classified literature according to the residual image, can deeply capture and understand the semantic associations of the sentences in the target classified literature, effectively distinguish and identify the unique attributes of different literatures, reduce the classification errors caused by approximate themes, improve the accuracy of classification, and thus can reduce the manpower for classifying the literature in the fishery professional field, and solve the problem of inconsistent standards caused by subjective evaluation in the process of human participation when classifying the literature in the fishery professional field.

[0079] To achieve the above object, the third aspect embodiment of the present invention discloses an electronic device, which includes: the subject classification system of fishery literature in the second aspect embodiment of the present invention; or, a processor, a memory, and a subject classification program of fishery literature stored on the memory and executable on the processor, and when the subject classification program of fishery literature is executed by the processor, it implements the subject classification method of fishery literature as in the first aspect embodiment of the present invention.

[0080] According to the electronic device of the embodiment of the present invention, by constructing a thesaurus mask image and a subject relationship clustering image, and superimposing the two to obtain a residual image, it is convenient to accurately determine the classification type of the target classified literature according to the residual image, can deeply capture and understand the semantic associations of the sentences in the target classified literature, effectively distinguish and identify the unique attributes of different literatures, reduce the classification errors caused by approximate themes, improve the accuracy of classification, and thus can reduce the manpower for classifying the literature in the fishery professional field, and solve the problem of inconsistent standards caused by subjective evaluation in the process of human participation when classifying the literature in the fishery professional field.

[0081] To achieve the above object, the fourth aspect embodiment of the present invention discloses a computer-readable storage medium, on which a subject classification program of fishery literature is stored, and when the subject classification program of fishery literature is executed by a processor, it implements the subject classification method of fishery literature as in the first aspect embodiment of the present invention.

[0082] When the subject classification program of fishery literature stored on a computer-readable storage medium according to an embodiment of the present invention is executed by a processor, by constructing a thesaurus mask image and a subject relationship clustering image and superimposing the two to obtain a residual image, it is convenient to accurately determine the classification type of the target classification literature according to the residual image, deeply capture and understand the semantic associations of the sentences in the target classification literature, effectively distinguish and identify the unique attributes of different literatures, reduce classification errors caused by approximate topics, improve the accuracy of classification, thereby reducing the manpower for classifying literatures in the fishery professional field and solving the problem of inconsistent standards caused by subjective evaluation in the process of human participation when classifying literatures in the fishery professional field.

[0083] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example.

[0084] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A subject classification method for fishery literature, characterized in that: The following steps are involved: Obtain thesaurus information of fishery literature and target classification literature; constructing a thesaurus mask image according to the thesaurus information; Constructing a subject relationship clustering image according to the target classification documents; Superimposing the thesaurus mask image onto the subject relationship cluster image to obtain a residual image; The classification type corresponding to the target classification document is determined according to the residual image.

2. The subject classification method of fishery literature according to claim 1 is characterized in that: The step of constructing a thesaurus mask image according to the thesaurus information includes: Marking the subject relationships of the words in the thesaurus information; Get the similarity of the parts of speech of the annotated words; Clustering the annotated words according to the similarity relationship between the parts of speech; The thesaurus mask image is constructed based on the clustered vocabulary.

3. The subject classification method of fishery literature according to claim 2 is characterized in that: The step of constructing the thesaurus mask image according to the clustered vocabulary includes: Obtaining a normalized value of the distance between the clustered vocabulary and the cluster center point to which the corresponding subject belongs; Drawing a corresponding cluster image based on the clustered vocabulary, and using the normalized value as a transparency index of the cluster image; The transparency index is kept unchanged, and the cluster image is converted into a grayscale image to obtain the descriptor mask image.

4. The subject classification method of fishery literature according to claim 1 is characterized in that: The step of constructing a subject relationship clustering image according to the target classification documents includes: Preprocessing the target classified documents, wherein the preprocessing includes word segmentation and stop word removal; Construct a subject relationship clustering image based on the preprocessed target classification documents.

5. The subject classification method of fishery literature according to claim 4 is characterized in that: The step of constructing a subject relationship clustering image based on the preprocessed target classification documents includes: Extracting subject words from the preprocessed target classified documents; Obtaining the subject topic corresponding to the extracted subject word and the similarity between the extracted subject word and the subject topic; A subject relationship clustering image is constructed according to the proximity.

6. The subject classification method of fishery literature according to claim 1 is characterized in that: The step of superimposing the thesaurus mask image onto the subject relationship cluster image to obtain a residual image includes: The grayscale space of the thesaurus mask image is multiplied by the color space of the subject relationship cluster image, and the inverse value of the transparency space of the thesaurus mask image is multiplied by all channel values ​​of the subject relationship cluster image to obtain the residual image.

7. The subject classification method of fishery literature according to claim 1 is characterized in that: The step of determining the classification type corresponding to the target classification document according to the residual image includes: Determine a corresponding closed area image according to the residual image; Determine the maximum area cluster corresponding to the residual image according to the closed area image; The classification type corresponding to the target classification document is determined according to the maximum area cluster corresponding to the residual image.

8. The subject classification method of fishery literature according to claim 1 is characterized in that: The step of determining the classification type corresponding to the target classification document according to the maximum area cluster corresponding to the residual image includes: Obtaining the cluster center of the largest area cluster corresponding to the residual image; The classification type corresponding to the cluster center in the thesaurus mask image is used as the classification type corresponding to the target classification document.

9. A subject classification system for fishery literature, the system comprising: An acquisition module is used to obtain thesaurus information of fishery literature and target classification literature; A first construction module, used for constructing a thesaurus mask image according to the thesaurus information; A second construction module is used to construct a subject relationship clustering image according to the target classification documents; A third construction module is used to superimpose the thesaurus mask image onto the subject relationship cluster image to obtain a residual image; A determination module is used to determine the classification type corresponding to the target classification document determined by the residual image.

10. An electronic device, comprising the subject classification system for fishery documents as described in claim 9; or, a processor, a memory, and a subject classification program for fishery documents stored in the memory and executable on the processor, wherein the subject classification program for fishery documents, when executed by the processor, implements the subject classification method for fishery documents as described in the embodiment of the first aspect of the present invention.

11. A computer-readable storage medium, on which a subject classification program for fishery literature is stored, and when the subject classification program for fishery literature is executed by a processor, the subject classification method for fishery literature as described in any one of claims 1 to 8 is implemented.