Low-altitude economics teaching text classification method and system based on topic model
By constructing a low-altitude economics teaching terminology association network and subject structure mining, an optimized subject classification model is generated, which solves the accuracy and consistency problems of low-altitude economics teaching text classification and achieves efficient text classification and resource utilization.
Patent Information
- Application Number
- CN202511052745.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Low-altitude economics teaching texts lack a unified and standardized classification standard. Traditional classification methods are inefficient and difficult to guarantee accuracy and consistency. Existing text classification methods fail to fully explore the deep semantic information of teaching texts, resulting in inaccurate classification, which affects the integration of teaching resources and the quality of talent training.
Construct a low-altitude economic teaching terminology association network, conduct topic structure mining and semantic vector representation, generate an optimized topic classification model, including a topic feature library and classification rules, and achieve accurate classification through topic feature matching and decision processing.
It significantly improves the accuracy, objectivity and efficiency of low-altitude economic teaching text classification, ensures the accuracy and consistency of classification results, and supports the efficient use of teaching resources and talent training.
Smart Images

Figure CN120561291B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a low-altitude economics teaching text classification method and system based on a topic model. Background Art
[0002] As education in the field of low-altitude economy continues to develop, teaching text resources are becoming increasingly abundant, covering a wide range of knowledge content from basic theory to practical applications. However, the current management and utilization of low-altitude economy teaching texts face many challenges. On the one hand, since low-altitude economy is an emerging field, its teaching texts lack unified and standardized classification standards. Traditional manual classification methods are not only inefficient, but the classification results are easily affected by the subjective factors of the classifier, making it difficult to ensure the accuracy and consistency of the classification. On the other hand, existing text classification methods are mostly based on simple keyword matching or shallow semantic analysis. They do not fully consider the complex associations between terms in low-altitude economy teaching texts and the deep semantic information of the texts. They are unable to accurately explore the thematic features contained in the texts, making it difficult to achieve accurate classification of low-altitude economy teaching texts. This is not conducive to the effective integration and efficient utilization of teaching resources, and also restricts the quality and efficiency of talent training in the field of low-altitude economy. Summary of the Invention
[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, a method for classifying low-altitude economic teaching texts based on a topic model is provided, the method comprising:
[0004] Obtaining a low-altitude economics teaching text collection, performing term extraction and association analysis on the low-altitude economics teaching text collection, and constructing a low-altitude economics teaching term association network, wherein the nodes of the low-altitude economics teaching term association network represent teaching terms extracted from the low-altitude economics teaching text collection, and the edges of the low-altitude economics teaching term association network represent the co-occurrence association strength between different teaching terms;
[0005] Performing topic structure mining based on the low-altitude economics teaching terminology association network to generate an initial low-altitude economics teaching topic set, wherein the initial low-altitude economics teaching topic set includes multiple topic units with term aggregation characteristics, and each topic unit is composed of a teaching term set with a close co-occurrence correlation strength in the low-altitude economics teaching terminology association network;
[0006] Performing semantic vector representation processing on text units in the low-altitude economics teaching text set to obtain a text semantic vector set, and performing topic semantic consistency optimization processing on the initial low-altitude economics teaching topic set using the text semantic vector set to generate an optimized low-altitude economics teaching topic set;
[0007] Constructing a low-altitude economics teaching text topic classification model based on the optimized low-altitude economics teaching topic set, wherein the low-altitude economics teaching text topic classification model includes a topic feature library for topic matching and a topic classification rule set for classification decision-making;
[0008] The low-altitude economics teaching text to be classified is input into the low-altitude economics teaching text subject classification model, subject feature matching processing is performed through the subject feature library, and classification decision processing is performed in combination with the subject classification rule set to generate a subject classification result corresponding to the low-altitude economics teaching text to be classified.
[0009] On the other hand, the present invention also provides a low-altitude economic teaching text classification system based on a topic model, including a processor and a machine-readable storage medium, the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0010] Based on the above aspects, the beneficial effects of the present invention are as follows:
[0011] By obtaining a collection of low-altitude economics teaching texts and performing term extraction and association relationship analysis, a low-altitude economics teaching term association network is constructed, which accurately captures the co-occurrence association strength between teaching terms. Based on the low-altitude economics teaching term association network, the topic structure mining is performed. The generated initial low-altitude economics teaching topic set can reasonably aggregate teaching terms with close associations to form topic units with clear semantic orientations. The text units are represented by semantic vectors and the initial topic set is optimized for semantic consistency, which further improves the accuracy and representativeness of the topics and makes the generated topics more consistent with the actual semantic characteristics of low-altitude economics teaching texts. The text topic classification model constructed based on the optimized topic set includes a topic feature library and a topic classification rule set, which can achieve accurate topic matching and classification decisions for low-altitude economics teaching texts. The overall method, from term association to topic mining, to semantic optimization and classification modeling, significantly improves the accuracy, objectivity and efficiency of low-altitude economics teaching text classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 The present invention provides a method for classifying low-altitude economic teaching texts based on a topic model.
[0013] Figure 2 Schematic diagram of exemplary hardware and software components of a low-altitude economics teaching text classification system based on a topic model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a method for classifying low-altitude economic teaching texts based on a topic model provided by an embodiment of the present invention. The method for classifying low-altitude economic teaching texts based on a topic model is introduced in detail below.
[0015] Step S110: Obtain a low-altitude economics teaching text collection, perform term extraction and association analysis on the low-altitude economics teaching text collection, and construct a low-altitude economics teaching term association network, wherein the nodes of the low-altitude economics teaching term association network represent teaching terms extracted from the low-altitude economics teaching text collection, and the edges of the low-altitude economics teaching term association network represent the co-occurrence association strength between different teaching terms.
[0016] In this embodiment, obtaining a collection of low-altitude economic teaching texts is the starting step of the entire method. These text collections cover multiple aspects of teaching in the field of low-altitude economics, such as drone system structure, low-altitude flight weather conditions, low-altitude airspace management regulations, drone control technology, low-altitude logistics and distribution processes, etc. The channels for acquisition can be professional low-altitude economic teaching resource databases, authorized teaching platform document libraries, and related academic research databases. During the acquisition process, it is necessary to perform preliminary format unification processing on the texts, converting texts in different formats, such as PDF format, Word format, TXT format, etc., into a unified processable text format.
[0017] Term extraction and association analysis of the acquired low-altitude economics teaching text collection are key to constructing a low-altitude economics teaching terminology association network. Term extraction requires accurately identifying specialized terms characteristic of low-altitude economics teaching from the text, while association analysis explores the co-occurrence of these terms to determine the strength of their associations. This series of processes ultimately creates a low-altitude economics teaching terminology association network that clearly displays the relationships between individual teaching terms.
[0018] Step S111: performing word segmentation processing on each text unit in the low-altitude economics teaching text set to obtain a text word segmentation sequence set.
[0019] When processing the low-altitude economic teaching text collection, each text unit must first be segmented. A text unit can be a complete teaching article, a teaching chapter, or a teaching paragraph. The purpose of word segmentation is to split the continuous text into independent words for subsequent operations such as term extraction. For example, for the text unit "The flight control system of the drone is composed of multiple sensors", after word segmentation processing, the resulting word segmentation sequence is "drone", "of", "flight", "control", "system", "by", "multiple", "sensor", and "composed".
[0020] The tool used for word segmentation needs to be able to adapt to the characteristics of professional terms in the low-altitude economy field to ensure the accuracy of word segmentation. During the processing, for some compound terms, such as "low-altitude airspace management", "UAV communication link", etc., the word segmentation tool should be able to correctly split them to avoid incorrect word segmentation results. At the same time, for punctuation marks and special characters appearing in the text, they will be treated as delimiters during word segmentation and will not be included in the word segmentation sequence. Integrate the word segmentation results of all text units to form a set of text word segmentation sequences.
[0021] Step S112: Perform stop word filtering on the set of text word segmentation sequences to remove lexical units without actual teaching term meaning and retain candidate term units with teaching professional meaning.
[0022] After obtaining the set of text word segmentation sequences, stop word filtering needs to be performed. Stop words refer to those words that frequently appear in the text but have no actual teaching term meaning, such as "of", "in", "and", "for", "a", etc. The existence of these words will interfere with subsequent term extraction and association relationship analysis, so they must be removed.
[0023] During the filtering process, it can be carried out according to a pre-constructed stop word list. The stop word list contains common meaningless words in low-altitude economy teaching texts. For each word segmentation sequence in the set of text word segmentation sequences, check the lexical units one by one. If the lexical unit exists in the stop word list, it will be deleted from the word segmentation sequence; if it does not exist, it will be retained as a candidate term unit with teaching professional meaning. For example, in the word segmentation sequence "The flight control system of the UAV consists of multiple sensors", "of", "by", "multiple" belong to stop words, and the candidate term units obtained after filtering are "UAV", "flight", "control system", "sensor", "consist". Through the above processing, the text content can be effectively streamlined and the term units with professional meaning can be highlighted.
[0024] Step S113: Perform term recognition on the candidate term units to extract teaching terms with the characteristics of low-altitude economy field teaching. The teaching terms include noun terms representing low-altitude economy concepts, verb terms representing operation processes, and adjective terms representing attribute characteristics.
[0025] The candidate term units obtained after stop word filtering still need to undergo term recognition processing to extract teaching terms with teaching characteristics in the low-altitude economy field. Term recognition processing will combine professional knowledge in the low-altitude economy field to screen and judge the candidate term units. Noun terms representing low-altitude economy concepts are an important part of teaching terminology, such as "drone," "low-altitude airspace," "flight control system," "sensor," "meteorological data," etc. These terms clearly refer to various concepts in the low-altitude economy field.
[0026] Verb terms that represent operational procedures, such as "take-off", "landing", "navigation", "monitoring", "transmission", etc., describe the processes and actions of various operations in the low-altitude economy field. Adjective terms that represent attribute characteristics, such as "high precision", "stable", "real-time", "miniaturization", etc., are used to describe the attributes and characteristics of various things in the low-altitude economy field. In the recognition process, it is possible to determine whether the candidate term unit belongs to the above three categories of teaching terms by comparing the professional low-altitude economic terminology library and combining the context of the term in the text. For candidate term units that are determined to be teaching terms, they are extracted; those that do not belong are excluded.
[0027] Step S114: Counting the co-occurrence frequencies of the teaching terms in the low-altitude economics teaching text set, and calculating the co-occurrence association strength between different teaching terms based on the co-occurrence frequencies, wherein the co-occurrence association strength is positively correlated with the co-occurrence frequency.
[0028] After extracting the teaching terms, we need to count their co-occurrence frequency within the low-altitude economics teaching text collection. Co-occurrence frequency refers to the number of times two teaching terms appear together within the same text unit. For example, if "drone" and "sensor" frequently appear together across multiple teaching texts, their co-occurrence frequency will be relatively high; on the other hand, if "low-altitude airspace" and "weather data" appear together less frequently, their co-occurrence frequency will be low.
[0029] To calculate co-occurrence frequencies, we can iterate through all text units in the low-altitude economics teaching text collection and, for each unit, record the number of times each pair of teaching terms co-occurs. Once the statistics are complete, we calculate the co-occurrence association strength between different teaching terms based on the co-occurrence frequencies. The co-occurrence association strength is calculated to ensure a positive correlation with the co-occurrence frequency: higher co-occurrence frequencies increase the co-occurrence association strength, while lower co-occurrence frequencies decrease the co-occurrence association strength. For example, if the co-occurrence frequency of "drone" and "sensor" is A, and the co-occurrence frequency of "drone" and "navigation" is B, and A is greater than B, then the co-occurrence association strength between "drone" and "sensor" is greater than the co-occurrence association strength between "drone" and "navigation."
[0030] Step S115: Using the teaching terms as nodes and the co-occurrence association strength as edge weights, construct the low-altitude economics teaching term association network, wherein the node attributes of the low-altitude economics teaching term association network include the domain attribution labels of the teaching terms and the term occurrence frequencies.
[0031] In this low-altitude economics teaching terminology association network, each teaching term is considered a node. For example, "drone," "low-altitude airspace," "flight control system," and "sensor" are all nodes in the network. The edges between nodes represent the associations between different teaching terms, and the edge weights are the co-occurrence association strengths calculated previously.
[0032] The node attributes of the low-altitude economics teaching terminology association network include the domain attribution label of the teaching term and the term occurrence frequency. The domain attribution label is used to identify the specific low-altitude economics sub-field to which the teaching term belongs, such as "unmanned aerial vehicle system", "airspace management", "flight operations", etc. The term occurrence frequency refers to the total number of times the teaching term appears in the entire low-altitude economics teaching text collection. For example, the domain attribution label of the "unmanned aerial vehicle" node may be "unmanned aerial vehicle system", and its occurrence frequency may be multiple times in the entire text collection. By constructing the above-mentioned low-altitude economics teaching terminology association network, the degree of correlation and importance between various teaching terms can be intuitively displayed.
[0033] Step S120: Perform topic structure mining based on the low-altitude economics teaching terminology association network to generate an initial low-altitude economics teaching topic set, wherein the initial low-altitude economics teaching topic set contains multiple topic units with term aggregation characteristics, and each topic unit is composed of a set of teaching terms with close co-occurrence association strength in the low-altitude economics teaching terminology association network.
[0034] After constructing a low-altitude economics teaching terminology association network, we can perform topic structure mining based on this network. The goal of topic structure mining is to identify sets of teaching terms with strong co-occurrence correlations within the association network, aggregate them into topic units, and generate an initial set of low-altitude economics teaching topics. These topic units reflect the underlying thematic content within the low-altitude economics teaching text. Each topic unit revolves around a core theme and is composed of a series of related teaching terms.
[0035] For example, in an association network, the co-occurrence correlation strength between teaching terms such as "drone," "flight control system," "sensor," and "navigation" is high, and they may be aggregated to form a topic unit about "drone flight control." Thematic structure mining uses specific algorithms and methods to analyze the nodes and edges in the association network and identify groups of terms with clustering characteristics. This process can transform complex term association networks into a collection of topic units with clear thematic meaning.
[0036] Step S121: Calculate the network structure characteristics of the low-altitude economic teaching terminology association network to obtain the degree centrality, betweenness centrality and closeness centrality parameters of each node. The degree centrality represents the number of direct connections between a node and other nodes, the betweenness centrality represents the frequency of the node as the shortest path intermediary between other nodes, and the closeness centrality represents the average shortest path length from the node to all other nodes.
[0037] To mine the topic structure of the low-altitude economics teaching terminology association network, we first need to calculate the network structure characteristic parameters of each node in the network, including degree centrality, betweenness centrality, and closeness centrality. Degree centrality indicates the number of direct connections a node has with other nodes. The higher the degree centrality of a node, the more direct connections it has with other nodes. For example, the node "drone" may be directly connected to multiple nodes such as "sensor," "flight control system," "navigation," and "takeoff," so its degree centrality is relatively high.
[0038] Betweenness centrality indicates how often a node acts as an intermediary in the shortest paths between other nodes. If a node frequently appears on the shortest paths between other nodes, its betweenness centrality is high. For example, the node "Flight Control" may frequently appear in the shortest paths between "UAV" and nodes such as "Landing" and "Navigation," thus having a high betweenness centrality. Closeness centrality indicates the average shortest path length from a node to all other nodes. The shorter the average shortest path length, the higher the closeness centrality. For example, if the node "Low-Altitude Flight" has a short average shortest path length to nodes such as "UAV," "Airspace Management," and "Weather Conditions," its closeness centrality is high. By calculating these parameters, the importance and position of each node in the network can be assessed from different perspectives.
[0039] Step S122: Calculate the comprehensive node importance index based on the degree centrality, betweenness centrality and closeness centrality parameters, sort the nodes in the low-altitude economics teaching terminology association network in descending order according to the comprehensive node importance index, and select the top-ranked nodes as the topic core candidate terms.
[0040] After obtaining the degree centrality, betweenness centrality, and closeness centrality parameters for each node, it is necessary to calculate a comprehensive node importance index based on these parameters. When calculating the comprehensive node importance index, these three parameters can be comprehensively considered, assigned certain weights, and then the comprehensive importance of each node can be obtained through weighted calculation. For example, the weights of degree centrality, betweenness centrality, and closeness centrality are determined based on their roles in assessing node importance. The comprehensive node importance index of each node is then obtained by multiplying the degree centrality by its weight, adding the betweenness centrality multiplied by its weight, and then adding the closeness centrality multiplied by its weight.
[0041] After the calculations are complete, all nodes in the low-altitude economics teaching terminology network are sorted in descending order based on their comprehensive node importance index. Nodes ranked higher in the network are generally more important and have a greater impact on other nodes, so these nodes are selected as candidate core terms. For example, after sorting, nodes such as "drone," "low-altitude airspace," and "flight control" might appear at the top of the list, becoming candidate core terms.
[0042] Step S123: using a community discovery algorithm to perform community division processing on the low-altitude economics teaching terminology association network, clustering teaching terms with significant co-occurrence association strength to form multiple terminology community units, each terminology community unit corresponding to a potential topic.
[0043] After selecting the core candidate terms for the subject, a community discovery algorithm is used to divide the low-altitude economics teaching terminology association network into communities. The community discovery algorithm can identify groups of nodes in the network that have close internal connections and relatively sparse external connections. These groups are terminology community units.
[0044] During the processing, the community discovery algorithm will cluster based on the co-occurrence correlation strength between nodes. Teaching terms with significant co-occurrence correlation strength will be clustered together to form a term community unit. Each term community unit corresponds to a potential topic, reflecting the teaching content of a specific aspect of low-altitude economics. For example, the co-occurrence correlation strength between teaching terms such as "drone", "sensor", "flight control system", and "navigation" is significant, and they may be divided into the same term community unit, corresponding to the potential topic of "drone navigation and control"; while teaching terms such as "low-altitude airspace", "airspace division", "management regulations", and "application for use" may form another term community unit, corresponding to the potential topic of "low-altitude airspace management". Through community division, the terms in the association network can be effectively grouped.
[0045] Step S124: Match the core candidate subject terms with the term community units, determine the core teaching terms of each term community unit, and take the core teaching terms as the center to select the teaching terms in the term community units whose co-occurrence correlation strength with the core teaching terms exceeds a preset threshold to form an initial subject term set.
[0046] After obtaining the term community units, it is necessary to match the subject core candidate terms with these term community units. For each term community unit, determine whether there is a subject core candidate term in it. If so, the subject core candidate term is the core teaching term of the term community unit. If not, the teaching term with the highest comprehensive node importance index is selected from the term community unit as the core teaching term.
[0047] After determining the core teaching terms, take them as the center and select the teaching terms in the term community unit whose co-occurrence correlation strength with the core teaching terms exceeds the preset threshold. The preset threshold is set according to the distribution of co-occurrence correlation strength of the entire association network, and is used to screen out teaching terms that are more closely related to the core teaching terms. These screened teaching terms and the core teaching terms together constitute the initial subject term set. For example, in the "UAV Navigation and Control" term community unit, the core teaching term is "UAV", and the teaching terms whose co-occurrence correlation strength with "UAV" exceeds the preset threshold include "sensor", "flight control system", "navigation", "take-off", "landing", etc. These terms together constitute the initial subject term set of the community unit.
[0048] Step S125: performing term semantic similarity analysis on the initial subject term sets, merging semantically similar initial subject term sets, and generating an initial low-altitude economics teaching subject set comprising multiple subject units, each subject unit comprising a core teaching term and multiple related teaching terms.
[0049] After forming the initial subject term set, it is necessary to perform a semantic similarity analysis on it. Semantic similarity analysis is performed by comparing the semantic meanings of teaching terms in different initial subject term sets. For teaching terms with similar semantics, it can be assumed that the initial subject term sets to which they belong are also similar.
[0050] For example, if there are two initial subject term sets, one contains terms such as "UAV navigation", "satellite positioning", and "path planning", and the other contains terms such as "UAV route", "global positioning", and "route design". After semantic similarity analysis, it is found that the terms in the two sets are semantically similar, then they are merged into one initial subject term set. Through the above merging process, duplication and redundancy of topics can be avoided. After the merger is completed, an initial low-altitude economic teaching subject set containing multiple subject units is generated. Each subject unit contains a core teaching term and multiple related teaching terms that are closely related to the core teaching term. For example, the core teaching term of a subject unit is "UAV", and related teaching terms include "sensor", "flight control system", "navigation", "take-off", "landing", etc.
[0051] Step S130: Perform semantic vector representation processing on the text units in the low-altitude economics teaching text set to obtain a text semantic vector set, and use the text semantic vector set to perform topic semantic consistency optimization processing on the initial low-altitude economics teaching topic set to generate an optimized low-altitude economics teaching topic set.
[0052] After generating an initial set of low-altitude economics teaching topics, we need to perform semantic vector representation on the text units in the low-altitude economics teaching text set. Semantic vector representation converts text units into computer-processable numerical values, facilitating subsequent semantic analysis and comparison. By converting each text unit into a corresponding semantic vector, we form a set of text semantic vectors.
[0053] Using the text semantic vector set, the initial low-altitude economics teaching topic set can be optimized for semantic consistency. This process primarily adjusts and refines the initial topic set by analyzing the consistency between the text semantic vectors and the semantics of the topic units. Topic units that don't closely match the text semantics are modified or deleted, while those with similar semantics are merged to ultimately generate an optimized low-altitude economics teaching topic set. This optimized topic set more accurately reflects the thematic content of the low-altitude economics teaching text, improving the accuracy of subsequent classification models.
[0054] Step S131: performing sentence segmentation processing on the text units in the low-altitude economics teaching text set, and segmenting each text unit into multiple sentence units.
[0055] When processing the semantic vector representation of text units in the low-altitude economics teaching text collection, sentence segmentation is first performed. Each text unit usually contains multiple sentences, and the purpose of sentence segmentation is to split the text unit into independent sentence units.
[0056] The segmentation process uses punctuation marks such as periods, question marks, and exclamation points as sentence end markers. For example, the text unit "The drone's flight relies on a flight control system. The flight control system receives data from sensors, analyzes and processes it, and then adjusts the drone's flight attitude based on the results." can be segmented into three sentence units: "The drone's flight relies on a flight control system.", "The flight control system receives data from sensors, analyzes and processes it," and "The drone's flight attitude is then adjusted based on the results." Sentence segmentation breaks down text units into smaller processing units, facilitating more detailed semantic encoding.
[0057] Step S132: Perform word-level semantic encoding processing on each sentence unit, and convert each word in the sentence unit into a word semantic vector, wherein the word semantic vector is generated by a pre-trained language model.
[0058] The pre-trained language model is trained on a large-scale text corpus and can capture the semantic information and contextual associations of words. In this embodiment, the selected pre-trained language model needs to be fine-tuned with the corpus in the low-altitude economy field to better adapt to the semantic encoding requirements of low-altitude economy teaching terms.
[0059] For each word in a sentence unit, it is input into a pre-trained language model, which outputs the corresponding word semantic vector. These word semantic vectors are multi-dimensional numerical vectors, with each dimension representing the characteristics of the word in a specific semantic space. For example, the words "drone," "through," "sensor," "obtain," and "flight data" in the sentence unit "drone obtains flight data through sensors" will all generate corresponding word semantic vectors after semantic encoding. Each vector contains multiple numerical values to reflect its unique semantic characteristics. This word-level semantic encoding process can convert each word in a sentence into a semantic vector form that can be understood by a computer.
[0060] Step S133: performing sequence concatenation processing on the word semantic vectors in the sentence unit to generate a sentence semantic vector sequence, wherein the length of the sentence semantic vector sequence is consistent with the number of words in the sentence unit.
[0061] After obtaining the word semantic vectors for each word in a sentence unit, these word semantic vectors need to be concatenated. Sequence concatenation involves concatenating the word semantic vectors in the order in which they appear in the sentence unit to form a continuous vector sequence, i.e., a sentence semantic vector sequence.
[0062] For example, the words in the sentence unit "The sensor monitors the flight state of the drone" are, in order, "sensor", "monitors", "drone", "of", "flight state", and their corresponding word semantic vectors are V1, V2, V3, V4, V5 respectively. After sequence concatenation processing, the generated sentence semantic vector sequence is [V1, V2, V3, V4, V5], and the length of this sentence semantic vector sequence is 5, which is the same as the number of words in the sentence unit.
[0063] The sentence semantic vector sequence retains the order information of the words in the sentence, which is crucial for understanding the semantics of the sentence. Because different word orders may lead to significant differences in the meaning expressed by the sentence. The sentence semantic vector sequence generated through sequence concatenation can fully reflect the arrangement order of the words in the sentence and their respective semantic vectors.
[0064] Step S134: Use an attention mechanism to perform weighted aggregation processing on the sentence semantic vector sequence to generate a sentence-level semantic vector. The attention mechanism assigns different weight parameters according to the semantic contribution degree of the words in the sentence unit.
[0065] After obtaining the sentence semantic vector sequence, use an attention mechanism to perform weighted aggregation processing on it. The core idea of the attention mechanism is to assign corresponding weight parameters to each word semantic vector according to the contribution degree of each word in the sentence unit to the overall semantics of the sentence. Words with a high contribution degree will be assigned higher weight parameters and occupy a larger proportion in the aggregation process; words with a low contribution degree will be assigned lower weight parameters and have relatively less influence.
[0066] For example, in the sentence unit "The flight control system adjusts the flight attitude of the drone", the words "flight control system", "adjusts", "drone", "flight attitude" have a relatively high contribution degree to the sentence semantics, while words such as "of" have a relatively low contribution degree. The attention mechanism will assign higher weight parameters to the word semantic vectors corresponding to "flight control system", "adjusts", "drone", "flight attitude", and assign lower weight parameters to the word semantic vector corresponding to "of".
[0067] Weighted aggregation processing is to perform weighted calculation on each word semantic vector and its corresponding weight parameter, and then splice and integrate all the weighted vectors to generate a sentence-level semantic vector. This sentence-level semantic vector synthesizes the semantic information of all the words in the sentence and highlights the semantic contribution of the key words. For example, after weighted aggregation processing, the sentence unit "The flight control system adjusts the flight attitude of the drone" will generate a sentence-level semantic vector, which can better represent the semantic content of the entire sentence.
[0068] Step S135: performing average pooling processing on all sentence-level semantic vectors in the text unit to generate a text semantic vector at the text unit level, and combining the text semantic vectors of all text units to form the text semantic vector set.
[0069] A text unit consists of multiple sentence units, each of which corresponds to a sentence-level semantic vector. To obtain a text unit-level semantic vector, average pooling is performed on all sentence-level semantic vectors within the text unit. Average pooling averages the corresponding dimensions of all sentence-level semantic vectors to obtain the text unit's semantic vector.
[0070] For example, a text unit contains three sentence units, whose sentence-level semantic vectors are S1, S2, and S3, respectively, each with multiple dimensions. During average pooling, the values of each corresponding dimension of S1, S2, and S3 are added together and divided by 3. The result is used as the value of the corresponding dimension of the text semantic vector. Through this processing method, the text semantic vector can comprehensively reflect the semantic information of all sentences in the text unit and provide a holistic representation of the semantics of the entire text unit.
[0071] The text semantic vectors of all text units in the low-altitude economics teaching text set are collected and arranged in the order of the text units to form a text semantic vector set. Each text semantic vector in the text semantic vector set corresponds to a text unit.
[0072] While the initial low-altitude economics teaching topic set can reflect certain thematic content, some topic units may not match the text semantics well, and there may be overlap or redundancy between topics. Using the text semantic vector set to optimize the topic semantic consistency is to improve the consistency between the topic set and the text semantics, making the topic units more accurate and standardized.
[0073] The optimization process primarily involves filtering, merging, and adjusting topic units. The rationality of topic units is determined by analyzing the correlation between the text semantic vectors and the semantic vectors of the topic units. The resulting optimized low-altitude economics teaching topic set more accurately captures the thematic content of the low-altitude economics teaching text.
[0074] Step S136: Performing subject semantic vector representation processing on each subject unit in the initial low-altitude economics teaching subject set, averaging and fusing the semantic vectors of the core teaching terms and related teaching terms in the subject unit to generate a subject unit semantic vector.
[0075] Each topic unit consists of a core teaching term and multiple related teaching terms. To facilitate comparative analysis with the text semantic vector, each topic unit needs to be represented by a topic semantic vector. This processing first obtains the word semantic vectors for the core teaching term and all related teaching terms in the topic unit. These word semantic vectors were generated using the pre-trained language model in step S132.
[0076] Then, these word semantic vectors are averaged and fused. The average fusion process is to average the corresponding dimensions of the word semantic vectors of the core teaching terms and related teaching terms respectively to obtain the topic unit semantic vector of the topic unit. For example, the core teaching term of a topic unit is "drone", and the related teaching terms include "sensor", "flight control system", and "navigation". Their word semantic vectors are W1, W2, W3, and W4 respectively. During the average fusion process, the values of the corresponding dimensions of these four vectors can be added and divided by 4. The result is the topic unit semantic vector of the topic unit. The topic unit semantic vector generated in the above manner can comprehensively reflect the semantic information of all teaching terms in the topic unit, which is convenient for the subsequent calculation of the similarity with the text semantic vector.
[0077] Step S137: Calculate the cosine similarity between each text semantic vector and each topic unit semantic vector in the text semantic vector set to obtain a text topic association matrix, where the rows of the text topic association matrix represent text units, the columns represent topic units, and the matrix elements represent the semantic association strength between text units and topic units.
[0078] Cosine similarity measures the directional similarity between two vectors. A larger cosine similarity indicates a closer semantic relationship between the two vectors. After obtaining the set of text semantic vectors and the topic unit semantic vector for each topic unit, we need to calculate the cosine similarity between each text semantic vector and each topic unit semantic vector.
[0079] During the calculation process, the cosine similarity formula is used to calculate the similarity between each text semantic vector T and each topic unit semantic vector M in the text semantic vector set. All similarity values are arranged according to the correspondence between text units and topic units to form a text topic association matrix.
[0080] For example, the text semantic vector set contains three text semantic vectors, T1, T2, and T3, corresponding to three text units respectively; the initial low-altitude economics teaching topic set contains two topic unit semantic vectors, M1 and M2. The calculated similarity between T1 and M1 is 0.8, and the similarity between T1 and M2 is 0.3; the similarity between T2 and M1 is 0.2, and the similarity between T2 and M2 is 0.7; the similarity between T3 and M1 is 0.6, and the similarity between T3 and M2 is 0.5. The text topic association matrix is then a 3-row, 2-column matrix, where the matrix elements are the similarity values calculated above, where rows correspond to text units and columns correspond to topic units. Each element represents the strength of the semantic association between the corresponding text unit and the topic unit.
[0081] Step S138: Based on the text topic association matrix, count the number of text units associated with each topic unit and the average semantic association strength, and delete topic units whose number of associated text units is lower than a preset number threshold or whose average semantic association strength is lower than a preset strength threshold.
[0082] The text topic association matrix reflects the strength of the semantic association between text units and topic units, which allows for further screening of each topic unit. First, the number of text units associated with each topic unit is counted. That is, the number of elements with a similarity value greater than 0 in each column (corresponding to the topic unit) in the text topic association matrix is counted. The text units corresponding to these elements are the text units associated with the topic unit.
[0083] At the same time, the average semantic relevance strength of each topic unit is calculated, that is, the similarity values of all elements in each column are averaged. The preset quantity threshold is set based on the size of the text collection and the expected coverage of the topic unit, and the preset strength threshold is set based on the overall distribution of similarity values in the text topic relevance matrix.
[0084] For topic units whose number of associated text units is lower than the preset number threshold, it means that the topic unit lacks sufficient text support in the text collection and may be a redundant or unimportant topic, which needs to be deleted. For topic units whose average semantic association strength is lower than the preset strength threshold, it means that the semantic association between the topic unit and the related text units is not close enough, and it also needs to be deleted. For example, if the number of text units associated with a topic unit is 5, which is lower than the preset number threshold of 8, or its average semantic association strength is 0.3, which is lower than the preset strength threshold of 0.5, then the topic unit will be deleted. Through the above screening process, those topic units that lack text support or have loose semantic associations can be eliminated, thereby improving the quality of the topic collection.
[0085] Step S139: Calculate the topic similarity of the remaining topic units, merge the topic units whose cosine similarity between the semantic vectors exceeds the preset merging threshold, and retain the core teaching terms with the highest semantic contribution in the merged topic units.
[0086] After screening, some of the remaining topic units may have semantic similarities, requiring topic similarity calculation. This topic similarity calculation is achieved by calculating the cosine similarity between the topic unit semantic vectors of the remaining topic units. This calculation method is the same as the cosine similarity calculation of the text semantic vector and the topic unit semantic vector in step S142.
[0087] The preset merging threshold is set based on the allowed semantic similarity between topic units. When the cosine similarity between the semantic vectors of two topic units exceeds this threshold, it indicates that the two topic units are semantically similar and need to be merged. During the merging process, the teaching terms contained in the two topic units are integrated, and duplicate terms are removed to form a new topic unit.
[0088] In the merged subject unit, the core teaching term with the highest semantic contribution needs to be retained. The semantic contribution is determined based on the comprehensive importance of the core teaching term in the subject unit before the merger, and the strength of its semantic association with all the teaching terms after the merger. For example, there are two subject units. The core teaching term of subject unit A is "UAV navigation", and the core teaching term of subject unit B is "UAV route planning". The cosine similarity of their subject unit semantic vectors exceeds the preset merger threshold. After the merger, a new subject unit is formed. After evaluation, "UAV navigation" has a higher semantic contribution, so it is used as the core teaching term of the merged subject unit. Through the above merging process, the redundancy of the subject unit can be reduced and the subject set can be made more refined.
[0089] Step S1310: reconstructing the subject term set based on the merged subject units, and performing semantic hierarchical sorting on the teaching terms in the subject term set to generate an optimized low-altitude economic teaching subject set with clear semantic boundaries.
[0090] The merged subject units require a new subject term set to be reconstructed. This involves organizing the core teaching terms and all related teaching terms contained in the merged subject unit to form a new subject term set. The teaching terms in the subject term set are then semantically organized to clarify the semantic relationships between them.
[0091] Semantic hierarchical organization divides teaching terms into different levels based on their semantic connotations and interrelationships. For example, in a thematic unit with "UAV navigation" as the core teaching term, teaching terms such as "satellite positioning," "inertial navigation," and "path planning" are specific implementation methods of "UAV navigation" and belong to the next level of terms. Terms such as "navigation accuracy" and "navigation error" describe the performance of "UAV navigation" and belong to another level of terms.
[0092] By organizing the semantic hierarchy, the teaching terms in the subject term set can be organized into a system with a clear structure and well-defined semantic boundaries. This further clarifies the position and role of each term within the subject unit. The resulting optimized low-altitude economics teaching subject set features clear semantic boundaries between subject units and precise subject content, better reflecting the thematic structure of the low-altitude economics teaching text.
[0093] Step S140: constructing a low-altitude economics teaching text topic classification model based on the optimized low-altitude economics teaching topic set, wherein the low-altitude economics teaching text topic classification model includes a topic feature library for topic matching and a topic classification rule set for classification decision-making.
[0094] The purpose of constructing the low-altitude economics teaching text subject classification model is to realize the automatic subject classification of low-altitude economics teaching text. The low-altitude economics teaching text subject classification model includes two core parts: subject feature library and subject classification rule set.
[0095] The topic feature library is used to store the feature information of topic units and provide a basis for topic matching. The topic classification rule set is used to formulate the basis and standards for classification decisions and ensure the accuracy of classification results. By converting the optimized topic set into features and rules that can be used by the model, the low-altitude economics teaching text topic classification model can be equipped with the ability to thematically classify new low-altitude economics teaching texts.
[0096] Step S141: performing feature extraction processing on each topic unit in the optimized low-altitude economics teaching topic set, converting the teaching terms contained in the topic unit into a topic feature vector, wherein the dimension of the topic feature vector is consistent with the dimension of the text semantic vector.
[0097] Performing feature extraction on each topic unit in the optimized low-altitude economics teaching topic set is a key step in building the topic feature library. This feature extraction process is similar to the process of generating the topic unit semantic vector in step S141, converting the teaching terms contained in the topic unit into vector form. However, here, a topic feature vector is generated, whose dimensions are consistent with the text semantic vector, facilitating subsequent topic matching calculations.
[0098] Specifically, for the core teaching terms and related teaching terms in each topic unit, their word semantic vectors are obtained, and then the topic feature vector is generated by average fusion. The average fusion method is the same as the method of generating the topic unit semantic vector in step S136, that is, the corresponding dimensions of all word semantic vectors are averaged. For example, after the teaching terms contained in a certain topic unit are converted and averaged, the generated topic feature vector has the same number of dimensions as the text semantic vector, and each dimension represents the characteristics of the topic unit in the corresponding semantic space. Through the above processing, each topic unit is converted into a topic feature vector.
[0099] Step S142: The topic feature vectors of all topic units are combined to form a topic feature library, which contains identification information, core teaching terms and corresponding topic feature vectors of each topic unit.
[0100] After obtaining the topic feature vectors for each topic unit, they are combined to form a topic feature library. This library contains not only the topic feature vectors but also the identification information and core teaching terms for each topic unit. Identification information is a symbol or code used to uniquely distinguish different topic units, such as "T001" and "T002," facilitating the indexing and management of topic units in the model.
[0101] Core teaching terms intuitively reflect the subject content of a topic unit and correspond one-to-one with the topic feature vectors. For example, a record in the topic feature database might contain the identifier "T001," the core teaching term "UAV flight control," and the corresponding topic feature vector. This information from all topic units is organized according to a specific structure to form a complete topic feature database. The topic feature database is a key data structure used for topic matching in the topic classification model for low-altitude economics teaching texts.
[0102] Step S143: collecting a set of low-altitude economics teaching text samples with subject labels, wherein each sample in the set of low-altitude economics teaching text samples includes text content and a corresponding subject unit label.
[0103] Constructing a set of topic classification rules requires a collection of low-altitude economics teaching text samples with topic annotations. These sample collections are obtained through manual annotation. The annotators perform topic annotations on the low-altitude economics teaching texts based on the topic units in the optimized low-altitude economics teaching topic collection, assigning corresponding topic unit labels to each text sample.
[0104] For example, a text sample about the working principle of a drone flight control system can be labeled with the subject unit label "drone flight control"; a text sample introducing the low-altitude airspace demarcation standards can be labeled with the subject unit label "low-altitude airspace management." The sample set needs to cover all the subject units in the optimized subject set, and the number of samples corresponding to each subject unit should be as balanced as possible to ensure that the subject classification rule set obtained from subsequent training has good classification performance. At the same time, the size of the sample set should be large enough to ensure that the trained rules have generalization capabilities and can adapt to different low-altitude economic teaching texts.
[0105] Step S144: performing text semantic vector representation processing on the sample texts in the low-altitude economics teaching text sample set to obtain a sample semantic vector set.
[0106] Each sample text in the low-altitude economics teaching text sample set is processed for text semantic vector representation. The processing process is exactly the same as the semantic vector representation process for text units in steps S131 to S135. First, the sample text is segmented into sentence units, and each sentence unit is semantically encoded at the word level to generate a word semantic vector. Then, the word semantic vectors are sequentially spliced to generate a sequence of sentence semantic vectors. Then, an attention mechanism is used for weighted aggregation to obtain a sentence-level semantic vector. Finally, all sentence-level semantic vectors are average-pooled to generate the text semantic vector of the sample text, namely the sample semantic vector.
[0107] Combining the sample semantic vectors of all sample texts forms a sample semantic vector set. Each sample semantic vector in this sample semantic vector set is associated with the corresponding sample text and its topic unit label. For example, if the topic unit label of a sample text is "UAV communication technology," its corresponding sample semantic vector contains various semantic information related to drone communication in the text, such as the semantic features of terms such as "data transmission," "signal strength," and "communication protocol."
[0108] Step S145: Based on the sample semantic vector set and the corresponding topic unit labels, a supervised learning algorithm is used to train a topic classification rule set, where the topic classification rule set includes a mapping relationship between the sample semantic vector and the topic unit label and a classification decision threshold.
[0109] Once we have a set of sample semantic vectors and corresponding topic unit labels, we can use a supervised learning algorithm to train a set of topic classification rules. This algorithm can learn the mapping relationship between sample semantic vectors and topic unit labels from labeled sample data and determine the corresponding classification decision threshold.
[0110] During the training process, the sample semantic vector is used as input, and the corresponding topic unit label is used as the desired output. The algorithm continuously adjusts the model parameters so that the model can accurately predict the corresponding topic unit label based on the input sample semantic vector. Determining the classification decision threshold is a key step in the training process. It is used to determine whether the text corresponding to the sample semantic vector belongs to a specific topic unit. For example, when the degree of match between the sample semantic vector and the feature vector of a topic unit exceeds the classification decision threshold corresponding to the topic unit, the sample is determined to belong to the topic unit. The set of topic classification rules obtained through training can effectively guide subsequent classification decisions for the text to be classified.
[0111] Step S146: Integrate the subject feature library and the subject classification rule set into the low-altitude economics teaching text subject classification model, and the low-altitude economics teaching text subject classification model further includes a model input interface and a classification result output interface.
[0112] After the topic feature library and topic classification rule set are constructed, they need to be integrated into the low-altitude economics teaching text topic classification model. This integration process involves establishing a link between the topic feature library and the topic classification rule set, ensuring that when performing topic classification, the model can quickly retrieve relevant topic feature vectors from the topic feature library and apply the topic classification rule set to make decisions.
[0113] The low-altitude economics teaching text topic classification model also includes a model input interface and a classification result output interface. The model input interface is used to receive the low-altitude economics teaching text to be classified. During the reception process, the input text can be format checked and pre-processed to ensure that the input text meets the requirements of the model processing. The classification result output interface is used to output the final topic classification results. The output results include information such as the topic unit identifier, core teaching terms, and the matching confidence of the corresponding topic. They can also be converted into different output formats as needed, such as JSON format, XML format, etc., to facilitate subsequent processing by other systems or modules.
[0114] Step S150: input the low-altitude economics teaching text to be classified into the low-altitude economics teaching text subject classification model, perform subject feature matching processing through the subject feature library, and perform classification decision processing in combination with the subject classification rule set to generate a subject classification result corresponding to the low-altitude economics teaching text to be classified.
[0115] When a low-altitude economics teaching text needs to be classified, it is input into the low-altitude economics teaching text topic classification model. The model first performs topic feature matching on the text to be classified using the topic feature library to identify topic units that are semantically related to the text to be classified. It then makes a classification decision based on the matching results and the classification rules, ultimately determining the topic classification result for the text to be classified.
[0116] This process fully leverages the previously constructed topic feature library and trained classification rules to ensure accurate and reliable classification results. For example, if the text to be classified is about drone data transmission technology, the model can classify it into the topic unit "drone communication technology" through topic feature matching and classification decisions, and output the corresponding classification results.
[0117] Step S151: performing text preprocessing on the low-altitude economics teaching text to be classified, performing semantic vector representation processing on the preprocessed low-altitude economics teaching text to be classified, and generating a semantic vector of the text to be classified.
[0118] The low-altitude economy teaching text to be classified undergoes text preprocessing. This process is similar to steps S111 to S113 and includes word segmentation, stop word filtering, and term identification. First, the text to be classified is segmented to obtain a word sequence. Stop words are then filtered out, retaining candidate term units. Finally, from these candidate term units, teaching terms with characteristics specific to the low-altitude economy field are identified.
[0119] After preprocessing, the processed low-altitude economics teaching text to be classified is processed for semantic vector representation. The process is the same as steps S131 to S135. The text is first segmented into sentence units, and the words in each sentence unit are semantically encoded to obtain word semantic vectors. Then, through sequence concatenation, weighted aggregation using an attention mechanism, and average pooling, a semantic vector for the text to be classified is generated. This semantic vector contains the overall semantic information of the text to be classified.
[0120] Step S152: Read the topic feature vectors of all topic units from the topic feature library, and calculate the cosine similarity between the semantic vector of the text to be classified and each topic feature vector to obtain the topic matching degree between the text to be classified and each topic unit.
[0121] The topic feature vectors for all topic units are read from the topic feature library. Each topic feature vector corresponds to the semantic features of a topic unit. The cosine similarity between the semantic vector of the text to be classified and each topic feature vector is then calculated. The cosine similarity calculation reflects the directional similarity between two vectors; a larger value indicates a higher semantic similarity between the two vectors.
[0122] The calculated cosine similarity is the degree of topic matching between the text to be classified and each topic unit. For example, a high cosine similarity between the semantic vector of the text to be classified and the topic feature vector of the topic unit "UAV Navigation Technology" indicates a high degree of topic matching between the text to be classified and the topic unit.
[0123] Step S153: sorting the subject matching degrees in descending order, and selecting a preset number of subject units with the highest ranking as candidate subject units.
[0124] After obtaining the thematic match between the document to be classified and each topic unit, these thematic matches are sorted from highest to lowest. After the sorting is complete, a preset number of thematic units with the highest rankings are selected as candidate topic units. The preset number is set based on actual application requirements and the total number of topic units. Its purpose is to screen out the topic units that are most likely to be relevant to the document to be classified, reducing the processing load for subsequent classification decisions.
[0125] For example, the preset number is set to a number, and then a number of topic units with the highest topic matching degree are selected as candidate topic units. These candidate topic units are the topics to which the text to be classified may belong.
[0126] Step S154: input the semantic vector of the text to be classified into the topic classification rule set, calculate the rule matching score between the semantic vector of the text to be classified and each topic unit through the feature weight matrix, compare the rule matching score with the classification decision threshold vector, and determine the preliminary judgment result of whether the text to be classified belongs to the corresponding topic unit.
[0127] The semantic vector of the text to be classified is input into the topic classification rule set. The vector is processed using the feature weight matrix contained therein, and the rule matching score between the semantic vector of the text to be classified and each topic unit is calculated. The weight parameters in the feature weight matrix reflect the importance of different semantic features in the classification decision. The rule matching score obtained through weighted calculation can reflect the degree of match between the text to be classified and each topic unit at the classification rule level.
[0128] After calculating the rule matching score, it is compared with the corresponding threshold in the classification decision threshold vector. If the rule matching score exceeds the corresponding classification decision threshold, the document to be classified is preliminarily determined to belong to the topic unit; otherwise, it is preliminarily determined not to belong to the topic unit. Through this process, a preliminary determination result is obtained as to whether the document to be classified belongs to each topic unit.
[0129] Step S155: Based on the subject matching degree and the preliminary determination result, a comprehensive score is given to the candidate subject unit, where the comprehensive score is calculated by a weighted sum of the subject matching degree and the preliminary determination result.
[0130] The candidate thematic units are comprehensively scored based on the theme matching and preliminary judgment results. The comprehensive score is calculated by taking the weighted sum of the theme matching and preliminary judgment results, where the weights of the theme matching and preliminary judgment results are set according to their importance in the classification decision.
[0131] If the preliminary judgment result is "yes," the corresponding value is higher; if the preliminary judgment result is "no," the corresponding value is lower. For example, the weight of the topic match is set to a certain ratio, and the weight of the preliminary judgment result is set to another ratio. The comprehensive score is the topic match multiplied by its weight plus the value corresponding to the preliminary judgment result multiplied by its weight. This comprehensive score can more comprehensively evaluate the degree of match between the candidate topic unit and the text to be classified.
[0132] Step S156: Select the subject unit with the highest comprehensive score as the main classification subject of the low-altitude economic teaching text to be classified, and check whether there are other candidate subject units with comprehensive scores exceeding the secondary subject threshold. If so, use them as secondary classification subjects.
[0133] After obtaining the comprehensive scores of the candidate subject units, the subject unit with the highest comprehensive score is selected as the main classification subject of the low-altitude economic teaching text to be classified. The main classification subject is the subject unit to which the text to be classified is most likely to belong and can reflect the core content of the text.
[0134] At the same time, the overall score of other candidate topic units is checked to see if it exceeds the secondary topic threshold. The secondary topic threshold is used to determine whether a candidate topic unit should be considered a secondary classification topic. If a candidate topic unit has an overall score exceeding this threshold, it will be considered a secondary classification topic. Secondary classification topics can reflect other important thematic content contained in the document to be classified.
[0135] Step S157: combining the main classification topics and the secondary classification topics to form a topic classification result corresponding to the low-altitude economics teaching text to be classified, wherein the topic classification result includes a topic unit identifier, core teaching terms and a matching confidence of the corresponding topic.
[0136] The primary and secondary classification themes are combined to form the thematic classification results for the low-altitude economics teaching texts to be classified. The thematic classification results include the identifier of each thematic unit, which uniquely identifies the thematic unit; core teaching terms that can intuitively reflect the core content of the thematic unit; and the matching confidence of the corresponding thematic unit. The higher the confidence value, the greater the probability that the text to be classified belongs to the thematic unit.
[0137] For example, the main classification topic of the text to be classified is "UAV flight control" and the secondary classification topic is "UAV sensor technology". The classification results will include the identifiers of these two topic units, the core teaching terms "UAV flight control" and "UAV sensor" and their respective matching confidence levels.
[0138] Step S158: collecting the subject classification result set generated by the low-altitude economics teaching text subject classification model and the corresponding low-altitude economics teaching texts to be classified.
[0139] In order to continuously optimize the low-altitude economics teaching text topic classification model, it is necessary to collect the topic classification results generated by the model and the corresponding low-altitude economics teaching texts to be classified. The collection includes all the texts to be classified that have been classified by the model and their corresponding topic classification results.
[0140] During the collection process, data can be organized and stored, and a dedicated database can be established for management to ensure data integrity and traceability. At the same time, information such as the generation time and model version of each classification result can be recorded to facilitate subsequent analysis of model performance changes.
[0141] Step S159: manually review the subject classification result set, mark misclassified samples, and form a misclassified sample set.
[0142] The collected subject classification results are manually reviewed by a panel of researchers with expertise and teaching experience in the field of low-altitude economics. During the review process, the subject classification results are compared with the actual content of the text to be classified to determine whether the classification results are accurate.
[0143] If the classification result is found to be inconsistent with the actual content of the text, it is considered a classification error and the corresponding sample is marked. All samples marked as misclassified are collected to form a set of misclassified samples. The set of misclassified samples can reflect problems in the model's classification process and provide specific directions for model optimization.
[0144] Step S1510: re-processing the texts in the misclassified sample set with term extraction and association analysis to update the low-altitude economics teaching term association network.
[0145] For the text in the misclassified sample set, term extraction and association analysis are performed again. The processing process is the same as steps S111 to S115, including word segmentation, stop word filtering, term identification, co-occurrence frequency statistics, and co-occurrence association strength calculation.
[0146] Based on the reprocessed teaching terms and their co-occurrence strengths, the original low-altitude economics teaching terminology network was updated. This update included adding new teaching term nodes, adjusting edge weights (co-occurrence strengths) between nodes, and deleting no longer applicable term nodes or edges. This update enabled the network to more accurately reflect the relationships between low-altitude economics teaching terms.
[0147] Step S1511: Based on the updated low-altitude economics teaching terminology association network, re-execute the topic structure mining process and the topic semantic consistency optimization process to generate an updated optimized low-altitude economics teaching topic set.
[0148] On the basis of the updated low-altitude economics teaching terminology association network, the subject structure mining process is re-executed, that is, the processing process from step S121 to step S125, including calculating the network structure characteristic parameters, determining the core candidate terms of the subject, performing community division, matching the core terms and forming the initial subject term set, and merging semantically similar sets, etc., to obtain a new initial low-altitude economics teaching subject set.
[0149] Then, the original text semantic vector set or the regenerated text semantic vector set is used to optimize the semantic consistency of the new initial topic set, delete the topic units that do not meet the requirements, merge the semantically similar topic units, etc., and generate an updated optimized low-altitude economic teaching topic set.
[0150] Step S1512: Reconstruct the topic feature library and the topic classification rule set using the updated optimized low-altitude economics teaching topic set to achieve iterative optimization of the low-altitude economics teaching text topic classification model.
[0151] According to the updated optimized low-altitude economic teaching theme set, the theme feature library is reconstructed, that is, according to the process from step S151 to step S152, the theme feature vector of each theme unit is generated and combined to form a new theme feature library.
[0152] At the same time, a new set of low-altitude economics teaching text samples with topic annotations is collected, or the original sample set is combined with the misclassified sample set to retrain the topic classification rule set, i.e., the process from step S153 to step S155. The reconstructed topic feature library and topic classification rule set are integrated into the low-altitude economics teaching text topic classification model to achieve iterative optimization of the model.
[0153] Through the above iterative optimization process, the model can continuously improve the classification performance, improve the accuracy of low-altitude economic teaching text classification, and better adapt to different classification needs.
[0154] Figure 2A schematic diagram illustrates exemplary hardware and software components of a topic-model-based low-altitude economics teaching text classification system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, the processor 120 can be used in the topic-model-based low-altitude economics teaching text classification system 100 and be used to perform the functions of the present application.
[0155] The low-altitude economics teaching text classification system 100 based on the topic model can be a general-purpose server or a special-purpose server, both of which can be used to implement the low-altitude economics teaching text classification method based on the topic model of the present application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0156] The low altitude economics teaching text classification system 100 of the subject model can comprise the network port 110 that is connected to the network, one or more processors 120 that are used to carry out program instructions, communication bus 130 and storage medium 140 of different forms, for example, disk, ROM, or RAM, or its arbitrary combination. Exemplarily, the low altitude economics teaching text classification system 100 of the subject model can also comprise the program instruction in the non-transitory storage medium or its arbitrary combination that is stored in ROM, RAM or other types.Can realize the method for the present application according to these program instructions.The low altitude economics teaching text classification system 100 of the subject model also comprises the I / O interface 150 between computer and other input and output equipment based on the subject model.
[0157] For ease of explanation, only one processor is described in the low-altitude economics teaching text classification system 100 based on the topic model. However, it should be noted that the low-altitude economics teaching text classification system 100 based on the topic model in the present application can also include multiple processors, so the steps performed by a processor described in the present application can also be performed jointly or individually by multiple processors. For example, if the processor of the low-altitude economics teaching text classification system 100 based on the topic model executes step A and step B, it should be understood that step A and step B can also be performed jointly by two different processors or performed individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.
[0158] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned low-altitude economic teaching text classification method based on the topic model is implemented.
[0159] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A low-altitude economics teaching text classification method based on topic model, characterized by: The method comprises: Obtaining a low-altitude economics teaching text collection, performing term extraction and association analysis on the low-altitude economics teaching text collection, and constructing a low-altitude economics teaching term association network, wherein the nodes of the low-altitude economics teaching term association network represent teaching terms extracted from the low-altitude economics teaching text collection, and the edges of the low-altitude economics teaching term association network represent the co-occurrence association strength between different teaching terms; Performing topic structure mining based on the low-altitude economics teaching terminology association network to generate an initial low-altitude economics teaching topic set, wherein the initial low-altitude economics teaching topic set includes multiple topic units with term aggregation characteristics, and each topic unit is composed of a teaching term set with a close co-occurrence correlation strength in the low-altitude economics teaching terminology association network; Performing semantic vector representation processing on text units in the low-altitude economics teaching text set to obtain a text semantic vector set, and performing topic semantic consistency optimization processing on the initial low-altitude economics teaching topic set using the text semantic vector set to generate an optimized low-altitude economics teaching topic set; Constructing a low-altitude economics teaching text topic classification model based on the optimized low-altitude economics teaching topic set, wherein the low-altitude economics teaching text topic classification model includes a topic feature library for topic matching and a topic classification rule set for classification decision-making; The low-altitude economics teaching text to be classified is input into the low-altitude economics teaching text subject classification model, subject feature matching processing is performed through the subject feature library, and classification decision processing is performed in combination with the subject classification rule set to generate a subject classification result corresponding to the low-altitude economics teaching text to be classified.
2. The low-altitude economic teaching text classification method based on the topic model according to claim 1 is characterized in that: The step of obtaining a low-altitude economics teaching text set, performing term extraction and association analysis on the low-altitude economics teaching text set, and constructing a low-altitude economics teaching term association network includes: Performing word segmentation processing on each text unit in the low-altitude economics teaching text set to obtain a text word segmentation sequence set; Performing stop word filtering on the text segmentation sequence set to remove vocabulary units without actual teaching terminology meaning and retaining candidate terminology units containing teaching professional meaning; Performing term recognition processing on the candidate term units to extract teaching terms with teaching characteristics in the field of low-altitude economy, wherein the teaching terms include noun terms representing the concept of low-altitude economy, verb terms representing operation procedures, and adjective terms representing attribute characteristics; Counting the co-occurrence frequencies of the teaching terms in the low-altitude economics teaching text set, and calculating the co-occurrence correlation strength between different teaching terms based on the co-occurrence frequencies, wherein the co-occurrence correlation strength is positively correlated with the co-occurrence frequencies; The low-altitude economics teaching terminology association network is constructed with the teaching terms as nodes and the co-occurrence association strength as edge weights. The node attributes of the low-altitude economics teaching terminology association network include the domain attribution labels of the teaching terms and the term occurrence frequencies.
3. The low-altitude economic teaching text classification method based on the topic model according to claim 1 is characterized in that, The subject structure mining process is performed based on the low-altitude economics teaching term association network to generate an initial low-altitude economics teaching subject set, including: The network structure characteristics of the low-altitude economic teaching terminology association network are calculated to obtain the degree centrality, betweenness centrality and closeness centrality parameters of each node. The degree centrality represents the number of direct connections between a node and other nodes, the betweenness centrality represents the frequency of a node as an intermediary in the shortest path between other nodes, and the closeness centrality represents the average shortest path length from a node to all other nodes. Calculate a comprehensive node importance index based on the degree centrality, betweenness centrality, and closeness centrality parameters, sort the nodes in the low-altitude economics teaching term association network in descending order according to the comprehensive node importance index, and select the top-ranked nodes as the topic core candidate terms; A community discovery algorithm is used to perform community division processing on the low-altitude economics teaching terminology association network, and teaching terms with significant co-occurrence association strength are clustered to form multiple terminology community units, each of which corresponds to a potential topic; Matching the core candidate terms of the subject with the term community units, determining the core teaching terms of each term community unit, and taking the core teaching terms as the center, selecting the teaching terms in the term community units whose co-occurrence correlation strength with the core teaching terms exceeds a preset threshold to form an initial subject term set; The initial subject term set is subjected to a term semantic similarity analysis, and initial subject term sets with similar semantics are merged to generate an initial low-altitude economic teaching subject set containing multiple subject units, each subject unit containing a core teaching term and multiple related teaching terms.
4. The low-altitude economic teaching text classification method based on the topic model according to claim 1 is characterized in that: The processing of semantic vector representation of text units in the low-altitude economics teaching text set to obtain a text semantic vector set includes: Performing sentence segmentation processing on the text units in the low-altitude economics teaching text set, dividing each text unit into multiple sentence units; Perform word-level semantic encoding on each sentence unit, converting each word in the sentence unit into a word semantic vector generated by a pre-trained language model; Performing sequence splicing processing on the word semantic vectors in the sentence unit to generate a sentence semantic vector sequence, wherein the length of the sentence semantic vector sequence is consistent with the number of words in the sentence unit; An attention mechanism is used to perform weighted aggregation processing on the sentence semantic vector sequence to generate a sentence-level semantic vector, wherein the attention mechanism assigns different weight parameters according to the semantic contribution of words in the sentence unit; All sentence-level semantic vectors in the text unit are average-pooled to generate a text-unit-level text semantic vector, and the text semantic vectors of all text units are combined to form the text semantic vector set.
5. The low-altitude economic teaching text classification method based on the topic model according to claim 1 is characterized in that: The method of performing topic semantic consistency optimization processing on the initial low-altitude economics teaching topic set by using the text semantic vector set to generate an optimized low-altitude economics teaching topic set includes: Performing topic semantic vector representation processing on each topic unit in the initial low-altitude economics teaching topic set, averaging and fusing the semantic vectors of the core teaching terms and related teaching terms in the topic unit to generate a topic unit semantic vector; Calculating the cosine similarity between each text semantic vector and each topic unit semantic vector in the text semantic vector set to obtain a text topic association matrix, wherein the rows of the text topic association matrix represent text units, the columns represent topic units, and the matrix elements represent the semantic association strength between the text units and the topic units; Based on the text topic association matrix, the number of text units associated with each topic unit and the average semantic association strength are counted, and topic units whose number of associated text units is lower than a preset number threshold or whose average semantic association strength is lower than a preset strength threshold are deleted; Calculate the topic similarity of the remaining topic units, merge the topic units whose cosine similarity between the semantic vectors exceeds the preset merging threshold, and retain the core teaching terms with the highest semantic contribution in the merged topic units; The subject term set is reconstructed based on the merged subject units, and the teaching terms in the subject term set are semantically sorted to generate an optimized low-altitude economic teaching subject set with clear semantic boundaries.
6. The low-altitude economic teaching text classification method based on the topic model according to claim 1 is characterized in that: The method of constructing a low-altitude economics teaching text topic classification model based on the optimized low-altitude economics teaching topic set includes: Performing feature extraction processing on each topic unit in the optimized low-altitude economics teaching topic set, converting the teaching terms contained in the topic unit into a topic feature vector, wherein the dimension of the topic feature vector is consistent with the dimension of the text semantic vector; Combining the subject feature vectors of all subject units to form a subject feature library, wherein the subject feature library includes identification information, core teaching terms and corresponding subject feature vectors of each subject unit; Collecting a set of low-altitude economics teaching text samples with subject annotations, wherein each sample in the low-altitude economics teaching text sample set includes text content and a corresponding subject unit label; Performing text semantic vector representation processing on sample texts in the low-altitude economics teaching text sample set to obtain a sample semantic vector set; Based on the sample semantic vector set and the corresponding topic unit labels, a supervised learning algorithm is used to train a topic classification rule set, wherein the topic classification rule set includes a mapping relationship between the sample semantic vector and the topic unit label and a classification decision threshold; The subject feature library and the subject classification rule set are integrated into the low-altitude economics teaching text subject classification model, and the low-altitude economics teaching text subject classification model further comprises a model input interface and a classification result output interface.
7. The low-altitude economic teaching text classification method based on the topic model according to claim 6 is characterized in that: The method of training a topic classification rule set based on the sample semantic vector set and the corresponding topic unit labels using a supervised learning algorithm includes: Dividing the sample semantic vector set and the corresponding topic unit labels into a training sample subset and a verification sample subset; Initialize the parameters of the topic classification rule set, including the feature weight matrix and the classification decision threshold vector; Inputting the sample semantic vectors in the training sample subset into the classification rule training model, and calculating the matching score between the sample semantic vectors and each topic unit through the feature weight matrix; Comparing the matching score with a classification decision threshold vector to generate a predicted topic unit label for the sample; Calculating a classification loss value between the predicted topic unit label and the actual topic unit label in the training sample subset, wherein the classification loss value is calculated using a cross entropy loss function; Based on the classification loss value, a back propagation algorithm is used to update the parameters of the feature weight matrix and the classification decision threshold vector; Verify the updated topic classification rule set using the verification sample subset, and calculate the classification accuracy and F1 score indicators; When the classification accuracy and F1 score indicators reach the preset convergence conditions, the training is stopped and the current feature weight matrix and classification decision threshold vector are saved as the final topic classification rule set.
8. The low-altitude economic teaching text classification method based on the topic model according to claim 1 is characterized in that: The low-altitude economics teaching text to be classified is input into the low-altitude economics teaching text subject classification model, subject feature matching processing is performed through the subject feature library, and classification decision processing is performed in combination with the subject classification rule set to generate a subject classification result corresponding to the low-altitude economics teaching text to be classified, including: Performing text preprocessing on the low-altitude economics teaching text to be classified, performing semantic vector representation processing on the preprocessed low-altitude economics teaching text to be classified, and generating a semantic vector of the text to be classified; Reading the topic feature vectors of all topic units from the topic feature library, and calculating the cosine similarity between the semantic vector of the text to be classified and each topic feature vector, to obtain the topic matching degree between the text to be classified and each topic unit; Sorting the subject matching degrees in descending order, and selecting a preset number of subject units with the highest ranking as candidate subject units; Inputting the semantic vector of the text to be classified into the topic classification rule set, calculating the rule matching score between the semantic vector of the text to be classified and each topic unit through the feature weight matrix, comparing the rule matching score with the classification decision threshold vector, and determining a preliminary determination result of whether the text to be classified belongs to the corresponding topic unit; Combining the subject matching degree and the preliminary determination result, a comprehensive score is given to the candidate subject unit, wherein the comprehensive score is calculated by a weighted sum of the subject matching degree and the preliminary determination result; The subject unit with the highest comprehensive score is selected as the main classification subject of the low-altitude economic teaching text to be classified, and it is checked whether there are other candidate subject units with comprehensive scores exceeding the secondary subject threshold. If there are other candidate subject units, they are selected as secondary classification subjects. The main classification topics and the secondary classification topics are combined to form a topic classification result corresponding to the low-altitude economic teaching text to be classified. The topic classification result includes a topic unit identifier, core teaching terms and a matching confidence of the corresponding topic.
9. The low-altitude economic teaching text classification method based on the topic model according to claim 1 is characterized in that: The method further comprises: Collect the subject classification result set generated by the low-altitude economics teaching text subject classification model and the corresponding low-altitude economics teaching text to be classified; Manually review the subject classification result set, mark misclassified samples, and form a misclassified sample set; Re-processing the texts in the misclassified sample set with term extraction and association analysis to update the low-altitude economics teaching term association network; Based on the updated low-altitude economics teaching terminology association network, the topic structure mining process and topic semantic consistency optimization process are re-executed to generate an updated and optimized low-altitude economics teaching topic set; The updated optimized low-altitude economics teaching topic set is used to reconstruct a topic feature library and a topic classification rule set, thereby achieving iterative optimization of the low-altitude economics teaching text topic classification model.
10. A low-altitude economic teaching text classification system based on topic model, characterized by: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the low-altitude economic teaching text classification method based on the topic model described in any one of claims 1 to 9.