Label determination method and system for campus platform content
By constructing a multi-dimensional tagging system and a multi-modal Transformer model, the problem of detailed classification on campus content platforms was solved, achieving deep integration and accurate classification of content and tags, and improving the correlation between tags and content.
Patent Information
- Application Number
- CN202511631034.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-01-09
AI Technical Summary
Existing campus content platforms cannot perform detailed categorization, resulting in poor correlation between tags and content, and failing to meet the diverse needs of users.
By acquiring text, image, and video features from campus content platforms, a multi-dimensional tagging system is constructed, encompassing student basic information, school scenarios, interactive behaviors, and trending tags. The multi-modal Transformer model is then used for feature fusion and tag annotation, achieving deep integration of multi-dimensional tags and multi-modal data.
It improves the accuracy of content tag classification, enhances the correlation between content and tags, and supports structured storage and accurate classification of multi-dimensional tags and multi-modal data.
Smart Images

Figure CN121301933A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of platform content data processing, and in particular to a label determination method and system for campus platform content. BACKGROUND
[0002] With the advancement of digital campus construction, the content creation of the university student group presents the characteristics of "multi-modal, scene fragmentation, and strong attribute correlation". The campus content has long exceeded the single text form and presents a situation of deep integration of multiple modalities such as text, image, and video. Specifically, students conduct in-depth academic discussions and publish club notices through text, display experimental data charts and campus activity highlights through images, and share course learning experiences and graduation design demonstration processes through videos. These multi-modal content collectively constitutes an ecology of campus knowledge sharing and cultural exchange, and the content is deeply bound to the professional attributes, campus scenes, and behavior characteristics of the university students (such as computer professional students discussing "deep learning framework" and literature society members sharing "poetry creation"). The existing content in the campus content platform is classified using general tags such as entertainment and technology, which cannot be classified in detail, resulting in low classification accuracy, poor relevance of tags and content, and failure to meet the diversified needs of users. SUMMARY
[0003] To solve the above problems in the prior art, the present application provides a label determination method and system for campus platform content. The technical problem to be solved by the present application is solved by the following technical scheme: The first aspect of the embodiment of the present application provides a label determination method for campus platform content, comprising the following steps: obtaining at least one corresponding text feature, image feature, and video feature of text information, image, and video of each existing content of a campus content platform; matching the text feature, the image feature, and the video feature to corresponding content labels to generate an association database; wherein the content labels include at least one of student basic information labels and weights, school scene labels and weights, interactive behavior labels and weights, and hot topic labels and weights; performing feature enhancement and / or label correction on the text feature, the image feature, and the video feature and the corresponding content labels to obtain enhanced text features, enhanced image features, enhanced video features, and corresponding enhanced content labels; performing feature fusion on the enhanced text features, the enhanced image features, and the enhanced video features to obtain fused features; The fusion feature and the enhanced content label are input into a multi-modal Transformer model for training to obtain a label annotation model; wherein a self-attention sublayer of an encoder of the multi-modal Transformer model performs vector fusion according to the fusion feature and the enhanced content label to obtain a fusion vector, and in the training process, a output layer outputs a label matching result and a content review result based on the fusion vector; At least one of text features, image features and video features corresponding to text information, images and videos of the current edited content of the campus content platform is acquired, and the acquired features are fused and input into the label annotation model to output a current label matching result and a current content review result.
[0004] In an embodiment of the present application, the student basic information label includes a plurality of course majors, a plurality of grades and a plurality of education types; wherein the weight of the course major is determined according to the number of times of occurrence of the course major term in the text features, the image features and the video features; The school scene label includes laboratory operation, club activity, postgraduate review, course learning, graduation design, social practice, daily life and leisure entertainment; wherein the weight of the school scene label is determined according to the scene target, the course major element target, the experimental operation target and the corresponding scene position in the image features and the video features; The interactive behavior label includes original content, forwarding and commenting, asking for help, controversial discussion and academic sharing; The hot spot label includes a campus hot spot event and a social hot spot event; wherein the weight of the hot spot label is determined according to the degree of association with the campus hot spot event or the social hot spot event.
[0005] In an embodiment of the present application, the calculation formula of the fusion vector is: wherein, represents the fusion vector, represents the relevance of the enhanced content label and one of the enhanced text features, the enhanced image features and the enhanced video features in the fusion features, represents one of the enhanced text features, the enhanced image features and the enhanced video features in the fusion features, represents the weight of the label embedding vector of the enhanced content label, represents the label embedding vector of the enhanced content label.
[0006] In an embodiment of the present application, the feature enhancement and / or label correction of the text features, the image features and the video features and the corresponding content labels to obtain enhanced text features, enhanced image features, enhanced video features and corresponding enhanced content labels comprises: According to the weight of the professional terms in the text features and the weight of the corresponding student basic information label multiplied by a weight coefficient greater than 1 according to the professional term dictionary, an enhanced content label and weight are obtained; According to the context semantic analysis of the text features according to the professional term dictionary, the corresponding content label is replaced to obtain an enhanced content label; The region in the image features associated with the content label uses a super-resolution reconstruction feature to obtain enhanced image features; According to the content label, the key frames of the corresponding video features are extracted; The optical flow method is used to extract motion features from the key frames, and then the corresponding content label is matched to generate enhanced video features and enhanced content labels.
[0007] The second aspect of the embodiment of the present application provides a label determination system for a campus platform content, comprising: The acquisition module is used to acquire the text information, images and videos of each existing content of the campus content platform, and at least one corresponding text feature, image feature and video feature; The matching module is used to match the text features, the image features and the video features to the corresponding content labels to generate an association database; wherein the content labels comprise at least one of a student basic information label and weight, a school scene label and weight, an interactive behavior label and weight, and a hot spot label and weight; The enhancement module is used to perform feature enhancement and / or label correction on the text features, the image features and the video features and the corresponding content labels to obtain enhanced text features, enhanced image features, enhanced video features and corresponding enhanced content labels; The fusion module is used to perform feature fusion on the enhanced text features, the enhanced image features and the enhanced video features to obtain fusion features; The training module is used to input the fusion features and the enhanced content labels into a multi-modal Transformer model for training to obtain a label annotation model; wherein the self-attention sublayer of the encoder of the multi-modal Transformer model performs vector fusion according to the fusion features and the enhanced content labels to obtain a fusion vector, and in the training process, the output layer outputs a label matching result and a content review result based on the fusion vector; The labeling module is configured to acquire at least one of text features, image features and video features in current edited content of a campus content platform, fuse the acquired features, input the fused features into a label labeling model, and output a current label matching result and a current content review result.
[0008] In an embodiment of the present application, the student basic information label includes a plurality of course majors, a plurality of grades and a plurality of educational background types; wherein the weight of the course major is determined according to the number of times the course major term appears in the text features, the image features and the video features; The school scene label includes laboratory operation, club activity, postgraduate review, course learning, graduation design, social practice, daily life and leisure entertainment; wherein the weight of the school scene label is determined according to the scene target, the course major element target, the experimental operation target and the corresponding scene position in the image features and the video features; The interactive behavior label includes original content, forwarding and commenting, asking for help, controversial discussion and academic sharing; The hot topic label includes a campus hot topic event and a social hot topic event; wherein the weight of the hot topic label is determined according to the degree of association with the campus hot topic event or the social hot topic event.
[0009] In an embodiment of the present application, the calculation formula of the fusion vector is: wherein, represents the fusion vector, represents the relevance of the enhanced content label to one of the enhanced text features, the enhanced image features and the enhanced video features in the fusion features, represents one of the enhanced text features, the enhanced image features and the enhanced video features in the fusion features, represents the weight of the label embedding vector of the enhanced content label, represents the label embedding vector of the enhanced content label.
[0010] In an embodiment of the present application, the feature enhancement and / or label correction of the text features, the image features and the video features and the corresponding content labels to obtain enhanced text features, enhanced image features, enhanced video features and corresponding enhanced content labels include: According to the weight of the professional term in the text features and the weight of the corresponding student basic information label multiplied by a weight coefficient greater than 1 according to the professional term dictionary, the enhanced content label and the weight are obtained; According to the context semantic analysis of the text features according to the professional term dictionary, the corresponding content label is replaced to obtain the enhanced content label. The region associated with the content label in the image feature adopts super-resolution reconstruction features to obtain enhanced image features; According to the content label, a corresponding key frame of the video feature is extracted; Motion features are extracted from the key frame by using an optical flow method, and then corresponding content labels are matched to generate enhanced video features and enhanced content labels.
[0011] The third aspect of the embodiment of the application provides an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to realize the method for determining the label of the campus platform content provided in the first aspect of the embodiment of the application.
[0012] The fourth aspect of the embodiment of the application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the method for determining the label of the campus platform content provided in the first aspect of the embodiment of the application.
[0013] The beneficial effects of the application are as follows: The application covers the full attribute dimension of the college student content by constructing the multi-dimensional label of the student basic information label, the school scene label, the interactive behavior label and the hot spot label, automatically binds the multi-modal features of the text features, the image features and the video features with the multi-dimensional label, realizes the structured storage of the multi-modal content and the multi-dimensional label of the college student, enhances the relevance of the content and the label, further constructs the label annotation model based on the multi-modal Transformer, realizes the deep fusion of the multi-dimensional label and the multi-modal data, outputs the accurate content classification result, and improves the accuracy of the label classification of the content.
[0014] Other features and advantages of the application will be described in the following description, and some will become apparent from the description, or will be understood from the practice of the application. The purpose and other advantages of the application can be achieved and obtained by the structure specially pointed out in the written description, claims and drawings.
[0015] The technical solutions of the application will be further described in detail below with the help of the drawings and embodiments. DETAILED DESCRIPTION
[0016] The accompanying drawings are used to provide a further understanding of the application, and constitute a part of the specification, and are used to explain the application together with the embodiments of the application, and do not constitute a limitation on the application. In the drawings: Figure 1 A flowchart of the method for determining the label of the campus platform content provided by the embodiment of the application is shown in the figure. Figure 2A block diagram of a campus platform content label determination system provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0017] The present application will be further described in detail below in conjunction with specific embodiments, but the embodiments of the present application are not limited thereto.
[0018] As shown in Figure 1 the first aspect of the embodiment of the present application provides a campus platform content label determination method, comprising the following steps: Step 11, obtaining at least one corresponding text feature, image feature and video feature of the text information, image and video of each existing content of the campus content platform.
[0019] Step 12, matching the text feature, image feature and video feature with the corresponding content label to generate an association database.
[0020] Among them, the content label includes at least one of student basic information label and weight, school scene label and weight, interactive behavior label and weight, and hot spot label and weight.
[0021] Step 13, performing feature enhancement and / or label correction on the text feature, image feature and video feature and the corresponding content label to obtain enhanced text feature, enhanced image feature, enhanced video feature and corresponding enhanced content label.
[0022] Step 14, performing feature fusion on the enhanced text feature, enhanced image feature and enhanced video feature to obtain fused features.
[0023] Step 15, inputting the fused features and enhanced content label into a multi-modal Transformer model for training to obtain a label annotation model.
[0024] Among them, the self-attention sublayer of the encoder of the multi-modal Transformer model performs vector fusion according to the fused features and the enhanced content label to obtain a fused vector, and in the training process, the output layer outputs label matching results and content review results based on the fused vector.
[0025] Step 16, obtaining at least one corresponding text feature, image feature and video feature of the text information, image and video of the current edited content of the campus content platform, and inputting the obtained features into the label annotation model after fusion to output current label matching results and current content review results.
[0026] In this embodiment, the full attribute dimension of the college student content is covered by constructing the multi-dimensional tags of the student basic information tag, the school scene tag, the interactive behavior tag and the hot tag, the multi-modal features of the text features, the image features and the video features are automatically bound based on the multi-dimensional tags, the structured storage of the multi-modal content of the college student and the multi-dimensional tags is realized, the relevance of the content and the tags is enhanced, the label labeling model is further constructed based on the multi-modal Transformer, the deep fusion of the multi-dimensional tags and the multi-modal data is realized, the accurate content classification result is output, and the accuracy of the label classification of the content is improved.
[0027] On the basis of the first aspect of the embodiment of the present application, the second aspect of the embodiment of the present application further details a label determination method of a campus platform content, and the second aspect of the embodiment of the present application provides a label determination method of a campus platform content, including the following steps: Step 21, constructing four-layer three-order dynamic content tags.
[0028] The content tags include the student basic information tag, the school scene tag, the interactive behavior tag and the hot tag. The student basic information tag includes a plurality of course majors, a plurality of grades and a plurality of education types; wherein the weight of the course major is determined according to the number of times of occurrence of the course major term in the text features, the image features and the video features.
[0029] The school scene tag includes laboratory operation, community activity, postgraduate review, course learning, graduation design, social practice, daily life and leisure entertainment.
[0030] The weight of the school scene tag is determined according to the scene target, the course major element target, the experimental operation target and the corresponding scene position in the image features and the video features.
[0031] The interactive behavior tag includes original content, forwarding and commenting, asking for help, controversial discussion and academic sharing.
[0032] The hot tag includes a campus hot event and a social hot event; wherein the weight of the hot tag is determined according to the degree of association with the campus hot event or the social hot event.
[0033] For example, the content tags are shown in Table 1: Table 1 Here, the weight calculation of each tag can combine multiple information (such as term frequency, scene matching degree, user interaction), and through normalization processing (such as Softmax), the sum is ensured to be 1.
[0034] The content tags described above have a three-order dynamic adjustment mechanism: I. Periodic adjustment (quarterly): According to the university student academic cycle (e.g. September "freshmen enrollment", June "graduation season"), add / delete temporary tags (e.g. add "thesis guidance" tag in graduation season, delete after graduation), adjust the weight of basic attribute layer tags (e.g. increase the weight of "course learning" tag to 0.4 in the semester, and reduce it to 0.15 in the holiday).
[0035] II. Hot spot triggered adjustment (real-time): When a campus hotspot breaks out (e.g. "campus singer contest", "subject competition registration"), capture keywords (e.g. "singer contest registration channel") through the hotspot monitoring module, add corresponding hotspot tags within 1 hour, and set the initial correlation strength to "strong".
[0036] III. Data feedback adjustment (monthly): Based on classification error cases (e.g. mislabeling "mechanical design drawing" as "architectural design"), correct the label feature mapping relationship (e.g. add "mechanical parts labeling" as a core feature of "mechanical engineering professional" label), and adjust the label weight coefficient.
[0037] Step 22, obtain at least one corresponding text feature, image feature, video feature of the text information, image and video of each existing content of the campus content platform.
[0038] In this step, collect text information, images and videos, and capture data by professional classification direction—computer professional: CSDN campus special zone, MOOC "artificial intelligence" course comment area; clinical medicine professional: medical MOOC platform "pathophysiology" discussion area, campus hospital public number message; Chinese language and literature professional: campus literature forum, "contemporary literature" course assignment submission system, ensure that the data volume of each type of professional is ≥50,000. Collect student basic information such as grade segment (from freshman to graduate three) and education type (undergraduate, graduate).
[0039] Based on the content publishing source associated with the campus geographical location (such as laboratory, library, community activity center), collect "experiment record upload area" of laboratory management system, "self-study seat reservation" message of library public number, and "activity promotion content" of community recruitment platform.
[0040] For each content, at least one of "text + image + video + metadata" is collected—such as collecting "electronic information professional experiment report", synchronously obtaining report text (including experiment steps), experiment data image (oscilloscope waveform), experiment operation video, and publisher's grade / professional (metadata).
[0041] After collecting the above information, the corresponding text features, image features, and video features in the information are extracted. The domain-adapted BERT (such as BERT-College) is used to extract text features, and the pre-training is fine-tuned on the college professional corpus (such as MOOC comments, experiment reports) to enhance the semantic understanding of professional terms. The example of the text "using the PID controller to adjust the reaction temperature" will generate a high-dimensional vector containing professional terms such as "PID controller" and "reaction temperature". The corresponding image information is extracted through ResNet50 to obtain global visual features (such as object contours, color distribution). For video information, 3D-CNN is used to extract action sequence features (such as "stirring solution" and "recording data" experimental operations).
[0042] Step 23, match the text features, image features, and video features with the corresponding content labels to generate an association database, and store the matched features and labels in a structured manner.
[0043] Among them, the content label includes at least one of student basic information label and weight, school scene label and weight, interactive behavior label and weight, and hot spot label and weight.
[0044] In this step, based on the pre-set "feature-label" mapping rule (such as text containing "PID controller" → binding "automation major" label; image containing "PCR instrument" → binding "biological engineering major + laboratory scene" label), a multi-dimensional label matrix (such as [computer major, course learning scene, original content, no hot spot association]) is automatically generated for the collected data. The multi-dimensional label matrix stores labels and weights in the form of a vector (such as "computer major" weight 0.35, "course learning scene" weight 0.28). The multi-label matrix is the core of the association between content and multi-dimensional labels, used to store the label type and its weight of the content, and reflects the matching degree of the content and the label.
[0045] The YOLOv9 model is used to load the "university student exclusive visual target library" (containing 200+ campus scene targets: such as experimental instruments, school badges, and club flags; 500+ professional related targets: such as mechanical parts and chemical reagent bottles), identify the targets in the image features and associate the corresponding labels (such as identifying "Fourier infrared spectrometer" → associating "material science major + laboratory scene" label). The "titration operation" segment in the video feature is encoded as a spatiotemporal feature vector, which is used to match the "chemical experiment" label.
[0046] The features and labels stored in the association database are structured to form a three-level index table: content ID-multi-label matrix-multi-modal feature vector, for example Content ID: { "label_matrix": {label1: weight1, label2: weight2,...}, "multimodal_features": { "text": text_embedding, "image": image_features, "video": video_features } }, The associated database features tag-driven retrieval optimization capabilities. For example, when a user searches for "computer science postgraduate entrance exam materials," the system prioritizes returning content with high tag weights (e.g., "computer science" weight 0.4 + "postgraduate entrance exam review" weight 0.35). It also has dynamic update capabilities; when "AI large-scale models" become a new hot topic, new tags are added and the index is updated in real time without reconstructing the entire database. Furthermore, it possesses cross-modal association capabilities, supporting complex queries (e.g., "videos containing chemical experiment operations and text discussing PID control"), improving accuracy through multimodal feature cross-validation. It supports fast queries by tag dimension (e.g., retrieving all images for "clinical medicine + experimental scenarios"). Multimodal feature complementarity is also supported (e.g., when text descriptions are ambiguous, image / video features are used to complete the tags).
[0047] Example of a database query statement: SELECT * FROM content_index WHERE label_matrix['Clinical Medicine Major']>0.5 AND image_scene_feature LIKE '%laboratory%' Examples of practical application scenarios: Campus forum moderation: Automatically blocks videos containing "dangerous experimental operations" (by matching safety rules based on image features).
[0048] Personalized recommendations: Push high-weight content related to "deep learning frameworks" to computer science students (based on tag matrix and matching user's professional tags).
[0049] Step 24: Enhance the text features, image features, and video features and their corresponding content tags, and / or correct the tags to obtain enhanced text features, enhanced image features, enhanced video features, and their corresponding enhanced content tags.
[0050] Step 24 includes steps 241-245: Step 241: Based on the professional terminology dictionary, the weights of the professional terms in the text features and the corresponding student basic information tags are multiplied by a weight coefficient greater than 1 to obtain the enhanced content tags and weights.
[0051] Based on the "Dictionary of College Professional Terms" (containing 13 major disciplines, 8000+ professional terms, such as "finite element analysis" "apoptosis"), when the professional terms in the text features are the same as the professional terms in the professional term dictionary, the weight of the corresponding course professional label is multiplied by 1.5 to get the new weight, avoiding the loss of label association features due to term filtering.
[0052] Step 242, according to the professional term dictionary, the text features are analyzed by context semantic analysis, and the corresponding content label is replaced to get the enhanced content label.
[0053] If the text features contain both "quantum mechanics" (physics major) and "quantum communication" (communication major) labels, through context semantic analysis (such as "discussing quantum bit transmission" in the content → tending to communication major), the original physical major + communication major label is modified to communication major, solving the label conflict caused by cross-professional terms.
[0054] Step 243, the region in the image feature associated with the content label is enhanced by super-resolution reconstruction feature to get the enhanced image feature.
[0055] In this step, the region in the image associated with the label (such as the "building structure diagram" region corresponding to the "civil engineering major" label) is enhanced by super-resolution reconstruction technology to enhance the feature clarity and improve the recognition accuracy of the subsequent model for the label.
[0056] Step 244, according to the content label, extract the key frames of the corresponding video features.
[0057] Step 245, extract motion features from key frames using optical flow method, and then match corresponding content labels to generate enhanced video features and enhanced content labels.
[0058] For video features, based on the label type, extract key frames - such as "experimental operation" label video, extract "experiment start" "reagent addition" "data record" key frames (extract 1 frame every 5 seconds, additional key operation frames are reserved); "social activities" label video, extract "opening ceremony" "interactive session" key frames. For key frame sequence, extract motion features using optical flow method, and combine with the pre-set "feature-label" mapping rule (such as "code typing action" → associated with "computer major + programming scene" label), to generate the fusion vector of "spatial-temporal features + behavior label features".
[0059] Step 25, fuse the enhanced text features, enhanced image features, and enhanced video features to get the fusion features. Concatenate or weightedly fuse the enhanced text, image, and video features to form a unified multi-modal feature vector.
[0060] Step 26, after splitting the fusion features into enhanced text features, enhanced image features, enhanced video features, and enhanced content labels, input them into a multi-modal Transformer model for training to obtain a label annotation model.
[0061] In the multi-modal Transformer model, the self-attention sublayer of the encoder performs vector fusion based on the fusion features and the enhanced content labels to obtain a fusion vector. During the training process, the output layer outputs label matching results and content review results based on the fusion vector.
[0062] In this step, the structure of the model includes three layers, as follows: Input layer: fusion feature input and multi-label embedding input (convert multi-dimensional label matrix to label embedding vector, dimension = 128).
[0063] Label-guided attention fusion layer, i.e., encoder layer (encoder layer = self-attention sublayer + feedforward neural network sublayer): with "label-modal" cross-attention mechanism, the model prioritizes attention to multi-modal features associated with labels - for example, when the label is "clinical medicine professional + experimental scene", the attention weight tilts towards "experimental procedure description in text", "experimental instrument region in image", and "operation action frame in video". The fusion formula is as follows: where, represents the fusion vector, represents the modal attention weight, i.e., the correlation calculation between the enhanced content label and one of the enhanced text features, enhanced image features, and enhanced video features in the fusion features, represents one of the enhanced text features, enhanced image features, and enhanced video features in the fusion features, i =1 enhanced text feature, 2 enhanced image feature, 3 enhanced video feature), represents the overall importance weight of the label embedding vector, which can be determined by the weight of the student basic information label and the dynamic factor weight of the enhanced content label. It acts like an "attention scaling factor" to control the strength of label information in fusion. Dynamic factors such as forwarding volume, comment volume, and like volume, dynamic factor weights are the label weights corresponding to forwarding, commenting, and liking. The specific calculation is the same as . represents the label embedding vector of the enhanced content label.
[0064] Classification and review output layer: Classification output: based on the fusion vector, output "multi-dimensional label matching results" (such as label matching degree: computer professional 92%, course learning scene 88%); Audit output: The results of the classification output are combined with the preset "label-compliance rule" library (such as "Chemical Engineering Major + Experimental Scene" content needs to be audited "whether it contains dangerous reagent operation description"; "Club Activities" label content needs to be audited "whether it contains illegal gathering information"), output "compliance / suspected violation / violation" conclusion and violation reason (such as "suspected violation: contains experimental data graph without source annotation").
[0065] The model training adopts a three-stage process of "basic training-cross label fine-tuning-adversarial training". The cross label fine-tuning strategy: for the cross scene of multiple dimensions of college students (such as "Computer Major + Exam Review + Hotspot Association (Exam Enrollment Season)"), a "label cross sample set" (sample size of each cross scene ≥2000) is constructed, and the attention layer of the model is fine-tuned, to improve the classification accuracy of cross label scene. The adversarial sample is a "label confusion sample" (such as "mechanical parts diagram" labeled as "building structure diagram").
[0066] In step 27, at least one of the text features, image features, and video features corresponding to the text information, images, and videos of the current edited content of the campus content platform is obtained, and the obtained features are fused and input into the label annotation model to output the current label matching result and the current content audit result.
[0067] In a feasible implementation, the model can be fine-tuned during its use, as follows: Incremental fine-tuning mechanism: when a new label is added to the label system (such as "AI large model course discussion"), only the label embedding layer and the attention layer of the model are fine-tuned, without the need for full retraining, and the fine-tuning time is shortened to 1 / 5 of the full training time.
[0068] In a feasible implementation, the "label system-database-model" linkage update is realized to adapt to the changes in the attributes of college students. Specifically, the trigger conditions for label addition and update are divided into two categories: Hotspot trigger: the discussion volume of a campus hotspot topic breaks through 10,000 in 1 hour (such as "University Student Innovation and Entrepreneurship Competition Enrollment"), and the corresponding hotspot label is automatically added; Data feedback trigger: the classification error rate of a certain unannotated content is greater than or equal to 15% for three consecutive times (such as "New Energy Major Course Discussion" being mislabeled as "Mechanical Major"), and after manual audit, the "New Energy Major" label is added.
[0069] During the use of the model and the database, the weight of the content label is updated according to real-time dynamic factors (such as forwarding volume, comment volume, and like volume, etc.): the label weight information update is calculated by using "user interaction + content frequency" double factors, and the formula is as follows: wherein: represents the new weight of a certain content label after updating, represents the original weight of the content label, is the number of contents containing the label, is the total number of contents, represents the original weight of the retweet comment label, is the user interaction amount (likes + comments) of the content containing the label, is the total interaction amount.
[0070] Example of updating weight calculation: A "Computer Professional Postgraduate Review" content contains "Postgraduate Mathematics" and "Data Structure" terms (frequency 12 times, total professional term frequency 5000), interaction amount 280 times (total interaction amount 5000 times), then the "Computer Professional" label updating weight = 0.6 x (12 / 5000) + 0.4 x (280 / 5000) = 0.0288 + 0.0224 = 0.0512, the "Postgraduate Review scene" updating weight = 0.5 x (8 scene description word frequency / 3000 scene description word total frequency) + 0.4 x (280 / 5000) = 0.0013 + 0.0224 = 0.024.
[0071] Model iteration update during use: I. Select contents with new / updated labels in the database within the last 30 days (sample size ≥3000) to construct an incremental training set, and simultaneously load historical "label-feature" mapping knowledge (retained through knowledge distillation).
[0072] II. After each model iteration, evaluate the "new label classification accuracy" and "cross-label matching degree" indicators. If the new label classification accuracy is less than 85%, supplement samples for fine-tuning again to ensure that the model's recognition accuracy of new labels meets the standard after iteration.
[0073] Model iteration only fine-tunes the label embedding layer and attention layer, with high iteration efficiency.
[0074] In a feasible implementation, an evaluation system centered on "multi-dimensional labels" is constructed, and closed-loop optimization is realized in combination with user feedback: The core evaluation indicators (new label related indicators) are shown in Table 2: Table 2 College student user feedback: Embed a "label feedback portal" on the campus platform, and college students can score (1-5) the label matching results of contents or provide correction suggestions (such as "the content should be 'Electronic Information Major' instead of 'Automation Major'"), and the feedback data is automatically synchronized to the "label correction sample library".
[0075] Auditors feedback: Develop an audit assistance tool, auditors can mark "tag-guided audit error" cases (such as "missing 'bio-experiment' tag matching 'hazardous operation audit rule'"), the system automatically updates the "tag-compliance rule" mapping library, and optimizes the audit logic.
[0076] The method of the embodiment significantly improves the accuracy of tag matching. Through the fusion model of the multi-dimensional tag system of college students and tag guidance, the multi-tag matching accuracy rate reaches 91.5%, which is 34.2% higher than that of the general multi-tag model (68.2%); the cross-tag recognition rate reaches 89.3%, solving the identification problem of cross-professional and cross-scene tags.
[0077] The audit efficiency and accuracy are optimized, the matching rate of the "tag-compliance rule" reaches 93.1%, and the artificial review amount of auditors is reduced by 60%; the violation identification rate of the professional-related content of college students (such as "experimental data falsification" and "academic misconduct discussion") reaches 92.7%, and the missed audit rate is reduced by 55%.
[0078] The dynamic adaptation capability is enhanced, the response delay of new tags is ≤1.2 hours, and the model can adapt to the changes of the college students' academic cycle and campus hotspots in real time; after the iteration of the tag system, the time consumption of model incremental fine-tuning is only 2.5 hours, which saves 79.2% of the computing resources compared with full training (12 hours).
[0079] The scene adaptation is expanded, covering college students' content in 13 major disciplines, 8 types of campus scenes, and 5 types of behavior characteristics, supporting three mainstream modalities of text, image, and video, and can be directly deployed in 10+ types of college students' high-frequency use scenes such as campus social platforms, MOOC platforms, and campus forums.
[0080] The specific implementation of the present application is illustrated as follows: (I) Implementation environment 1. Hardware environment: distributed collection cluster (15 servers, 8 cores and 32G per server), GPU training node (4 NVIDIA A100 graphics cards), database server (2, storage capacity 20TB, supporting Redis cache); 2. Software environment: Python 3.10, PyTorch 2.1, MySQL 8.0, Elasticsearch 8.6 (used for tag-content indexing), YOLOv9 (target recognition), BERT-College (domain-adapted pre-training model).
[0081] (II) Multi-dimensional tag system embodiment 1. Tag hierarchy and subclass: Student basic information layer: 13 major categories (computer, clinical medicine, Chinese language and literature, etc.), 5 grade segments (freshman to postgraduate), 2 education types (undergraduate, postgraduate); School scene layer: 8 types of scenes (course learning, experiment operation, club activity, postgraduate examination review, graduation design, etc.); Interactive behavior layer: 5 types of behaviors (original content, forwarding comments, asking for help, controversial discussion, academic sharing); Hotspot association layer: 4 types of association strength (strong, medium, weak, none), including 20+ types of high-frequency campus hotspot tags such as "postgraduate examination season", "new student enrollment", "academic competition", etc.
[0082] 2. Label weight calculation example: The weight of the student basic information label: The weight is determined by the frequency of professional terms (such as "TensorFlow" appearing ≥3 times in the text), user identity (such as the school affiliation of the student ID), etc.
[0083] The weight of the school scene label: The weight is calculated by target detection in the image (such as recognizing "microscope" to associate with "laboratory scene") or metadata (such as the release location being in the laboratory).
[0084] The potential of interactive behavior label: According to interactive data (such as "forwarding + critical comments" triggering "controversial discussion" label) or content metadata (such as "asking" action) to dynamically adjust the weight.
[0085] The weight of the hotspot label: Based on real-time hotspot word matching (such as "postgraduate examination English true question" triggering "postgraduate examination season" label), the initial association strength is set to "strong".
[0086] The weight calculation of each label can combine various information (such as term frequency, scene matching degree, user interaction, etc.) based on the preset original weight, and through normalization processing (such as Softmax) to ensure the sum is 1.
[0087] (Three) Model training and testing embodiments 1. Training data: Construct a data set containing 1.2 million college students' multi-modal content (60 million text, 350,000 images, 25 million videos), each content is labeled with 5-8 multi-dimensional labels; divide the training set (80%), the validation set (10%), and the test set (10%).
[0088] 2. Training parameters: batch size = 32, initial learning rate = 1e-4, using AdamW optimizer, label embedding vector dimension = 128, attention fusion layer hidden layer dimension = 512; training period = 100 rounds, including 60 rounds of basic training, 20 rounds of cross-label fine-tuning, and 20 rounds of adversarial training (adversarial samples are "label confusion samples", such as labeling "mechanical part diagram" as "building structure diagram").
[0089] 3. Test results: The overall method of the present application is summarized as follows: S1: Construct a "four-layer three-order" college student multi-dimensional label system, determine the label type, weight calculation rule and dynamic update trigger condition; S2: Collect college student multi-modal content according to label dimension, realize the binding storage of "content-multi-label matrix-multi-modal feature"; S3: Guided by multi-dimensional labels, customize the preprocessing of text, image and video, and enhance the label associated features; S4: Input the preprocessed multi-modal features and label embedding vectors into the multi-label cross-fusion model, output the classification results and review conclusions; S5: Update the label system and model parameters based on evaluation indicators and user feedback, realize the linkage iteration of "label-database-model".
[0090] As shown in Figure 2 The third aspect of the embodiment of the present application provides a label determination system for campus platform content, comprising: An acquisition module 31 is configured to acquire at least one of text features, image features and video features corresponding to text information, images and videos of each existing content of a campus content platform; A matching module 32 is configured to match the text features, image features and video features with corresponding content labels to generate an association database; wherein the content labels include at least one of student basic information labels and weights, school scene labels and weights, interactive behavior labels and weights, and hot topic labels and weights; An enhancement module 33 is configured to perform feature enhancement and / or label correction on the text features, image features and video features and the corresponding content labels to obtain enhanced text features, enhanced image features, enhanced video features and corresponding enhanced content labels; A fusion module 34 is configured to perform feature fusion on the enhanced text features, enhanced image features and enhanced video features to obtain fusion features; The training module 35 is configured to input the fusion feature and the enhanced content label into a multi-modal Transformer model for training to obtain a label labeling model; wherein a self-attention sublayer of an encoder of the multi-modal Transformer model performs vector fusion according to the fusion feature and the enhanced content label to obtain a fusion vector, and in the training process, an output layer outputs a label matching result and a content review result based on the fusion vector; The labeling module 36 is configured to obtain at least one of text information, image and video of the current edited content of the campus content platform, corresponding text features, image features and video features, and input the obtained features into the label labeling model after fusion to output a current label matching result and a current content review result.
[0091] In an embodiment of the present application, the student basic information label includes a plurality of course majors, a plurality of grades and a plurality of education types; wherein the weight of the course major is determined according to the number of times of occurrence of the course major term in the text feature, the image feature and the video feature; The school scene label includes laboratory operation, club activity, postgraduate review, course learning, graduation design, social practice, daily life and leisure entertainment; wherein the weight of the school scene label is determined according to the scene target, the course major element target, the experimental operation target and the corresponding scene position in the image feature and the video feature; The interactive behavior label includes original content, forwarding and commenting, asking for help, controversial discussion and academic sharing; The hot topic label includes a campus hot topic event and a social hot topic event; wherein the weight of the hot topic label is determined according to the degree of association with the campus hot topic event or the social hot topic event.
[0092] In an embodiment of the present application, the calculation formula of the fusion vector is: wherein, represents the fusion vector, represents the relevance of the enhanced content label to one of the enhanced text feature, the enhanced image feature and the enhanced video feature in the fusion feature, represents one of the enhanced text feature, the enhanced image feature and the enhanced video feature in the fusion feature, represents the weight of the label embedding vector of the enhanced content label, represents the label embedding vector of the enhanced content label.
[0093] In an embodiment of the present application, the text feature, the image feature and the video feature and the corresponding content label are subjected to feature enhancement and / or label correction to obtain the enhanced text feature, the enhanced image feature, the enhanced video feature and the corresponding enhanced content label, which includes: According to the weight of the professional term in the text feature multiplying the weight coefficient greater than 1 of the corresponding student basic information label in the professional term dictionary, the enhanced content label and weight are obtained; According to the context semantic analysis of the text feature in the professional term dictionary, the corresponding content label is replaced to obtain the enhanced content label; The region associated with the content label in the image feature adopts the super-resolution reconstruction feature to obtain the enhanced image feature; According to the content label, the key frame of the corresponding video feature is extracted; The motion feature is extracted by using the optical flow method on the key frame, and then the corresponding content label is matched to generate the enhanced video feature and the enhanced content label.
[0094] The fourth aspect of the embodiment of the application provides an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to realize the label determination method of the campus platform content provided by the embodiment of the application.
[0095] The fifth aspect of the embodiment of the application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by the processor to realize the steps of the label determination method of the campus platform content provided by the embodiment of the application.
[0096] The memory can include a random access memory (RAM) and a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage system located away from the aforementioned processor.
[0097] The aforementioned processor can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware system.
[0098] The method provided by the embodiment of the present application can be applied to an electronic device. Specifically, the electronic device can be a desktop computer, a portable computer, a smart mobile terminal, a server, etc. Herein, no limitation is made, and any electronic device that can implement the present application falls within the protection scope of the present application.
[0099] For the system / electronic device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant part can be referred to the part of the method embodiment.
[0100] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device implemented in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 a system with the function specified in one or more flows and / or blocks.
[0101] These computer program instructions can also be stored in a computer readable memory capable of guiding the computer or other programmable data processing devices to work in a specific way, so that the instructions stored in the computer readable memory produce a product including instruction devices, which implement the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 a system with the function specified in one or more flows and / or blocks.
[0102] These computer program instructions can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to produce a computer implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 a system with the function specified in one or more flows and / or blocks.
[0103] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A method for determining tags for content on a campus platform, characterized in that, Includes the following steps: Obtain at least one corresponding text feature, image feature, and video feature from the text information, image, and video of each existing piece of content on the campus content platform; The text features, image features, and video features are matched with corresponding content tags to generate an association database; wherein, the content tags include at least one of student basic information tags and weights, school scene tags and weights, interactive behavior tags and weights, and hot topic tags and weights; The text features, image features, and video features, along with the corresponding content tags, are enhanced and / or the tags are corrected to obtain enhanced text features, enhanced image features, enhanced video features, and corresponding enhanced content tags. Enhanced text features, enhanced image features, and enhanced video features are fused to obtain fused features; The fused features and enhanced content tags are input into a multimodal Transformer model for training to obtain a tag labeling model. The self-attention sublayer of the encoder of the multimodal Transformer model performs vector fusion based on the fused features and enhanced content tags to obtain a fused vector. During training, the output layer outputs tag matching results and content review results based on the fused vector. The system obtains at least one corresponding text feature, image feature, and video feature from the text information, image, and video of the currently edited content on the campus content platform. After fusing the obtained features, the system inputs them into a tag labeling model and outputs the current tag matching result and the current content review result.
2. The method as described in claim 1, characterized in that, The student basic information tags include: multiple course majors, multiple grade levels, and multiple academic degree types; wherein, the weight of the course major is determined based on the frequency of occurrence of course major terms in the text features, the image features, and the video features; The school scene tags include: laboratory operations, club activities, postgraduate entrance examination preparation, course learning, graduation project, social practice, daily life and leisure and entertainment; wherein, the weight of the school scene tags is determined based on the scene target, course professional element target, experimental operation target and corresponding scene location in the image features and video features; The interactive behavior tags include: original content, forwarding and commenting, asking questions and seeking help, controversial discussions, and academic sharing; The hot topic tags include: campus hot topics and social hot topics; the weight of the hot topic tags is determined according to the degree of relevance to campus hot topics or social hot topics.
3. The method as described in claim 1, characterized in that, The formula for calculating the fusion vector is: in, Represents the fusion vector. This indicates the correlation between the enhanced content tag and one of the three features in the fusion features: enhanced text features, enhanced image features, and enhanced video features. This indicates one of the following in the fusion feature set: enhanced text feature, enhanced image feature, or enhanced video feature. The weights of the tag embedding vectors representing the enhanced content tags. This represents the tag embedding vector of the enhanced content tag.
4. The method as described in claim 1, characterized in that, The step of performing feature enhancement and / or label correction on the text features, image features, video features, and corresponding content tags to obtain enhanced text features, enhanced image features, enhanced video features, and corresponding enhanced content tags includes: The enhanced content tags and their weights are obtained by multiplying the weights of the professional terms in the text features and their corresponding student basic information tags by a weight coefficient greater than 1 according to the professional terminology dictionary. Based on a terminology dictionary, the text features are analyzed for contextual semantics, and the corresponding content tags are replaced to obtain enhanced content tags. The regions in the image features associated with the content tags are reconstructed using super-resolution features to obtain enhanced image features. Extract keyframes corresponding to video features based on content tags; Motion features are extracted from the keyframes using optical flow, and then matched with corresponding content tags to generate augmented video enhancement features and enhanced content tags.
5. A tagging system for campus platform content, characterized in that, include: The acquisition module is used to acquire at least one corresponding text feature, image feature, and video feature from the text information, image, and video of each existing piece of content on the campus content platform. The matching module is used to match the text features, image features, and video features with corresponding content tags to generate an association database; wherein, the content tags include at least one of student basic information tags and weights, school scene tags and weights, interactive behavior tags and weights, and hot topic tags and weights; The enhancement module is used to enhance and / or correct the text features, image features, and video features and the corresponding content tags to obtain enhanced text features, enhanced image features, enhanced video features and corresponding enhanced content tags; The fusion module is used to fuse enhanced text features, enhanced image features, and enhanced video features to obtain fused features. The training module is used to input the fused features and enhanced content labels into the multimodal Transformer model for training to obtain a labeling model; wherein, the self-attention sublayer of the encoder of the multimodal Transformer model performs vector fusion based on the fused features and enhanced content labels to obtain a fused vector, and during the training process, the output layer outputs the label matching result and content review result based on the fused vector; The annotation module is used to obtain at least one corresponding text feature, image feature, and video feature from the text information, image, and video of the currently edited content on the campus content platform. After fusing the obtained features, the module inputs them into the tag annotation model and outputs the current tag matching result and the current content review result.
6. The system as described in claim 5, characterized in that, The student basic information tags include: multiple course majors, multiple grade levels, and multiple academic degree types; wherein, the weight of the course major is determined based on the frequency of occurrence of course major terms in the text features, the image features, and the video features; The school scene tags include: laboratory operations, club activities, postgraduate entrance examination preparation, course learning, graduation project, social practice, daily life and leisure and entertainment; wherein, the weight of the school scene tags is determined based on the scene target, course professional element target, experimental operation target and corresponding scene location in the image features and video features; The interactive behavior tags include: original content, forwarding and commenting, asking questions and seeking help, controversial discussions, and academic sharing; The hot topic tags include: campus hot topics and social hot topics; the weight of the hot topic tags is determined according to the degree of relevance to campus hot topics or social hot topics.
7. The system as described in claim 5, characterized in that, The formula for calculating the fusion vector is: in, Represents the fusion vector. This indicates the correlation between the enhanced content tag and one of the three features in the fusion features: enhanced text features, enhanced image features, and enhanced video features. This indicates one of the following in the fusion feature set: enhanced text feature, enhanced image feature, or enhanced video feature. The weights of the tag embedding vectors representing the enhanced content tags. This represents the tag embedding vector of the enhanced content tag.
8. The system as described in claim 5, characterized in that, The step of performing feature enhancement and / or label correction on the text features, image features, video features, and corresponding content tags to obtain enhanced text features, enhanced image features, enhanced video features, and corresponding enhanced content tags includes: The enhanced content tags and their weights are obtained by multiplying the weights of the professional terms in the text features and their corresponding student basic information tags by a weight coefficient greater than 1 according to the professional terminology dictionary. Based on a terminology dictionary, the text features are analyzed for contextual semantics, and the corresponding content tags are replaced to obtain enhanced content tags. The regions in the image features associated with the content tags are reconstructed using super-resolution features to obtain enhanced image features. Extract keyframes corresponding to video features based on content tags; Motion features are extracted from the keyframes using optical flow, and then matched with corresponding content tags to generate augmented video enhancement features and enhanced content tags.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for determining the tags of campus platform content as described in any one of claims 1 to 4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for determining the tags of the campus platform content as described in any one of claims 1 to 4.