Ai audio-visual intelligent fusion system and content pushing method for ideological and political propaganda

By using an AI-powered audio-visual intelligent fusion system, combined with campus network location awareness and learning trajectory tracking, personalized ideological and political content is generated and rendered in a differentiated manner. This solves the problems of duplicate pushes and content redundancy in the existing system, and enhances the immersive experience of ideological and political propaganda and the students' viewing completion rate.

CN122437882APending Publication Date: 2026-07-21HEILONGJIANG INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEILONGJIANG INST OF TECH
Filing Date
2026-04-29
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

The existing ideological and political education system lacks a personalized content delivery mechanism, resulting in duplicate pushes, information redundancy, and an inability to accurately identify high-quality content, which affects students' willingness to accept the content and their completion rate.

Method used

The campus network location sensing module determines students' locations, the push scheduling module controls the frequency, the AI ​​content generation module generates personalized content, the content review module performs automatic review, the audio-visual intelligent fusion module performs differentiated rendering, and the learning trajectory tracking module records viewing behavior to construct a personalized learning graph.

Benefits of technology

This approach achieves a deep integration of ideological and political education with the campus environment, enhancing students' sense of immersion and learning experience, avoiding repetitive push notifications, and prioritizing the delivery of highly completed content using peer group data, thereby increasing students' willingness to accept and watch the content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122437882A_ABST
    Figure CN122437882A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of ideological and political propaganda education, especially to an AI audio and picture intelligent fusion system for ideological and political propaganda and a content pushing method thereof, which comprises a campus network position sensing module, a pushing scheduling module, an AI content generating module, a content auditing module, an audio and picture intelligent fusion rendering module and a learning track tracking and popularizing module. The campus network position sensing module is used to estimate the current physical position of a student in a campus, and match it with a building scene area in a campus digital map to determine the current building scene type and the corresponding configuration parameters. When the student walks to a specific building area in the campus, the system actively pushes an AI ideological and political short video matching the building scene atmosphere to the mobile terminal of the student, so as to deeply integrate the ideological and political propaganda education with the actual scene in the campus, enhance the students' sense of identification and learning experience, and realize situational and immersive ideological and political propaganda.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ideological and political education technology, and in particular to an AI-powered audio-visual intelligent fusion system for ideological and political education and its content delivery method. Background Technology

[0002] With the rapid development of mobile internet, artificial intelligence, and campus information infrastructure, ideological and political education in universities is gradually transforming from traditional single forms such as classroom teaching, paper publicity materials, and broadcast notices to digitalization and intelligence based on mobile terminals, campus wireless networks, and intelligent algorithms.

[0003] Current ideological and political education often adopts a unified online video push or offline exhibition board publicity model. When students are active in different building areas on campus, the ideological and political content they receive lacks correspondence with their current physical environment, making it difficult to form a contextualized learning experience. This results in insufficient immersion and engagement in ideological and political education, and students have a low willingness to accept it.

[0004] Existing ideological and political education content delivery systems lack a mechanism for managing individual students' viewing history, making it easy for the same ideological and political education video to be repeatedly pushed to the same student, resulting in information redundancy and student resentment. Furthermore, existing systems fail to effectively utilize student viewing behavior data to evaluate the quality of ideological and political education content, and cannot rely on peer completion data to endorse content quality. This makes it difficult to accurately identify and prioritize high-quality ideological and political education content with high acceptance rates, leading to low efficiency in the allocation of ideological and political education resources and difficulty in improving the overall viewing completion rate.

[0005] Therefore, this invention proposes an AI-powered audio-visual intelligent fusion system for ideological and political education and its content delivery method. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an AI-powered audio-visual intelligent fusion system and its content delivery method for ideological and political education, thereby resolving the technical problems mentioned in the background section.

[0007] To achieve the above objectives, the present invention provides the following technical solution: An AI-powered audio-visual intelligent fusion system for ideological and political education includes: The campus network location awareness module, after connecting to the campus network, estimates the current physical location on campus and matches it with the building scene area in the campus digital map; The push scheduling module controls the frequency of ideological and political videos being pushed to students' mobile terminals to avoid excessive disruption. The AI ​​content generation module generates ideological and political propaganda content based on the architectural scene type and student identity identifier, and excludes content that has already been viewed to avoid duplicate pushes. The content moderation module automatically reviews and scores AI-generated content. The audio-visual intelligent fusion rendering module, based on the audio-visual fusion configuration parameters, performs differentiated rendering on the approved content and outputs short ideological and political propaganda films in video format; The learning trajectory tracking and promotion module records and promotes the content viewing records of each student's identity in different architectural scene types.

[0008] In one possible implementation, the campus network location awareness module includes: The network positioning unit obtains the campus network access information of the mobile terminal and estimates the student's current location coordinates; The scene matching unit matches the location coordinates with the building scene area, determines the building scene type, and loads the corresponding audio-visual fusion configuration parameters.

[0009] In one possible implementation, the push scheduling module includes: Frequency control unit: Check if there are any push records for this building scene type on the current day; The push trigger unit initiates a video push request to the mobile terminal.

[0010] In one possible implementation, the AI ​​content generation module includes: The content generation unit maintains a knowledge graph mapping campus architectural semantics, ideological and political themes, and video generation parameters, and generates multiple videos with different styles in the same scene. The deduplication filtering unit queries the list of viewed content identifiers based on the identity identifier and excludes viewed content from the content pool.

[0011] In one possible implementation, the content moderation module includes: The compliance review unit performs sensitive keyword scanning, historical fact knowledge base comparison and verification, and video frame content detection on the narrative script. The scoring and determination unit calculates the content review score based on the test results of the comprehensive compliance review unit.

[0012] In one possible implementation, the audio-visual intelligent fusion rendering module includes: The video rendering unit calls the video model to generate a short video that matches the atmosphere of the architectural scene, and overlays a prompt watermark corresponding to the architectural scene type on the corner area of ​​the video screen; The audio processing unit processes the audio track according to the audio switch status.

[0013] In one possible implementation, the learning trajectory tracking and generalization module: The viewing record unit records the viewing behavior when students finish playing the video, and constructs a personal ideological and political learning map for each student based on the viewing record. The promotion unit is used to calculate the comprehensive completion index of each content identifier based on the total viewing behavior data accumulated by the viewing record unit, generate a promotion priority queue, and push high completion content to the target students first. The completion calculation unit calculates the content viewing completion of each student's identity under the global and building scene types.

[0014] One possible implementation includes the following steps: S1. When a student's mobile terminal accesses the campus wireless network, the system estimates the location coordinates and matches the building scene type, and loads the audio-visual fusion configuration parameters. S2. The system queries the knowledge graph to generate ideological and political propaganda content, and performs deduplication filtering for viewing. S3. The system automatically reviews the generated narrative script, and after the review is passed, it enters the audio-visual fusion rendering process. S4. Perform audio-visual differentiation rendering based on the architectural scene type and output adapted videos; S5. Record the viewing behavior and the video with the highest completion rate, update the personal ideological and political learning map and calculate the learning completion rate.

[0015] In one possible implementation, step S1, estimating location coordinates and matching building scene type, specifically includes: Suppose the terminal detects k wireless access points, and the known deployment coordinates of each access point are: The received signal strength indication value is The network positioning unit then uses the weighted centroid algorithm to estimate the terminal's location coordinates. : in , For the first The absolute value of the signal strength of each access point, the scene matching unit will estimate the location coordinates. The scene matching unit performs point inclusion determination with the polygon boundaries of each building scene area marked in the campus digital map. The scene matching unit loads the audio-visual fusion configuration parameters corresponding to the scene from the configuration database, including playback mode identifier, content theme pool number, corner prompt text, audio switch status and video style template number.

[0016] In one possible implementation, step S5, which records the content viewing behavior and the video with the highest completeness, specifically includes: The video with the highest completion rate is selected, and this video can be given priority for students visiting this scene for the first time. The viewing behavior set of each content tag is obtained from the viewing record unit, and the comprehensive completion index is calculated for each content tag. : in The overall completion index for content identifier c ranges from 0 to 1.2. A higher value indicates that the content is more popular with students and the quality of viewing is better; n is the total number of unique students who have viewed content c. Let be the completion rate of the i-th student for playing content c, expressed as a percentage, with a value ranging from 0 to 100. The actual playback duration of content c for the i-th student, in seconds; The full standard duration of content c is in seconds; μ is the duration normalization weighting coefficient. The promotion unit will assign each content identifier according to the comprehensive completion index. Arrange in descending order to generate a priority promotion sequence.

[0017] Beneficial effects compared to existing technologies: 1. In this solution, the campus network location awareness module estimates the student's current physical location on campus and matches it with building scene areas in the campus digital map to determine the current building scene type and corresponding configuration parameters. When a student walks to a specific building area on campus, the system proactively pushes AI-generated ideological and political education short videos that are appropriate for the atmosphere of the building scene to their mobile terminal, deeply integrating ideological and political education with the actual campus scene, enhancing students' sense of immersion and learning experience, and achieving contextualized and immersive ideological and political education. 2. In this solution, the deduplication unit queries the list of content tags that have been viewed based on the student's identity and excludes viewed content from the content pool to ensure that the same student identity does not watch the same promotional video repeatedly under the same architectural scene type. At the same time, the promotion unit calculates the comprehensive completion index of each content tag based on the total viewing behavior data accumulated by the viewing record unit, generates a promotion priority queue, and pushes high completion content to target students first. The viewing behavior data of peer groups forms the endorsement of content quality, thereby increasing new students' willingness to accept and complete the viewing of short ideological and political videos. Attached Figure Description

[0018] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0019] Figure 1 This is a system framework diagram of the present invention; Figure 2This is a schematic diagram of the overall structure and flow of the present invention. Detailed Implementation

[0020] Preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. However, the present invention can also be implemented in various different forms, and therefore the present invention is not limited to the embodiments described below. In addition, for the purpose of more clearly describing the present invention, parts not connected to the invention will be omitted from the drawings. The technical solution in this application embodiment is to solve the problems mentioned in the background art, and the overall idea is as follows: Example

[0021] Please refer to Figure 1 and Figure 2 As shown, this embodiment of the invention provides an AI-powered audio-visual intelligent fusion system and its content delivery method for ideological and political education. This system and method are used in conjunction with a campus network and digital map positioning to proactively push AI-generated short videos tailored to the architectural scenes to students' mobile terminals when they walk through specific building areas on campus, achieving contextualized and immersive ideological and political education.

[0022] The AI-powered audio-visual intelligent fusion system for ideological and political education provided in this embodiment includes a campus network location awareness module, a push scheduling module, an AI content generation module, a content review module, an audio-visual intelligent fusion rendering module, and a learning trajectory tracking module.

[0023] The campus network location awareness module estimates a student's current physical location on campus after their mobile terminal connects to the campus network. It then matches this location with building scene areas on the campus digital map to determine the current building scene type and corresponding configuration parameters. This module includes a network positioning unit and a scene matching unit. The network positioning unit obtains the campus network access information of the student's mobile terminal, and estimates the student's current location coordinates based on the signal strength of multiple wireless access points received by the terminal and the deployment coordinates of the campus wireless access points. The scene matching unit stores the campus digital map and the geometric boundary information of the building scene area. It matches the location coordinates with the building scene area to determine the building scene type and loads the corresponding audio-visual fusion configuration parameters. The configuration parameters include playback mode identifier, content theme pool number, corner prompt text, audio on / off status and video style template number.

[0024] The push scheduling module controls the frequency of ideological and political videos pushed to students' mobile terminals to avoid excessive disruption. This module includes a frequency control unit and a push triggering unit. The frequency control unit checks whether the student has a push notification record for that building scene type on that day, based on the student's identity, building scene type, and current date. If a record exists, the push notification process is terminated; otherwise, the subsequent process is allowed to continue. After the frequency control unit approves the request, the push triggering unit sends a video push request to the student's mobile terminal.

[0025] The AI ​​content generation module generates ideological and political education content based on the architectural scene type and student identification, and excludes already viewed content to avoid duplicate pushes. This module includes a content generation unit and a deduplication filtering unit; The content generation unit maintains a knowledge graph mapping campus building semantics with ideological and political themes and video generation parameters. The knowledge graph contains four-dimensional nodes and their relationship edges, including spiritual lineage, historical events, geographical scenes, and personal deeds. It generates multiple videos with different styles (such as A1 and A2) in the same scene. The deduplication filtering unit queries the list of content identifiers that a student has already viewed based on the student's identity identifier, and excludes the viewed content from the content pool to ensure that the same student identity identifier does not watch the same promotional video repeatedly under the same architectural scene type.

[0026] The content moderation module automatically reviews and scores AI-generated content. This module includes a compliance review unit and a scoring determination unit. The compliance review unit performs sensitive keyword scanning, historical fact knowledge base comparison and verification, and video frame content detection on the narrative script. The scoring and judgment unit calculates the content review score based on the detection results of the comprehensive compliance review unit. When the review score is lower than the preset review threshold, the content is deleted directly and the rendering process is blocked. When the review score is higher than or equal to the preset review threshold, a review pass mark is output and the audio-visual fusion rendering process is entered.

[0027] The audio-visual intelligent fusion rendering module, based on the audio-visual fusion configuration parameters, performs differentiated rendering on the approved content, outputting short ideological and political propaganda videos. This module includes a video rendering unit and an audio processing unit; The video rendering unit selects the corresponding rendering strategy according to the type of building scene, calls the Wensheng video model to generate a short video that matches the atmosphere of the building scene, and overlays a prompt watermark corresponding to the type of building scene on the corner area of ​​the video screen. The audio processing unit processes the audio track according to the audio switch status. When the audio switch is off, it emptys the audio track of the output video and forces it to be muted. When the audio switch is on, it mixes and outputs background music, narration, and sound effects proportionally according to the ambient sound blending configuration.

[0028] The learning trajectory tracking and promotion module records and promotes the content viewing records of each student's identity in different architectural scene types, constructs a personal ideological and political learning map, and calculates the learning completion rate. This module includes a viewing record unit, a promotion unit, and a completion rate calculation unit; The viewing record unit records the viewing behavior when a student finishes playing a video, including student identity, architectural scene type, content identifier, playback duration, playback completion rate, and viewing timestamp. Based on the viewing record, a personal ideological and political learning map is constructed for each student identity. The map uses architectural scene type as the first-level node, spiritual lineage theme as the second-level node, and specific content identifier as the leaf node. The promotion unit is used to calculate the comprehensive completion index of each content identifier based on the total viewing behavior data accumulated by the viewing record unit, generate a promotion priority queue, and push high completion content to the target students first. The completion calculation unit calculates the content viewing completion of each student's identity under the global and building scene types.

[0029] The following describes in detail the AI-powered audio-visual intelligent fusion content push method for ideological and political education provided by this invention, based on the functional modules of the aforementioned system. This method corresponds to the specific processing procedures executed by each module in the aforementioned system during runtime, and the specific steps are as follows: S1. When a student's mobile terminal connects to the campus wireless network, the system estimates the location coordinates, matches the building scene type, and loads the audio-visual fusion configuration parameters. Student mobile terminals access the campus network via the campus wireless network. The network positioning unit obtains the signal strength values ​​of multiple wireless access points received by the terminal. Assume the terminal detects k wireless access points, and the known deployment coordinates of each access point are... The received signal strength indication value is The network positioning unit then uses the weighted centroid algorithm to estimate the terminal's location coordinates. : in , For the first The algorithm uses the absolute signal strength of each access point. It leverages the correlation between signal strength and distance to assign higher weights to access points closer to the terminal, thereby improving positioning accuracy.

[0030] The scene matching unit will estimate the position coordinates. Perform point inclusion determination between the polygon boundaries of each building scene area marked on the campus digital map. Let the sequence of polygon vertices of a certain building scene area be... Points are determined using the ray method. Whether it is inside the polygon: from the point Draw a horizontal ray to the right and calculate the number of intersections between the ray and each side of the polygon. If the number of intersections is odd, the point is considered to be within the region; otherwise, it is considered to be outside the region. Iterate through each building scene region sequentially to determine the type of building scene the student is currently in. For example, designate the pond as A, the library as B, and the bridge as C.

[0031] After determining the architectural scene type, the scene matching unit loads the corresponding audio-visual fusion configuration parameters from the configuration database, including the playback mode identifier, content theme pool number, corner prompt text, audio on / off status, and video style template number. For example, the architectural scene type number of the area around the pond is A, and the corresponding configuration parameters are as follows: the playback mode identifier is the revolutionary history documentary mode of the pond scene, the content theme pool number is the water culture theme pool, the corner prompt text is "Red Boat Spirit, Water Charm, Original Intention", the audio on / off status is "On", and the video style template number is the "Heavy Historical Tone Template". The video number is also specified, for example, Red Boat Spirit is A1, Flood Fighting Spirit is A2, etc.

[0032] S2. The system queries the knowledge graph to generate ideological and political propaganda content, and performs deduplication filtering for viewing. The system queries the knowledge graph based on the content theme pool number to determine the associated ideological and political theme pool and video generation parameters for the architectural scene type. The content generation unit, combined with student identification, retrieves students' interests and historical viewing records, and generates multimodal narrative prompts for a text-based video model based on a large language model. These multimodal narrative prompts include descriptions of historical scene shots, lighting and atmosphere instructions, color tone style instructions, narrative rhythm instructions, and audio style instructions. The narrative script is strictly based on Party history documents and ideological and political education materials.

[0033] The deduplication filtering unit queries the set of content identifiers that a student has viewed within the current architectural scene type, based on the student's identity identifier. Let the content pool corresponding to this building scene type be... The set of optional content for this time for: This means excluding viewed content from the content pool. If If it is an empty set, then the current push process will be terminated; if If not empty, select the ideological and political propaganda content to be pushed this time to ensure that the same student identity does not watch the same propaganda video repeatedly under the same architectural scene type.

[0034] S3. The system automatically reviews the generated narrative script. Once approved, it proceeds to the audio-visual fusion rendering process. The compliance review unit conducts a three-tiered review of the narrative script: the first tier involves scanning for sensitive keywords, and passing... The tree prefix matching algorithm searches the keyword database for inappropriate words; the second layer is historical fact verification, which extracts historical events, figures, times and places involved in the narrative script and compares them with the Party history knowledge base; the third layer is image content detection, which detects illegal images in the video frames generated by the Wensheng video model.

[0035] The scoring and evaluation unit calculates the content review score S based on the above three review results: S represents the overall content review score, ranging from 0 to 100 points. A higher score indicates that the content is more compliant. The keyword scanning score is determined by... The tree prefix matching algorithm outputs a score of 100 when no sensitive keywords are matched in the narrative script. For each sensitive keyword matched, the score is deducted according to the severity, with a minimum score of 0. The factual verification score is determined by comparing the historical factual descriptions in the narrative script with the Party history knowledge base. When all factual descriptions are consistent with the knowledge base, the score is 100 points. For each factual error found, the corresponding score is deducted according to the severity, with a minimum of 0 points. The image detection score is calculated by outputting video frames generated by the Wensheng video model through the image detection model. When there is no violation content in the image, the score is 100 points. For each violation detected, the corresponding score is deducted according to the severity of the violation, with a minimum of 0 points.

[0036] w1 is the weighting coefficient of the keyword scanning score, representing the proportion of compliance in the overall score; w2 is the weighting coefficient of the fact verification score, representing the proportion of historical accuracy in the overall score; w3 is the weighting coefficient of the image detection score, representing the proportion of image content compliance in the overall score; w1, w2, and w3 satisfy w1 + w2 + w3 = 1 to ensure the normalization of the overall score.

[0037] The calculated audit score S is compared with the preset audit threshold. Comparison, The minimum passing score for content review is set at 85 points, leaving very little room for error. If the threshold is too low (e.g., 60 points), content containing slightly sensitive expressions or factual deviations may enter the dissemination process, misleading students. If the threshold is too high (e.g., 95 points), the pass rate for AI-generated content will be too low, significantly reducing system usability, and AI-generated content will struggle to match the absolute perfection of human-written content in terms of expressive diversity. Through multiple rounds of testing, 85 points has been verified to maintain the system's operational efficiency while ensuring a minimum level of content safety; that is, it allows for very minor expressive flaws but resolutely blocks content with substantial errors. If so, the content will be deleted directly, the rendering process will be blocked, and a deletion log will be recorded for subsequent traceability and analysis; if If so, then output.

[0038] S4. Perform audio-visual differentiation rendering based on the architectural scene type and output adapted videos. The video rendering unit selects the corresponding rendering strategy according to the playback mode identifier and calls the Wensheng video model to generate short videos that match the atmosphere of the architectural scene. For example: when the architectural scene type is the internal area of ​​the library, the playback mode identifier is area B. The same ideological and political thought generates multiple short videos with different styles, such as: B1 (1) conveys ideological and political thought in a serious style, and B1 (2) conveys ideological and political thought in a more interesting style, giving students different viewing styles to increase the effect of ideological and political promotion. B1 uses classics, manuscripts, and historical video materials as the core visual image. The color tone of the picture adopts a low saturation warm gray academic tone, that is, the color purity is reduced to below 30% and a warm yellow color is superimposed on the gray base to simulate the reading environment of ancient paper; the light and shadow are soft and quiet, and diffuse reflection soft light is used to avoid hard edge shadows. The light intensity is controlled in the low illumination range; the narrative rhythm is slow and solemn, and the subtitles are presented in a line-by-line floating manner. The audio track is not output throughout the process to avoid disturbing other readers in the library; B2 uses macro documentary lens, knowledge graph three-dimensional space roaming, and history The core visual imagery is the dynamic reorganization of historical documents into subtitles. The video uses handheld camera footage of people walking through the bookshelf corridor and time-lapse photography to record real footage of the morning sunlight sweeping across the spines of classic Marxist-Leninist books. Combined with the video presentation of text nodes in three-dimensional space extending and connecting like a star map to form a knowledge graph, the abstract ideological and political concepts are transformed into a spatial narrative with a time dimension. The color tone of the video adopts a fresh wood-based academic color scheme, with light beige, natural wood brown, and light blue as the main tones, and the color purity is controlled at around 45%, simulating the fresh atmosphere of a bright reading area and natural wood bookshelves in a modern library.

[0039] The video rendering unit overlays a watermark on the corners of the video screen. The watermark content is obtained from the configuration parameters based on the building scene type. For example, the text "Please keep quiet inside the library" is overlaid on the area inside the library, and the text "Red Boat Spirit, Water Charm, Original Intention" is overlaid on the area around the pond.

[0040] The audio processing unit processes the audio track based on the audio switch status. When the audio switch status in the configuration parameters is off, the audio track of the output video is left empty, forcibly muted. When the audio switch status is on, the human voice and background sound are mixed proportionally according to the ambient sound blending configuration, and the output audio gain is calculated using the following formula: in This is the final audio gain after mixing, with a value ranging from 0 to 1; For voice gain, the recording volume is determined by the volume of the narration or motivational narration. The background sound gain is determined by the overall volume of the mixture of ambient sound effects and background music. The background sound includes ambient sound effects such as flowing water and wind, as well as background music such as string instruments and folk instruments.

[0041] This is the vocal mixing ratio coefficient, which indicates the weight of vocals in the final mix. This is the background sound mixing ratio coefficient, representing the weight of the background sound in the final mix. and satisfy To ensure that the total gain is normalized after mixing and to avoid overload distortion in the output audio, the mixing ratio is dynamically adjusted according to the characteristics of the architectural scene.

[0042] S5. Record viewing behavior and the video with the highest completion rate, update the personal ideological and political learning map, and calculate the learning completion rate. After students finish watching the video, the viewing record unit records the viewing behavior, including student identity, architectural scene type, content identifier, playback duration, completion rate, and viewing timestamp, and updates the individual ideological and political learning graph. Simultaneously, based on students with high completion rates, the video with the highest completion rate is selected, and this video can be prioritized for students visiting this scene for the first time. The viewing record unit obtains the viewing behavior set for each content identifier, and calculates its comprehensive completion index for each content identifier. : in The overall completion index for content identifier c ranges from 0 to 1.2. A higher value indicates that the content is more popular with students and the quality of viewing is better; n is the total number of unique students who have viewed content c. Let be the completion rate of the i-th student for playing content c, expressed as a percentage, with a value ranging from 0 to 100. The actual playback duration of content c for the i-th student, in seconds; , where is the complete standard duration of content c, in seconds, used to normalize the actual playback time of different students to a unified dimension; μ is the duration normalization weight coefficient, which is set to 0.2 in this embodiment because the playback completion rate can directly reflect the attractiveness of the content and should have the dominant weight, while the viewing time percentage is used as an auxiliary indicator to identify whether students are truly immersed in watching, to avoid the completion rate being distorted due to students accidentally opening the video, and to prevent the duration indicator from overshadowing the main indicator. The promotion unit will assign each content identifier according to the comprehensive completion index. Sort in descending order to generate a priority promotion sequence. For example, when pushing content to a student, the promotion unit first retrieves the student's unviewed content set under the current architectural scene type, and prioritizes selecting from that set. The content with the highest completion rate is pushed to students; if a student has already watched all the content in the priority promotion sequence for that scenario type, then the remaining ordinary content is randomly selected according to the deduplication filtering unit rules. Through this mechanism, high-completion content receives greater exposure weight, and the viewing behavior data of peer groups forms a guarantee of content quality, thereby increasing new students' willingness to accept and complete ideological and political short videos.

[0043] The completion calculation unit calculates the global completion score and the completion score for each student's identity and building scene type. Let the total amount of content in the system be N, and the total amount of content viewed by a student be n. Then, the student's global completion score is calculated as follows: for: Let M be the total number of content pools for a certain architectural scene type, and m be the number of content viewed by the student under this architectural scene type. Then, what is the completion rate of this architectural scene type? for: The system updates students' learning profiles based on the calculation results. When the overall completion rate or the completion rate for each building scene type reaches a preset threshold, the system will no longer push information to that student.

[0044] Finally, it should be noted that the above embodiments are merely examples for clearly illustrating the present invention and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. An AI-powered audio-visual intelligent fusion system for ideological and political education, characterized in that: include: The campus network location awareness module, after connecting to the campus network, estimates the current physical location on campus and matches it with the building scene area in the campus digital map; The push scheduling module controls the frequency of ideological and political videos being pushed to students' mobile terminals to avoid excessive disruption. The AI ​​content generation module generates ideological and political propaganda content based on the architectural scene type and student identity identifier, and excludes content that has already been viewed to avoid duplicate pushes. The content moderation module automatically reviews and scores AI-generated content. The audio-visual intelligent fusion rendering module, based on the audio-visual fusion configuration parameters, performs differentiated rendering on the approved content and outputs short ideological and political propaganda films in video format; The learning trajectory tracking and promotion module records and promotes the content viewing records of each student's identity in different architectural scene types.

2. The AI-powered audio-visual intelligent fusion system for ideological and political propaganda as described in claim 1, characterized in that, The campus network location awareness module includes: The network positioning unit obtains the campus network access information of the mobile terminal and estimates the student's current location coordinates; The scene matching unit matches the location coordinates with the building scene area, determines the building scene type, and loads the corresponding audio-visual fusion configuration parameters.

3. The AI-powered audio-visual intelligent fusion system for ideological and political propaganda as described in claim 1, characterized in that, The push scheduling module includes: Frequency control unit: Check if there are any push records for this building scene type on the current day; The push trigger unit initiates a video push request to the mobile terminal.

4. The AI-powered audio-visual intelligent fusion system for ideological and political propaganda as described in claim 1, characterized in that, The AI ​​content generation module includes: The content generation unit maintains a knowledge graph mapping campus architectural semantics, ideological and political themes, and video generation parameters, and generates multiple videos with different styles in the same scene. The deduplication filtering unit queries the list of viewed content identifiers based on the identity identifier and excludes viewed content from the content pool.

5. The AI-powered audio-visual intelligent fusion system for ideological and political propaganda as described in claim 1, characterized in that, The content moderation module includes: The compliance review unit performs sensitive keyword scanning, historical fact knowledge base comparison and verification, and video frame content detection on the narrative script. The scoring and determination unit calculates the content review score based on the test results of the comprehensive compliance review unit.

6. The AI-powered audio-visual intelligent fusion system for ideological and political propaganda as described in claim 1, characterized in that, The audio-visual intelligent fusion rendering module includes: The video rendering unit calls the video model to generate a short video that matches the atmosphere of the architectural scene, and overlays a prompt watermark corresponding to the architectural scene type on the corner area of ​​the video screen; The audio processing unit processes the audio track according to the audio switch status.

7. The AI-powered audio-visual intelligent fusion system for ideological and political propaganda as described in claim 1, characterized in that, The learning trajectory tracking and promotion module: The viewing record unit records the viewing behavior when students finish playing the video, and constructs a personal ideological and political learning map for each student based on the viewing record. The promotion unit is used to calculate the comprehensive completion index of each content identifier based on the total viewing behavior data accumulated by the viewing record unit, generate a promotion priority queue, and push high completion content to the target students first. The completion calculation unit calculates the content viewing completion of each student's identity under the global and building scene types.

8. A method for pushing AI audio-visual intelligent fusion content for ideological and political propaganda using the AI ​​audio-visual intelligent fusion system for ideological and political propaganda as described in any one of claims 1 to 7, characterized in that, Includes the following steps: S1. When a student's mobile terminal accesses the campus wireless network, the system estimates the location coordinates and matches the building scene type, and loads the audio-visual fusion configuration parameters. S2. The system queries the knowledge graph to generate ideological and political propaganda content, and performs deduplication filtering for viewing. S3. The system automatically reviews the generated narrative script, and after the review is passed, it enters the audio-visual fusion rendering process. S4. Perform audio-visual differentiation rendering based on the architectural scene type and output adapted videos; S5. Record the viewing behavior and the video with the highest completion rate, update the personal ideological and political learning map and calculate the learning completion rate.

9. The AI-powered audio-visual intelligent fusion content push method for ideological and political propaganda as described in claim 8, characterized in that, Step S1, estimating location coordinates and matching building scene type, specifically includes: Suppose the terminal detects k wireless access points, and the known deployment coordinates of each access point are: The received signal strength indication value is The network positioning unit then uses the weighted centroid algorithm to estimate the terminal's location coordinates. : in , For the first The absolute value of the signal strength of each access point, the scene matching unit will estimate the location coordinates. The scene matching unit performs point inclusion determination with the polygon boundaries of each building scene area marked in the campus digital map. The scene matching unit loads the audio-visual fusion configuration parameters corresponding to the scene from the configuration database, including playback mode identifier, content theme pool number, corner prompt text, audio switch status and video style template number.

10. The AI-powered audio-visual intelligent fusion content push method for ideological and political propaganda as described in claim 8, characterized in that, Step S5 records the viewing behavior and the video with the highest completeness, specifically including: The video with the highest completion rate is selected, and this video can be given priority for students visiting this scene for the first time. The viewing behavior set of each content tag is obtained from the viewing record unit, and the comprehensive completion index is calculated for each content tag. : in The overall completion index for content identifier c ranges from 0 to 1.

2. A higher value indicates that the content is more popular with students and the quality of viewing is better; n is the total number of unique students who have viewed content c. Let be the completion rate of the i-th student for playing content c, expressed as a percentage, with a value ranging from 0 to 100. The actual playback time of content c for the i-th student, in seconds; The full standard duration of content c is in seconds; μ is the duration normalization weighting coefficient. The promotion unit will assign each content identifier according to the comprehensive completion index. Arrange in descending order to generate a priority promotion sequence.