AI question answering method and system based on video key frame and progress association
By capturing the progress mark and key frame knowledge point coordinates when the video is paused in the intelligent education video platform, a spatial correlation map of touch hot spots and knowledge points is generated. Combined with the touch trajectory and voice semantic features, the question confidence is dynamically generated. This solves the problem of visual focus and user interaction intention being ignored in the existing technology, achieves precise positioning and active questioning, and improves the accuracy of answering questions and user experience.
Patent Information
- Application Number
- CN202511109433.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-08-08
AI Technical Summary
The AI question-answering technology of existing smart education video platforms relies on single text matching, ignoring visual focus and user interaction intentions, resulting in the inability to accurately locate question points in the screen and the lack of an active follow-up mechanism, leading to irrelevant or inaccurate answers.
By capturing the progress mark and key frame knowledge point coordinates when the video is paused, a spatial correlation map of touch hot spots and knowledge points is generated. The focus area markers are generated based on the touch trajectory. The spatial distribution of the map, the geometric properties of the markers, the coordinate topological relationship and the speech semantic features before the pause are integrated to dynamically generate the question confidence, realizing intelligent decision routing and multimodal information fusion.
It achieves precise question positioning, adaptive interactive questioning and knowledge-structured answers under multimodal information fusion, significantly improving the accuracy of answering questions in video scenes and user experience.
Smart Images

Figure CN120611031A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent educational video interaction technology, and in particular to an AI question-answering method and system based on the association between video key frames and progress. Background Art
[0002] On smart education video platforms, users often pause videos because they don't understand specific knowledge points. To provide efficient and accurate question-answering services, there is an urgent need for a technology that can automatically understand and intelligently respond to users' questions based on the current video progress, keyframe visual content, and possible user interaction intentions at the moment of pause. The core requirements of this technology are: accurately capturing the focus of the knowledge point at the moment of pause, understanding the context of the user's potential questions, and proactively initiating multiple rounds of context-related follow-up interactions when the AI itself cannot highly determine the details of the question, such as "Are you questioning the data formula in area A in the figure, or do you not understand the operating steps in area B?", so as to ultimately identify the problem and provide layered answers, thereby significantly improving the accuracy of the question answering and user experience.
[0003] Currently, one targeted technical solution uses video subtitles or automated speech recognition (ASR) near pause points to extract keywords and match them with a pre-set Q&A knowledge base. This solution analyzes the text content before and after the pause to identify possible keywords or phrases, compares them with Q&A entries organized by knowledge points, and returns the most matching entry directly to the user as the answer. Some implementations also incorporate simple progress markers to associate the general knowledge points.
[0004] However, this existing solution has obvious limitations: it relies heavily on text information and completely ignores the rich visual elements and their spatial relationships in the keyframes, resulting in the inability to locate the actual focus of the question in the picture; at the same time, the solution is essentially a passive response mechanism and lacks the ability to assess the confidence of the user's question. When the text keyword match is vague or ambiguous, the system cannot proactively initiate targeted multiple rounds of follow-up questions based on the current picture content to clarify the details; in addition, the understanding of the coherence of the speech semantics before the pause and the topological relationship of the picture elements is weak, making it difficult to support intelligent decision-making combined with multimodal context, ultimately leading to inaccurate or irrelevant answers to complex questions. Summary of the Invention
[0005] This application provides an AI question-answering method and system based on the association of video keyframes and progress, which is used to solve the problem in the existing technology that relies on single text matching and ignores visual focus and user interaction intention, resulting in the inability to accurately locate question points in the picture and the lack of an active questioning mechanism.
[0006] In the first aspect, the present application provides an AI question answering method based on the association between video keyframes and progress, including: When the user performs a video pause operation, the progress indicator and the corresponding key frame image are captured, and the spatial coordinate information representing the target area of the knowledge point in the key frame image is extracted; Capturing touch trajectory data that meets the pressure trigger condition through the touch screen, generating a spatial correlation map of touch hotspots and knowledge points based on the progress indicator and the spatial coordinate information, and generating a focus area mark based on the resident features of the touch trajectory data; Generate a question confidence by fusing the spatial distribution characteristics of the spatial association map, the geometric properties of the focus region marker, the topological relationship of the spatial coordinate information, and the semantic features of the associated speech segment before the pause operation; Decision routing is performed based on the quantified level of the question confidence, generating instant question answering content when a direct response threshold is reached, and activating a dynamic questioning engine when the direct response threshold is not reached; Parsing the semantic types of knowledge points in the spatial association graph by the dynamic questioning engine to generate a set of questioning options by matching the subject-specific questioning strategy library; Based on the user's selection operation of an option in the set of follow-up question options, question answering content including a core answer layer and an associated knowledge expansion layer is generated.
[0007] Optionally, the dynamic questioning engine is used to parse the semantic types of knowledge points in the spatial association graph to match the subject-specific questioning strategy library to generate a set of questioning options, including: Extracting target areas of knowledge points corresponding to strong correlation marks in the spatial correlation map, and classifying the target areas into formula, concept, or chart semantic types according to the content of the target areas; Matching the subject-specific questioning strategy library according to the semantic type, calling the step derivation option set corresponding to the formula class, the definition comparison option set corresponding to the concept class, or the data interpretation option set corresponding to the chart class as the basic questioning option set; Based on the average duration of all touch points in the touch trajectory data, the priority order of the options in the basic follow-up question option set is adjusted, and the high-frequency question options are placed at the top to generate a final follow-up question option set.
[0008] Optionally, adjusting the priority order of options in the basic set of follow-up question options based on the average duration of all touch points in the touch trajectory data, and placing high-frequency question options at the top to generate a final set of follow-up question options includes: Calculating the arithmetic mean of the durations of all touch points in the touch trajectory data as the average duration, and counting the number of times each option in the basic follow-up question option set has been historically selected from a historical question-answering database; Rearrange the order of the options in the basic question option set in descending order of the number of times the options were selected historically, and generate an adjusted set with high-frequency question options at the top; When the average duration exceeds the interaction depth threshold, adding a deep analysis option to the rearranged basic question option set to form a final question option set; When the average duration does not exceed the interaction depth threshold, the rearranged basic question option set is directly used as the final question option set.
[0009] Optionally, the generating of the question answering content including the core answer layer and the associated knowledge extension layer based on the user's selection operation of the option in the set of follow-up question options includes: In response to a user clicking on a specific option in the set of follow-up options, obtaining identification information of the specific option; Extracting the standard parsing content of the corresponding knowledge point from the pre-stored parsing text library as the core answer layer according to the identification information; Searching the video metadata database based on the progress indicator and the indicator information to obtain background knowledge modules and advanced knowledge modules associated with the current knowledge point, and combining them to form an associated knowledge expansion layer; The core answer layer is placed at the beginning of the question-answering content, and the related knowledge extension layer is attached to the end of the core answer layer in an expandable hierarchical structure to generate complete question-answering content.
[0010] Optionally, the capturing of touch trajectory data satisfying a pressure trigger condition through a touch screen, generating a spatial correlation map of touch hotspots and knowledge points in combination with the progress indicator and the spatial coordinate information, and generating a focus area marker based on a resident feature of the touch trajectory data, includes: Monitor the touch screen pressure sensor data. When the pressure value continuously exceeds the preset pressure threshold, record the touch point coordinate sequence and the duration of each touch point to form touch trajectory data. A closed polygon formed by continuous touch points is defined as a touch hot zone, the spatial coordinate information of the touch hot zone is compared with the target area of each knowledge point, and a spatial association map of the touch hot zone and the knowledge point is generated based on the comparison results; Touch points whose duration exceeds a dwell threshold in the touch trajectory data are extracted as dwell feature points, and a circular mark with a radius proportional to the duration is generated as a focus area mark with the coordinates of each dwell feature point as the center.
[0011] Optionally, comparing the spatial coordinate information of the touch hot zone with the target area of each knowledge point, and generating a spatial association map of the touch hot zone and the knowledge point according to the comparison result, includes: When the closed polygon formed by the touch hot zone completely covers the spatial coordinate information of the target area of a single knowledge point, a strong association mark between the touch hot zone and the corresponding knowledge point is established; When the closed polygon formed by the touch hot zone covers the spatial coordinate information of multiple knowledge point target areas at the same time, a weak association mark is established between the touch hot zone and the corresponding multiple knowledge points; The strong association marks and weak association marks are aggregated to generate a spatial association map including a set of association relationship marks between touch hot areas and knowledge points.
[0012] Optionally, the generating of the question confidence by fusing the spatial distribution characteristics of the spatial association map, the geometric properties of the focus region marker, the topological relationship of the spatial coordinate information, and the semantic features of the associated speech segment before the pause operation includes: Counting the ratio of the number of strongly associated markers to the total number of markers in the spatial association map as a spatial distribution characteristic; Calculate the straight-line distance between the center coordinates of each focus area mark and the center coordinates of the nearest knowledge point target area as the geometric attribute; Traversing the spatial coordinate information of all knowledge point target areas, establishing a connection relationship when the distance between any two area bounding boxes is less than a distance threshold, and counting the total number of the connection relationships as a topological relationship; The speech waveform within the set time before the pause operation is intercepted and converted into text, and the frequency of occurrence of predefined query keywords is counted as semantic features; The spatial distribution characteristics, geometric attributes, topological relationships and semantic features are input into a weighted calculation model to perform comprehensive calculations to generate a query confidence.
[0013] Optionally, the decision routing is performed based on the quantified level of the question confidence, generating instant question answering content when a direct response threshold is reached, and activating a dynamic questioning engine when the direct response threshold is not reached, including: Compare the confidence level of the query with the preset direct response threshold; When the confidence level of the question is greater than or equal to the direct response threshold, the pre-stored parsed text corresponding to the target area of the current knowledge point is retrieved to generate instant answer content; When the confidence level of a question is less than the direct response threshold, the dynamic questioning engine is activated and the questioning option generation operation is executed, while the generation of instant question answering content is prohibited.
[0014] Optionally, capturing the progress indicator and the corresponding key frame image when the user performs a video pause operation, and extracting spatial coordinate information representing the target area of the knowledge point in the key frame image, includes: When the video playback engine receives a pause command, it reads the current time value on the video timeline as a progress indicator; Calling a frame buffer module of a video decoder to output a static image precisely aligned with the progress indicator as a key frame picture; Identify independent blocks in the key frame image that meet the characteristics of text cluster areas, formula cluster areas or graphic cluster areas as target areas of knowledge points, and simultaneously generate spatial coordinate information including the positioning coordinates, bounding box size and center point coordinates of each target area.
[0015] In a second aspect, this application provides an AI question-answering system based on the association of video keyframes and progress, including: An acquisition module is used to capture the progress indicator and the corresponding key frame when the user performs a video pause operation, and extract the spatial coordinate information representing the target area of the knowledge point in the key frame; a generation module, configured to capture touch trajectory data satisfying a pressure trigger condition via a touch screen, generate a spatial association map of touch hotspots and knowledge points based on the progress indicator and the spatial coordinate information, and generate a focus area marker based on the resident features of the touch trajectory data; A fusion module, configured to fuse the spatial distribution characteristics of the spatial association map, the geometric properties of the focus region markers, the topological relationship of the spatial coordinate information, and the semantic features of the associated speech segment before the pause operation to generate a question confidence; a decision module, configured to make a decision routing based on the quantified level of the confidence of the question, generate instant question answering content when a direct response threshold is reached, and activate a dynamic questioning engine when the direct response threshold is not reached; A parsing module, configured to parse the semantic types of knowledge points in the spatial association graph through the dynamic questioning engine, so as to match the subject-specific questioning strategy library and generate a set of questioning options; The output module is used to generate question answering content including a core answer layer and an associated knowledge expansion layer based on the user's selection operation of an option in the question option set.
[0016] In an example of the present application, when a user pauses a video, a progress indicator and a corresponding keyframe image are captured, and spatial coordinate information representing a target area of a knowledge point in the keyframe image is extracted. Touch trajectory data that meets a pressure trigger condition is captured via a touch screen, and a spatial association map of touch hotspots and knowledge points is generated by combining the progress indicator and the spatial coordinate information. A focus area marker is generated based on the resident features of the touch trajectory data. A question confidence level is generated by integrating the spatial distribution characteristics of the spatial association map, the geometric properties of the focus area marker, the topological relationship of the spatial coordinate information, and the semantic features of the associated voice segment before the pause operation. Decision routing is performed based on the quantified level of the question confidence level, and instant question-answering content is generated when a direct response threshold is reached. A dynamic questioning engine is activated when the direct response threshold is not reached. The semantic type of the knowledge point in the spatial association map is parsed by the dynamic questioning engine to generate a set of question options based on a subject-specific questioning strategy library. Based on the user's selection of an option in the set of question options, question-answering content comprising a core answer layer and an associated knowledge extension layer is generated.
[0017] The technical solution of this application has the following beneficial effects: This application captures the progress mark and key frame knowledge point coordinates when the video is paused, combines the touch trajectory to generate a spatial correlation map of touch hot spots and knowledge points and a mark of the focus area, and then dynamically generates the question confidence by integrating the spatial distribution of the map, the geometric properties of the mark, the coordinate topological relationship and the speech semantic features before the pause; based on the confidence level, intelligent decision routing is implemented, and the question and answer content is directly output when the threshold is reached. If it is not reached, the dynamic questioning engine is activated to generate question options by parsing the semantic type of the knowledge point and matching the subject strategy library; finally, based on the user's choice, layered question and answer content including the core answer layer and the related knowledge extension layer is output, thereby realizing precise question positioning, adaptive interactive questioning and knowledge structured answers under multimodal information fusion, and significantly improving the accuracy of video scene question answering and user experience.
[0018] The dynamic questioning engine further extracts knowledge point regions corresponding to strongly associated markers in the spatial association map and categorizes their content into formula, concept, or chart semantic types. It then matches the subject-specific questioning strategy library, invoking the formula derivation steps, concept definition comparison, or chart data interpretation option sets as basic questioning options. Finally, it dynamically adjusts the priority of each option in the basic options based on the average duration of the touch point, generating a final, optimized set of questioning options. This solution achieves precise customization of subject-related questioning strategies based on the semantic type of the knowledge point, and intelligently adjusts option priorities by analyzing user touch behavior, thereby generating a highly scenario-specific set of questioning options that aligns with the user's potential questioning tendencies. This significantly improves the relevance of questioning and user selection efficiency, effectively guiding users to quickly clarify core questions.
[0019] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A flowchart of an AI question answering method based on the association of video keyframes and progress provided by the present application is shown; Figure 2 A scene diagram showing an AI question-answering method based on the association of video keyframes and progress provided by the present application is shown; Figure 3 A structural diagram of an AI question-answering system based on the association of video key frames and progress provided by the present application is shown. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0023] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.
[0024] Research shows that the AI question-answering technology of current smart education video platforms generally relies on a keyword matching mechanism for subtitles and text near the pause point. This solution has fundamental flaws: it seriously ignores the core visual information in the key frames of the video and the clear intentions expressed by the user through touch interaction, resulting in an inability to accurately locate the actual object of the question in the picture; at the same time, the solution is essentially a passive response mode and lacks the ability to dynamically evaluate the confidence of the user's question; in addition, the understanding of the semantic coherence of the speech before the pause and the spatial topological relationship of the picture elements are separated, making it difficult to support multimodal correlation analysis of complex knowledge points. The above defects together lead to question-answering responses often deviating from the focus of the user's real question, answering broad or incorrect answers, and the interactive experience being stiff.
[0025] In response to the above problems, this application proposes an AI question-answering method based on the association between video keyframes and progress. The core of this method is: at the moment the user pauses, the progress mark, the spatial coordinates of the keyframe knowledge points and the touch trajectory data that meets the pressure conditions are synchronously captured, and the spatial association map of the touch hot zone and the knowledge point and the focus area markers are generated by fusion; then, the spatial distribution characteristics of the map, the geometric properties of the markers, the coordinate topological relationship and the semantic features of the voice segment before the pause are combined to dynamically generate the confidence level of the multimodal fusion question. Intelligent decision routing is performed based on the confidence quantification value. When the threshold is reached, accurate answers are directly generated. When it is not reached, the dynamic questioning engine is activated: the engine parses the semantic type of the knowledge point in the map, matches the subject-specific strategy library to generate a customized question option set, and outputs layered content containing the core answer layer and the associated knowledge extension layer after the user selects. This method fundamentally solves the three major pain points of existing technologies: first, by fusing spatial coordinates with touch trajectories, it accurately captures the visual focus and interaction intentions in the picture, eliminating "visual blind spots"; second, based on confidence routing and dynamic questioning mechanisms, it achieves a transition from passive response to active guided interaction, effectively clarifying ambiguous questions; third, by integrating multimodal context, it supports the associative understanding and structured answers of complex knowledge points, significantly improving the accuracy of answering questions and user experience.
[0026] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0027] Figure 1 A flowchart of an AI question answering method based on the association of video keyframes and progress is provided for the embodiment of this application, such as Figure 1 As shown, the method includes: 101. When the user performs a video pause operation, a progress indicator and a corresponding key frame image are captured, and spatial coordinate information representing a target area of a knowledge point in the key frame image is extracted; Optionally, step 101 may specifically include the following steps: 1011. When the video playback engine receives a pause instruction, it reads the current time value on the video timeline as a progress indicator; 1012. Calling a frame buffer module of a video decoder to output a static image precisely aligned with the progress indicator as a key frame; 1013. Identify independent blocks in the key frame image that meet the characteristics of text cluster areas, formula cluster areas, or graphic cluster areas as target areas of knowledge points, and simultaneously generate spatial coordinate information including the positioning coordinates, size, and center point coordinates of the bounding box of each target area.
[0028] In the above scheme, the progress identifier refers to the precise timestamp captured by the system from the video timeline when the user triggers the pause operation during video playback. It is used to uniquely identify the moment when the pause occurs and is the core index associated with the corresponding video content segment. The key frame picture refers to the static image frame that the system locates and extracts from the video stream according to the progress identifier at the moment when the user pauses the video, and which precisely corresponds to the moment. It completely preserves the visual scene seen by the user at the moment of pause. The knowledge point target area refers to the independent visual block that carries the core teaching information in the key frame picture and is identified by the system through analysis. These areas are the physical carriers of the user's potential question points. Spatial coordinate information refers to the data set used to quantitatively describe the specific position and range of the knowledge point target area in the key frame picture. It accurately depicts the geometric properties of the target area and provides a basis for the subsequent analysis of the spatial relationship between the user interaction focus and the knowledge point.
[0029] In an embodiment of the present application, first, through step 1011, when the user triggers a pause operation through the video playback interface, the system will monitor and capture this pause event. Subsequently, the application program interface API of the video playback engine is called to read the precise position value of the current video playback head on the timeline. This value, usually in milliseconds ms, is recorded by the system and defined as a progress identifier. For example, when the user clicks the pause button or presses the space bar through the built-in player of the APP, the system captures the pause event and captures the precise position value of the current position on the timeline 00:02:05.670 as a progress identifier. This progress identifier is the core basis for locating the video content in the subsequent steps.
[0030] Next, after obtaining the progress indicator generated in step 1011 through step 1012, the system invokes the frame positioning function provided by the underlying video decoder and passes the progress indicator as an input parameter to the decoder. The decoder then searches for the closest keyframe around that time point based on the video's encoding structure. After finding the corresponding keyframe, the decoder extracts the complete static image data corresponding to that keyframe from the keyframe buffer. This static image data is the keyframe image that exactly corresponds to the content at the moment the user paused. For example, the progress indicator 00:02:05.670 may actually correspond to frame 3012 of the video, and the decoder will output the image of frame 3012.
[0031] Finally, after obtaining the key frame image output in step 1012 through step 1013, the image is analyzed using a pre-trained target detection model. The target detection model is designed to identify specific visual blocks in the image that carry knowledge content. The target detection model scans the entire image and outputs all detected candidate blocks, their category confidences, and bounding box coordinates. Next, the system applies a rule engine to filter and screen all blocks detected by the model. Blocks containing dense, lined text content, and where the width of the text content is generally significantly greater than the height, are screened as text clusters; blocks containing a large number of mathematical symbols, operators, and variables are screened as formula clusters; blocks containing clear geometric figures, data charts, or schematic diagrams are screened as graphic clusters. Blocks identified by the model and meeting at least one of the above-mentioned feature rules are confirmed as target areas of knowledge points. And for the target area of each knowledge point, its bounding box information is calculated: the coordinates of the upper left corner pixel of the target area bounding box are obtained. and the coordinates of the lower right pixel As the bounding box positioning coordinates; based on the bounding box positioning coordinate information, calculate the width of the bounding box and height The bounding box size and center point coordinates The bounding box positioning coordinates, bounding box size and center point coordinates are packaged together to form spatial coordinate information describing the position and range of the target area, and the spatial coordinate information set of all target areas is output.
[0032] In a practical application, suppose a user is watching a physics video on the online learning platform "EduLearn" explaining "Newton's Second Law F=ma." When the video reaches the point where the derivation of the formula is explained and a slide containing the formula F=ma and a force analysis diagram appears, the user becomes confused and presses the pause button at 00:08:23.105. The system receives the pause command and immediately records this precise timestamp. The system then calls the platform's video decoding module and locates the keyframe corresponding to the timestamp 00:08:23.105, which is frame 150. The decoder successfully extracts a still image of this frame, which contains the slide containing the formula and force diagram. The system then feeds this slide image into an object detection model. The model identifies several candidate regions in the image: the title bar "Newton's Second Law" at the top of the page, the formula block in the middle, the force analysis block on the right, and the page number at the bottom. The rule engine begins its work: Although the title bar contains text, it's not typically considered a core knowledge point area; page number information is irrelevant and is directly filtered out; the formula block contains numerous mathematical symbols and is identified as a formula cluster; the force analysis diagram, which contains arrows representing force vectors and squares representing objects, is identified as a graphic cluster. These two areas are identified as target knowledge points. Their spatial coordinates are then calculated: the formula block's bounding box has the upper left corner at (120, 80) and the lower right corner at (400, 300), resulting in a width of 280 pixels, a height of 220 pixels, and a center at (260, 190); the force diagram's bounding box is (450, 150, 800, 400), with a width of 350 pixels, a height of 250 pixels, and a center at (625, 275).
[0033] The above-mentioned 101 overall solution realizes the automatic and precise positioning of the knowledge point area at the moment of pause: through the strict alignment of timestamps and key frames, it ensures that the captured image is consistent with the user's question scene; combined with target detection and rule filtering, it intelligently identifies core knowledge blocks and quantifies their spatial position, laying a data foundation for the subsequent integration of touch interaction and semantic analysis, avoiding the positioning deviation problem caused by relying on manual labeling or text matching in traditional solutions.
[0034] 102. Capturing touch trajectory data that meets the pressure trigger condition through the touch screen, generating a spatial correlation map of touch hotspots and knowledge points based on the progress indicator and the spatial coordinate information, and generating a focus area mark based on the resident features of the touch trajectory data; Optionally, step 102 may specifically include the following steps: 1021. Monitor the touch screen pressure sensor data. When the pressure value continuously exceeds the preset pressure threshold, record the touch point coordinate sequence and the duration of each touch point to form touch trajectory data. 1022. Define a closed polygon formed by continuous touch points as a touch hotspot, compare the spatial coordinate information of the touch hotspot with the target area of each knowledge point, and generate a spatial association map of the touch hotspot and the knowledge point based on the comparison results; Among them, step 1022 may specifically include the following processes: when the closed polygon formed by the touch hot zone completely covers the spatial coordinate information of a single knowledge point target area, a strong association mark between the touch hot zone and the corresponding knowledge point is established; when the closed polygon formed by the touch hot zone simultaneously covers the spatial coordinate information of multiple knowledge point target areas, a weak association mark between the touch hot zone and the corresponding multiple knowledge points is established; the strong association marks and the weak association marks are aggregated to generate a spatial association map containing a set of touch hot zone and knowledge point association relationship marks.
[0035] 1023. Extract touch points whose duration exceeds the dwell threshold from the touch trajectory data as dwell feature points, and generate a circular mark with a radius proportional to the duration as a focus area mark, with the coordinates of each dwell feature point as the center.
[0036] In the above scheme, the pressure trigger condition refers to the lower limit of the touch screen pressure sensor value preset by the system. When the pressure value generated by the user's touch operation continuously exceeds this lower limit, the system determines that the touch is a valid interactive behavior rather than an unintentional touch, which is used to filter out interference signals. Touch trajectory data refers to the dynamic record information generated by the user's touch behavior that meets the pressure trigger condition, including the continuous coordinate sequence of the touch point on the screen and the duration of each touch point being pressed continuously, which fully depicts the screen path that the user actively marks or pays attention to. A touch hot zone refers to a closed polygonal area formed by connecting the end to end of a continuous touch point coordinate sequence. This area intuitively represents the range of visual content that the user intends to focus on in the screen space. A spatial association map refers to a structured data model that calculates the inclusion relationship between the spatial coordinate information of the touch hot zone and the target area of the knowledge point, establishes a strong association mark or weak association mark between the touch hot zone and the knowledge point, and forms a mapping map that reflects the correspondence between the focus of the user's question and the position of the knowledge content. A dwell feature refers to a specific touch point behavior pattern identified from touch trajectory data. This pattern is characterized by a sustained press duration on a single touch point exceeding a preset dwell threshold, reflecting the user's deep attention to the corresponding location on the screen. A focus area marker is a circular visual identifier generated based on this dwell feature, used to spatially quantify the intensity and extent of a user's attention to a specific location.
[0037] In the embodiment of the present application, first, step 1021 continuously monitors the user's operating pressure value through the pressure sensor built into the touch screen. When it is detected that the pressure value of a touch point continues to exceed the preset threshold, the system begins to record the coordinate position (x, y) of the touch point and its duration from the start to the end of the press. As the user moves on the screen, the system will continuously capture multiple touch points that meet the pressure conditions to form a touch trajectory data set containing a coordinate sequence and duration. For example, when the user draws a circle on the screen with his finger, the system captures five consecutive points: the starting point (100, 100) is pressed for 1.0 second, moves to (100, 200) and presses for 1.1 seconds, moves to (200, 200) and presses for 0.9 seconds, moves to (200, 100) and presses for 1.2 seconds, and returns to the starting point (100, 100) and presses for 1.0 second, ultimately forming a complete trajectory data set containing five sets of coordinates and durations.
[0038] Next, through step 1022, the continuous coordinate points in the touch trajectory data outputted in step 1021 are connected in sequence, and a closed polygonal area is constructed by the convex hull algorithm; then, the spatial coordinate information of the polygon and the target area of the knowledge point are compared, and the coverage status of the polygon on the bounding box of each knowledge point is calculated by the ray method. When all the vertices of the bounding box of a knowledge point are located inside the polygon, a strong association mark between the touch hot zone and the corresponding knowledge point is generated; when the polygon only covers part of the bounding boxes of multiple knowledge points, a weak association mark between the touch hot zone and the set of related knowledge points is generated; finally, all the marks are integrated to form a structured spatial association map, which clearly records the mapping relationship between the user interaction area and the knowledge content. For example, if the user circles to form a pentagon, the spatial coordinate information of the pentagon and the target area of the knowledge point are compared. If the user's circled range completely covers the four corner points of the formula area (50, 100, 300, 300), a strong association mark is generated between the touch hot zone and the corresponding knowledge point. If the circled range includes both the upper left corner of the formula area and the lower right corner of the chart area, a weak association mark is generated between the touch hot zone and the set of related knowledge points, marking hot zone 1 as associated with [formula area + chart area].
[0039] Finally, the touch trajectory data output by scanning in step 1023 is filtered to select touch points whose duration exceeds the preset dwell threshold as dwell feature points; and for each feature point, with its coordinates (x, y) as the center of the circle, according to the formula: , dynamically calculates the circular area; ultimately, a circular focus area marker is generated, centered at the feature point and with a radius proportional to the focus duration, to quantify the user's focus intensity on a specific location. For example, if a continuous press at coordinates (500, 200) for 1.2 seconds exceeds the threshold of 0.8 seconds, with (500, 200) as the center and a scaling factor k set to 20 pixels / second, the calculated circular area is a red semi-transparent circle with a radius of 24 pixels and a center at (500, 200).
[0040] In practice, a user pauses while watching a physics tutorial video on an online learning platform. Two core knowledge areas appear on the screen: a formula derivation area located within the coordinate range of 50 horizontally and 100 vertically in the upper left corner of the screen, extending to 300 horizontally and 300 vertically in the lower right corner; and an experimental chart area located within the range of 350 horizontally and 150 vertically to 700 horizontally and 400 vertically. The user applies a pressure of approximately 400 millipasal (mPa) on the screen with their finger, moving it along the boundary of the formula derivation area to form a closed polygonal trajectory roughly enclosing the area. The user then presses a point within the experimental chart area for approximately 1.2 seconds. The system first captures the sequence of consecutive touch points and their respective press durations as the user circles the formula area. It also records the coordinates and duration of the long-press points within the chart area. The system then analyzes the polygonal area formed by the circle and confirms that it completely encompasses the entire boundary of the formula derivation area, thereby establishing a strong association between the touch hotspot and the formula derivation knowledge point. At the same time, the system detected that the long press on a point in the chart area lasted longer than the 0.8-second dwell threshold. It then generated a circular marker with a radius of approximately 24 pixels, centered on the coordinates of the long press point and proportional to the 1.2-second duration, to indicate the user's deep attention area at that location. Ultimately, the system integrated this information to generate a spatial correlation map showing a strong correlation between the touch hotspot and the formula derivation area, as well as a circular attention area marker located in the experimental chart area.
[0041] The above-mentioned 102 overall solution captures user interaction intentions through pressure touch trajectories, generates a spatial correlation map between touch hot spots and knowledge points, and dynamically constructs attention area markers based on dwell time, thereby achieving precise spatial positioning of the focus of questions, visual mapping of user intentions, and quantitative expression of attention intensity. It provides structured spatial relationship data support for subsequent multimodal fusion analysis, significantly improving the accuracy of question understanding and the efficiency of interaction guidance.
[0042] 103. Generate a question confidence by integrating the spatial distribution characteristics of the spatial association map, the geometric properties of the focus region marker, the topological relationship of the spatial coordinate information, and the semantic features of the associated speech segment before the pause operation; Optionally, step 103 may specifically include the following steps: 1031. Counting the ratio of the number of strongly associated markers to the total number of markers in the spatial association map as a spatial distribution characteristic; 1032. Calculate the straight-line distance between the center coordinates of each focus area mark and the center coordinates of the nearest knowledge point target area as a geometric attribute; 1033. Traverse the spatial coordinate information of all knowledge point target areas, establish a connection relationship when the distance between any two area bounding boxes is less than a distance threshold, and count the total number of the connection relationships as a topological relationship; 1034. Intercepting the speech waveform within the set time period before the pause operation and converting it into text, and counting the frequency of occurrence of predefined query keywords as semantic features; 1035. Input the spatial distribution characteristics, geometric attributes, topological relationships and semantic features into a weighted calculation model to perform comprehensive calculations to generate a query confidence level.
[0043] In the above scheme, spatial distribution characteristics refer to quantitative indicators that reflect the concentration of user interaction focus, characterizing the degree of user concentration on core knowledge points, and can be used to evaluate the clarity of question intent. Geometric properties refer to physical quantities that measure the degree of deviation between the user's deep focus point and the core position of the knowledge point. The smaller the distance value, the closer the focus point is to the knowledge core. Topological relationship refers to the connection network characteristics that describe the spatial proximity between the target areas of knowledge points. The larger the value, the closer the knowledge points are related. Semantic features refer to question tendency clues extracted from the associated voice segment before the pause. The higher the frequency, the more significant the question intent in the voice. Question confidence refers to a probabilistic score generated by combining the above four types of features to evaluate the clarity of the user's question. The larger the value, the more confident the system is in its judgment of the user's question point.
[0044] In the embodiment of the present application, first, step 1031 is used to read the spatial correlation map generated in step 102, extract all the marker types from it, and count the number of strong correlation markers as , and the number of weakly associated markers is recorded as Then calculate the spatial distribution characteristic index, the calculation formula is as follows: For example, if the map contains 2 strong correlation markers and 1 weak correlation marker, then , this value reflects the user's concentration on core knowledge points.
[0045] Then, in step 1032, for each area of interest mark generated in step 102, the center coordinates of the circle are obtained. Then traverse the spatial coordinate information of all knowledge point target areas and extract the center point coordinates of each knowledge point bounding box Calculate the Euclidean distance from the center of the circle to the center of each knowledge point , and select the minimum value as the geometric attribute value For example, the distance from the center of the circle (200,150) to the center of the nearest knowledge point (180,160) Pixels, the smaller the value, the more consistent the focus is with the knowledge core.
[0046] Then load the bounding box coordinates of all knowledge point target areas, traverse any two area combinations to calculate the minimum bounding box spacing: first calculate the horizontal spacing component and the vertical spacing component, and take the maximum value of the two as the actual spacing , negative values are considered 0; when When a connection relationship is established, the total number of connections that meet the conditions is finally counted as the topological relationship value For example, in area A (50, 100, 300, 300) and area B (350, 150, 700, 400), the horizontal distance between the right boundary 300 of area A and the left boundary 350 of area B is 350-300=50 pixels, which is equal to the threshold and is counted as one connection.
[0047] The original speech waveform data of a fixed length before the pause operation is intercepted and converted into text by calling the speech recognition engine. Then the predefined question keyword library Q = {"why", "how", "whether", ...} is matched and the frequency of occurrence of these words in the text is counted as the semantic feature value. For example, if the recognized text is "Why does this formula need to be transformed?" and the keyword "why" appears once, then .
[0048] Finally, step 1035 converts the four-dimensional features generated in the previous step into spatial distribution characteristics. , geometric properties , topological relationship and semantic features Input the pre-trained weighted model to perform fusion calculation. The calculation formula is as follows: , where the weight , attenuation coefficient k=0.02. The final output is the confidence level of the question For subsequent routing decisions. For example, spatial distribution characteristics 0.67, geometric properties 22.36 pixels, topological relationship 1 and semantic features is 1, and the confidence level of the query is obtained by performing the fusion calculation. .
[0049] In actual applications, when the user pauses the physics video on the online learning platform, the system detects that the spatial correlation map contains two strong correlation marks, which completely cover the formula derivation area and the experimental data area in the screen respectively, and another weak correlation mark involves the text description area. It is calculated that the proportion of strong correlation marks is about 0.67, which is two-thirds. At the same time, the circular area of interest generated by the user is identified, with its center located at a position of 200 pixels horizontally and 150 pixels vertically on the screen, and a straight-line distance of about 22.36 pixels from the center point of the nearest formula area, which is 180 pixels horizontally and 160 pixels vertically. In the knowledge point area, the boundary of the formula area ranges from 50 pixels horizontally and 100 pixels vertically in the upper left corner to 300 pixels horizontally and 300 pixels vertically in the lower right corner, and the boundary of the chart area ranges from 350 pixels horizontally and 150 pixels vertically to 700 pixels horizontally and 400 pixels vertically. The minimum horizontal spacing between the two areas is 50 pixels, which is equal to the preset threshold, so a topological connection relationship is established, and the total number of connections is recorded as 1. The 5-second speech segment before the pause is intercepted and converted into the text "Why does this formula need to be converted" after recognition. The question word "why" triggers the predefined keyword library with a statistical frequency of 1. The question confidence is calculated using a weighted fusion model: the spatial distribution characteristic value is 0.67, the geometric attribute distance value 22.36 pixels is subjected to the attenuation function e^(-0.02×22.36)≈0.64, the topological relationship value 1 is subjected to the hyperbolic tangent function tanh(1)≈0.76, and the semantic feature value is 1. The weight coefficients are linearly superimposed according to 0.3, 0.3, 0.2, and 0.2: 0.3×0.67+0.3×0.64+0.2×0.76+0.2×1=0.72, and the question confidence of 0.72 is finally generated to trigger an instant question answering response.
[0050] The above-mentioned 103 overall solution innovatively integrates four-dimensional information: spatial interaction intensity, visual attention accuracy, knowledge point association density, and voice question clues. It generates a comprehensive question confidence through a weighted model, realizes the quantitative evaluation and unified expression of multimodal intentions, provides an objective basis for subsequent decision-making routing, and significantly improves the comprehensiveness and reliability of question recognition.
[0051] 104. Decision routing is performed based on the quantified level of the question confidence, and instant question answering content is generated when a direct response threshold is reached, and a dynamic questioning engine is activated when the direct response threshold is not reached; Optionally, step 104 may specifically include the following steps: 1041. Compare the confidence level of the question with the preset direct response threshold; 1042. When the confidence level of the question is greater than or equal to the direct response threshold, the pre-stored parsed text corresponding to the target area of the current knowledge point is retrieved to generate instant answer content; 1043. When the confidence level of a question is less than the direct response threshold, the dynamic questioning engine is activated and the questioning option generation operation is executed, while the generation of instant question answering content is prohibited.
[0052] In the above scheme, the direct response threshold refers to the preset critical value for determining the clarity of the question. When the question confidence calculated by the system reaches or exceeds this value, it indicates that the user's question point is highly clear, which is used to trigger the direct question-answering mechanism. Instant question-answering content refers to the pre-stored structured answer text that the system retrieves based on the current knowledge point target area identifier when the question confidence meets the threshold condition, and is used to directly respond to user questions. The dynamic question-following engine refers to an interactive decision-making module that is activated when the question confidence is lower than the threshold. It generates a set of targeted question options by parsing the user's potential question context to guide the user to clarify ambiguous intentions. The generation prohibition mechanism refers to a control instruction that is triggered synchronously when the dynamic question-following engine is activated. It forcibly closes the direct question-answering content output channel to ensure that the system maintains only a single interaction path in scenarios with ambiguous intentions to avoid information conflicts.
[0053] In this embodiment of the present application, step 1041 obtains a question confidence value (e.g., 0.72) and compares it to a preset direct response threshold. This process is implemented using a numerical comparator: if the confidence value is greater than or equal to the threshold, a direct response flag is triggered; if the confidence value is less than the threshold, a follow-up question flag is triggered. For example, in a physics teaching scenario, if the user's question confidence value of 0.75 for an electromagnetic formula exceeds the threshold of 0.6, the system generates a direct response flag.
[0054] When the question confidence level is greater than or equal to the threshold, triggering the direct response flag, the system performs the following chain operations at step 1042: extracting the current main area identifier from the knowledge point target area information; then using this identifier as the index key to query the pre-stored parsed text database; and finally packaging the retrieved structured text into instant question answering content. For example, the pre-stored text retrieved is: "Faraday's law describes the relationship between the induced electromotive force and the rate of change of magnetic flux: , commonly seen in the working principle of generators...", directly output to the user interface.
[0055] When the question confidence level falls below the threshold, triggering the follow-up question flag, the system performs two operations in step 1043: first, activating the dynamic question engine, initializing the engine's core processor, and invoking the algorithm for generating follow-up question options; second, initiating a generation-blocking mechanism, forcibly locking the data transmission channel of the direct question-answering interface. For example, in a math video scenario, if the user's confidence level for a question about integral application is only 0.52 < 0.6, the system activates the follow-up question engine to generate options ["Do you need a sample problem?", "Or a concept explanation?"] and simultaneously blocks access to the pre-stored text library, ensuring that only follow-up question options are displayed on the interface.
[0056] In practice, when a user pauses while watching a circuit analysis video, the system calculates the question confidence. If the question confidence is 0.68, which is greater than the threshold of 0.6, the direct response flag is triggered. The system then searches the database for the current knowledge point ID "ohm_law" to retrieve the pre-stored parsed text: "In the Ohm's Law formula V=IR, V represents voltage..." and outputs it. If the question confidence is 0.55, which is less than the threshold of 0.6, the follow-up question flag is triggered, activating the follow-up question engine generation options ("Are you asking about the formula symbols?" or "Are you asking about the experimental steps?"), while disabling the direct question answering function.
[0057] The above-mentioned 104 overall solution achieves precise decision routing through intelligent comparison of confidence and thresholds: when the confidence level is high, structured question-answering content is directly output to ensure efficient response; when the confidence level is low, the follow-up question engine is activated to guide users to clarify their questions, while prohibiting the output of conflicting information, significantly improving the rigor of the interaction logic and the smoothness of the user experience.
[0058] 105. Parsing the semantic types of the knowledge points in the spatial association graph by the dynamic questioning engine to generate a set of questioning options by matching them with a subject-specific questioning strategy library; Optionally, step 105 may specifically include the following steps: 1051. Extract target regions of knowledge points corresponding to strong correlation marks in the spatial correlation map, and classify the target regions into formula, concept, or chart semantic types according to the content of the target regions; 1052. Match the subject-specific questioning strategy library according to the semantic type, and call the step derivation option set corresponding to the formula class, the definition comparison option set corresponding to the concept class, or the data interpretation option set corresponding to the chart class as the basic questioning option set; 1053. Based on the average duration of all touch points in the touch trajectory data, adjust the priority order of the options in the basic follow-up question option set, and arrange the high-frequency question options at the top to generate a final follow-up question option set.
[0059] Among them, step 1053 may specifically include the following processes: calculating the arithmetic mean of the durations of all touch points in the touch trajectory data as the average duration, and counting the historical number of selections of each option in the basic follow-up question option set from the historical question and answer database; rearranging the order of options in the basic follow-up question option set in descending order according to the historical number of selections, and generating an adjusted set with high-frequency question options at the top; when the average duration exceeds the interaction depth threshold, adding a deep analysis option to the rearranged basic follow-up question option set to form a final follow-up question option set; when the average duration does not exceed the interaction depth threshold, directly using the rearranged basic follow-up question option set as the final follow-up question option set.
[0060] In the above scheme, the semantic type of knowledge points refers to a classification system divided according to the visual content of the target area, including formula class, concept class, and chart class. This classification is used to accurately match the subject questioning strategy. The subject-specific questioning strategy library refers to a set of questioning option templates pre-built according to subject areas and semantic types, which is used to generate scenario-based questioning content. The average duration refers to the arithmetic mean of the duration of all touch points pressed in the touch track, reflecting the overall interaction depth of the user. The interaction depth threshold refers to the critical duration value that triggers the addition of deep analysis options. When the average duration exceeds this value, advanced entries are added to the questioning options. Deep analysis options refer to advanced questioning items added to the end of the basic option set to meet deep interaction needs. Pinning high-frequency question options means counting the number of times each option has been selected based on the historical database, and re-arranging the order of basic options in descending order by the number of times, in order to improve user selection efficiency.
[0061] In the embodiment of the present application, first, all target areas of knowledge points with strong association marks are extracted from the spatial association map through step 1051, and then the semantic type of each area is classified: the OCR engine is called to identify the text content in the area, and the symbol detection algorithm is run to scan mathematical symbols and graphic elements. If the density of mathematical symbols detected exceeds the threshold, it is marked as a formula type; if keywords such as "definition" and "theorem" are recognized and there are no dense symbols, it is marked as a concept type; if the proportion of visual elements detected exceeds 50%, it is marked as a chart type. For example, if the area contains the expression " ” is classified as a formula.
[0062] Then, based on the semantic type output by 1051, the system queries the subject-specific questioning strategy library to match the corresponding option template. First, determine the video subject label, and then call the corresponding option template according to the semantic type. The formula type activates ["derivation steps", "symbol meaning", "application scenario"], the concept type activates ["definition comparison", "example analysis", "common misunderstandings"], and the chart type activates ["data interpretation", "graphing principle", "trend analysis"]. For example, the mathematics subject + formula semantics outputs a basic questioning option set. ={"1. Derivation steps","2. Symbol meaning","3. Application scenarios"}.
[0063] Finally, combined with the touch trajectory data, a third-order optimization is performed on the basic question option set: the arithmetic mean of the duration of all touch points in the touch trajectory data is calculated using the following formula: , is the duration of the touch point, n is the number of touch points, is the sum of the duration of n touch points; then query the historical database to obtain The number of times each option has been selected in the past is used to re-arrange the options in descending order. For example, "Derivation steps": 120 times, "Symbol meaning": 80 times. ={"1. Derivation steps","2. Symbol meaning","3. Application scenario"}; Finally, determine whether the average duration exceeds the interaction depth threshold. If the average duration exceeds the interaction depth threshold, add an option. If the average duration does not exceed the interaction depth threshold, directly output , such as the average duration , and finally output the set of question options ={"Derivation steps","Symbol meaning","Application scenario"}; if the average duration Then append ={"Derivation steps","Symbol meaning","Application scenarios","Step-by-step explanation"}.
[0064] In practice, when watching a calculus video, a user circles an integral formula area, creating a strong association mark. The system uses OCR to identify the formula symbols and classifies them as formula semantic types. The system then matches the mathematical strategy library to generate a basic set of options: ["Derivation steps," "Symbol meaning," "Application scenarios"]. Based on the touch track durations of 1.4 seconds, 1.0 seconds, and 1.8 seconds at three points, the average duration is calculated to be 1.4 seconds. A historical data query shows that "Derivation steps" was selected 105 times and "Symbol meaning" 62 times, rearranging the options to ["Derivation steps," "Symbol meaning," "Application scenarios"]. Because the average duration of 1.4 seconds does not exceed the 1.5-second threshold, the final set of follow-up questions is ["1. Derivation steps," "2. Symbol meaning," "3. Application scenarios"]. If the durations were changed to 2.0 seconds, 1.5 seconds, and 1.7 seconds, the average duration of 1.73 seconds would exceed the threshold, and the option "4. Step-by-step explanation" would be added.
[0065] The overall solution of 105 mentioned above generates subject-customized follow-up options through semantic type matching, and dynamically adjusts the option priority and depth based on the average touch duration: based on historical selection data, high-frequency question items are placed at the top to improve selection efficiency, and when the interaction duration exceeds the threshold, in-depth analysis options are added to achieve precise adaptation of follow-up content to user behavior characteristics, significantly improving the pertinence of question clarification and the depth of interaction.
[0066] 106. Based on the user's selection operation of an option in the set of follow-up question options, generate question answering content including a core answer layer and an associated knowledge expansion layer.
[0067] Optionally, step 106 may specifically include the following steps: 1061. In response to a user clicking on a specific option in the query option set, obtaining identification information of the specific option; 1062. Extracting standard parsing content corresponding to the knowledge point from a pre-stored parsing text library as a core answer layer according to the identification information; 1063. Search the video metadata database based on the progress indicator and the indicator information to obtain background knowledge modules and advanced knowledge modules associated with the current knowledge point, and combine them to form an associated knowledge expansion layer; 1064. Place the core answer layer at the beginning of the question-answering content, and attach the related knowledge extension layer to the end of the core answer layer in an expandable hierarchical structure to generate complete question-answering content.
[0068] In the above scheme, the specific option identification information refers to the unique code captured by the system when the user clicks on the follow-up option, which includes the index key value of the option in the policy library, and is used to accurately locate the pre-stored analysis content. The core answer layer refers to the standard answer content extracted from the pre-stored analysis text library based on the identification information, including the direct analysis text for the user's selected question points, which is used to immediately answer the core questions. The associated knowledge expansion layer refers to the supplementary knowledge module obtained by retrieving the video metadata database based on the progress identifier and option identifier, including the background knowledge module and the advanced knowledge module, which is used to provide context extension and in-depth expansion. The expandable hierarchical structure refers to the front-end interaction design scheme, which encapsulates the associated knowledge expansion layer as a folding control that is collapsed by default, to balance information density and user cognitive load.
[0069] In this embodiment, first, in step 1061, when a user clicks a specific option on the query option set interface, the system captures this interaction through the front-end event monitoring module and parses the unique identification information of the clicked control. This identification information serves as the core index key for subsequent content retrieval. For example, if the user selects the "Symbol Meaning" option, the ID="symbol_meaning" is captured.
[0070] Next, in step 1062, the specific option identification information obtained in step 1061 is used as a query key to search the pre-stored parsed text library structured SQL database: the SQL statement SELECT content FROM knowledge_responses WHERE option_id = 'symbol_meaning' is executed to extract the standard parsed text for the matching field. This text is formatted and used as the core solution layer. For example, a 500-word parsing of the integral symbol "∫" is retrieved: "The integral symbol comes from the Latin summa... which represents an infinite accumulation process."
[0071] Then, in step 1063, the progress indicator and the identification information of the options are integrated to initiate a joint query to the video metadata database: first, the background knowledge module is searched to obtain the historical background and application scenarios, and then the advanced knowledge module is searched to obtain the extended content. For example, if the progress indicator is 00:15:30.500 and the identification information is "derivation_steps", the search is executed. AND option_id = 'derivation_steps' returns a 150-word text about the creation of the Newton-Leibniz formula. Then, searching the advanced knowledge module (SELECT advanced FROM metadata...) returns a 200-word text containing a formula about the relationship between double integrals and cumulative integrals, obtaining expanded content. Ultimately, the historical background and application scenarios obtained are combined with the expanded content to form a related knowledge expansion layer.
[0072] Finally, the core answer layer generated in step 1062 is placed at the head of the question-answering content, and the associated knowledge expansion layer output in step 1063 is encapsulated as an expandable control: <details>Label wrapping background module and advanced module text, setting <summary> ▼Related knowledge expansion< / summary> As a clickable title, it is initialized to the default collapsed state, forming a complete Q&A content structure of [Core Answer] + [▼Expandable Extension Layer].
[0073] In actual application, after the user selects the "derivation steps" option ID="derivation_steps", the system queries the parsing library based on the progress indicator 00:15:30.500 to obtain the 300-word core solution layer, and jointly queries the metadata library to obtain a 150-word background module on the history of the Newton-Leibniz formula and a double integral formula. A 200-word advanced module; finally assembled into a layered Q&A content with core text at the beginning and expandable controls at the end.
[0074] The above-mentioned 106 overall solutions achieve layered knowledge delivery through structured content assembly. The core answer layer accurately responds to users' immediate questions, and the related extension layer provides explorable background and advanced knowledge. The expandable design balances information density and cognitive load, forming an adaptive question-answering paradigm of "focusing on the core and expanding on demand", significantly improving knowledge transfer efficiency and deep learning support.
[0075] The following is a complete example for steps 101 to 106. Figure 2 As shown in the example, user Xiao Wang, watching a physics course on an online learning platform, presses the pause button at the time 00:12:30.500 in the "Derivation of Newton's Second Law" segment. The system captures this progress indicator, and the decoder outputs the current keyframe: the formula F=ma is displayed on the left side of the slide, and the inclined plane force analysis diagram is on the right. The object detection model identifies two core knowledge areas: the formula area's bounding box has coordinates of 50 horizontal and 100 vertical, extending to the lower right corner by 300 horizontal and 300 vertical. The diagram area's bounding box has coordinates of 350 horizontal and 150 vertical, extending to the lower right corner by 700 horizontal and 400 vertical.
[0076] Xiao Wang used a stylus to circle the formula area and applied a pressure of 480 millipacals, exceeding the preset threshold of 300 millipacals. The system recorded the closed trajectory points, forming a hotspot that completely covered the boundary of the formula area and generated a strong association marker. Simultaneously, he pressed and held for 1.8 seconds at the force diagram coordinate point 600 horizontally and 250 vertically, triggering the resident feature to generate a circular attention marker with a radius of 36 pixels. The system integrated multimodal features to calculate the question confidence score: the spatial distribution characteristic was a unique strong association marker, scoring 1.0; the Euclidean distance between the attention marker's center and the center of the diagram area was approximately 76 pixels; the horizontal spacing of 50 pixels between the formula area and the diagram area boundary met the connection threshold, scoring 1 point for topological relationship; and the semantic feature recognition of the question word "why" in the first 5 seconds of the pause, "Why is mass inversely proportional to acceleration?", scoring 1 point. The weighted model output confidence score was 0.92, exceeding the direct response threshold of 0.6.
[0077] After determining a high confidence level, the system skips the follow-up questioning process and directly generates layered Q&A content: First, based on the formula area identifier, the pre-stored database is searched to obtain the core answer layer "Mass m represents the inertial property of an object, and acceleration a is proportional to the force F..."; then, based on the progress indicator 00:12:30.500, the metadata database is searched to obtain the background module "In 1687, Newton first proposed this law in the Mathematical Principles of Natural Philosophy" and the advanced module "Correction formula of relativity" The final interface is assembled into a front-end interface: the core analytical text is displayed at the top, and an expandable control "▼Related Knowledge Extension" is attached at the end. Clicking it will expand the historical background and expanded content of relativity.
[0078] After quickly understanding the essential properties of F=ma through the core layer, Xiao Wang clicked the expand icon to delve deeper into the evolution of knowledge from classical mechanics to relativity. The system's adaptive mechanisms of multimodal perception, confidence-based decision-making, and hierarchical output enable a closed-loop learning experience: "Precisely respond to core questions and expand knowledge boundaries on demand."
[0079] Figure 3 The present application embodiment provides a structural diagram of an AI question answering system based on the association of video key frames and progress, such as Figure 3 As shown, the system includes: An acquisition module 31 is configured to capture a progress indicator and a corresponding key frame when the user pauses the video, and extract spatial coordinate information representing a target area of a knowledge point in the key frame; a generation module 32 for capturing touch trajectory data that meets the pressure trigger condition through the touch screen, generating a spatial association map of touch hotspots and knowledge points based on the progress indicator and the spatial coordinate information, and generating a focus area marker based on the resident features of the touch trajectory data; A fusion module 33 is configured to fuse the spatial distribution characteristics of the spatial association map, the geometric properties of the focus region marker, the topological relationship of the spatial coordinate information, and the semantic features of the associated speech segment before the pause operation to generate a question confidence level; A decision module 34 is configured to make a routing decision based on the quantified level of the question confidence, generate instant question answering content when a direct response threshold is reached, and activate a dynamic questioning engine when the direct response threshold is not reached; The parsing module 35 is configured to parse the semantic types of the knowledge points in the spatial association graph through the dynamic questioning engine to generate a set of questioning options by matching the subject-specific questioning strategy library; The output module 36 is configured to generate question answering content including a core answer layer and an associated knowledge expansion layer based on the user's selection operation of an option in the set of follow-up question options.
[0080] Figure 3 The AI question answering system based on the association of video key frames and progress can be executed Figure 1 The implementation principle and technical effects of the AI question-answering method based on the association of video keyframes with progress described in the illustrated embodiment will not be elaborated on here. The specific manner in which each module and unit performs operations in the AI question-answering system based on the association of video keyframes with progress in the above embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.< / details>
Claims
1. The AI question answering method based on the association between video key frames and progress is characterized by: include: When the user performs a video pause operation, the progress indicator and the corresponding key frame image are captured, and the spatial coordinate information representing the target area of the knowledge point in the key frame image is extracted; Capturing touch trajectory data that meets the pressure trigger condition through the touch screen, generating a spatial correlation map of touch hotspots and knowledge points based on the progress indicator and the spatial coordinate information, and generating a focus area mark based on the resident features of the touch trajectory data; Generate a question confidence by fusing the spatial distribution characteristics of the spatial association map, the geometric properties of the focus region marker, the topological relationship of the spatial coordinate information, and the semantic features of the associated speech segment before the pause operation; Decision routing is performed based on the quantified level of the question confidence, generating instant question answering content when a direct response threshold is reached, and activating a dynamic questioning engine when the direct response threshold is not reached; Parsing the semantic types of knowledge points in the spatial association graph by the dynamic questioning engine to generate a set of questioning options by matching the subject-specific questioning strategy library; Based on the user's selection operation of an option in the set of follow-up question options, question answering content including a core answer layer and an associated knowledge expansion layer is generated.
2. The method according to claim 1, characterized in that The dynamic questioning engine is used to analyze the semantic types of knowledge points in the spatial association graph to match the subject-specific questioning strategy library to generate a set of questioning options, including: Extracting target areas of knowledge points corresponding to strong correlation marks in the spatial correlation map, and classifying the target areas into formula, concept, or chart semantic types according to the content of the target areas; Matching the subject-specific questioning strategy library according to the semantic type, calling the step derivation option set corresponding to the formula class, the definition comparison option set corresponding to the concept class, or the data interpretation option set corresponding to the chart class as the basic questioning option set; Based on the average duration of all touch points in the touch trajectory data, the priority order of the options in the basic follow-up question option set is adjusted, and the high-frequency question options are placed at the top to generate a final follow-up question option set.
3. The method according to claim 2, characterized in that The step of adjusting the priority order of the options in the basic follow-up question option set based on the average duration of all touch points in the touch trajectory data, and placing the high-frequency question options at the top to generate a final follow-up question option set includes: Calculating the arithmetic mean of the durations of all touch points in the touch trajectory data as the average duration, and counting the number of times each option in the basic follow-up question option set has been historically selected from a historical question-answering database; Rearrange the order of the options in the basic question option set in descending order of the number of times the options were selected historically, and generate an adjusted set with high-frequency question options at the top; When the average duration exceeds the interaction depth threshold, adding a deep analysis option to the rearranged basic question option set to form a final question option set; When the average duration does not exceed the interaction depth threshold, the rearranged basic question option set is directly used as the final question option set.
4. The method according to claim 1, wherein The generating of the question answering content including the core answer layer and the associated knowledge extension layer based on the user's selection operation of the option in the question option set includes: In response to a user clicking on a specific option in the set of follow-up options, obtaining identification information of the specific option; Extracting the standard parsing content of the corresponding knowledge point from the pre-stored parsing text library as the core answer layer according to the identification information; Searching the video metadata database based on the progress indicator and the indicator information to obtain background knowledge modules and advanced knowledge modules associated with the current knowledge point, and combining them to form an associated knowledge expansion layer; The core answer layer is placed at the beginning of the question-answering content, and the related knowledge extension layer is attached to the end of the core answer layer in an expandable hierarchical structure to generate complete question-answering content.
5. The method according to claim 1, wherein The method captures touch trajectory data that meets the pressure trigger condition through the touch screen, generates a spatial correlation map of touch hot areas and knowledge points in combination with the progress indicator and the spatial coordinate information, and generates a focus area mark based on the resident features of the touch trajectory data, including: Monitor the touch screen pressure sensor data. When the pressure value continuously exceeds the preset pressure threshold, record the touch point coordinate sequence and the duration of each touch point to form touch trajectory data. A closed polygon formed by continuous touch points is defined as a touch hot zone, the spatial coordinate information of the touch hot zone is compared with the target area of each knowledge point, and a spatial association map of the touch hot zone and the knowledge point is generated based on the comparison results; Touch points whose duration exceeds a dwell threshold in the touch trajectory data are extracted as dwell feature points, and a circular mark with a radius proportional to the duration is generated as a focus area mark with the coordinates of each dwell feature point as the center.
6. The method according to claim 5, characterized in that The step of comparing the spatial coordinate information of the touch hot zone with the target area of each knowledge point and generating a spatial correlation map of the touch hot zone and the knowledge point according to the comparison result includes: When the closed polygon formed by the touch hot zone completely covers the spatial coordinate information of the target area of a single knowledge point, a strong association mark between the touch hot zone and the corresponding knowledge point is established; When the closed polygon formed by the touch hot zone covers the spatial coordinate information of multiple knowledge point target areas at the same time, a weak association mark is established between the touch hot zone and the corresponding multiple knowledge points; The strong association marks and weak association marks are aggregated to generate a spatial association map including a set of association relationship marks between touch hot areas and knowledge points.
7. The method according to claim 1, characterized in that The generating of the question confidence by fusing the spatial distribution characteristics of the spatial association map, the geometric properties of the focus region marker, the topological relationship of the spatial coordinate information, and the semantic features of the associated speech segment before the pause operation includes: Counting the ratio of the number of strongly associated markers to the total number of markers in the spatial association map as a spatial distribution characteristic; Calculate the straight-line distance between the center coordinates of each focus area mark and the center coordinates of the nearest knowledge point target area as the geometric attribute; Traversing the spatial coordinate information of all knowledge point target areas, establishing a connection relationship when the distance between any two area bounding boxes is less than a distance threshold, and counting the total number of the connection relationships as a topological relationship; The speech waveform within the set time before the pause operation is intercepted and converted into text, and the frequency of occurrence of predefined query keywords is counted as semantic features; The spatial distribution characteristics, geometric attributes, topological relationships and semantic features are input into a weighted calculation model to perform comprehensive calculations to generate a query confidence.
8. The method according to claim 1, characterized in that The decision routing is performed based on the quantified level of the question confidence, generating instant question answering content when a direct response threshold is reached, and activating a dynamic questioning engine when the direct response threshold is not reached, including: Compare the confidence level of the query with the preset direct response threshold; When the confidence level of the question is greater than or equal to the direct response threshold, the pre-stored parsed text corresponding to the target area of the current knowledge point is retrieved to generate instant answer content; When the confidence level of a question is less than the direct response threshold, the dynamic questioning engine is activated and the questioning option generation operation is executed, while the generation of instant question answering content is prohibited.
9. The method according to claim 1, characterized in that The step of capturing the progress indicator and the corresponding key frame image when the user performs a video pause operation, and extracting the spatial coordinate information representing the target area of the knowledge point in the key frame image, includes: When the video playback engine receives a pause command, it reads the current time value on the video timeline as a progress indicator; Calling a frame buffer module of a video decoder to output a static image precisely aligned with the progress indicator as a key frame picture; Identify independent blocks in the key frame image that meet the characteristics of text cluster areas, formula cluster areas or graphic cluster areas as target areas of knowledge points, and simultaneously generate spatial coordinate information including the positioning coordinates, bounding box size and center point coordinates of each target area.
10. The AI question answering system based on the association between video keyframes and progress is characterized by: include: An acquisition module, configured to capture a progress indicator and a corresponding key frame when the user pauses the video, and extract spatial coordinate information representing a target area of a knowledge point in the key frame; a generation module, configured to capture touch trajectory data satisfying a pressure trigger condition via a touch screen, generate a spatial association map of touch hotspots and knowledge points based on the progress indicator and the spatial coordinate information, and generate a focus area marker based on the resident features of the touch trajectory data; A fusion module, configured to fuse the spatial distribution characteristics of the spatial association map, the geometric properties of the focus region markers, the topological relationship of the spatial coordinate information, and the semantic features of the associated speech segment before the pause operation to generate a question confidence; a decision module, configured to make a decision routing based on the quantified level of the confidence of the question, generate instant question answering content when a direct response threshold is reached, and activate a dynamic questioning engine when the direct response threshold is not reached; A parsing module, configured to parse the semantic types of knowledge points in the spatial association graph through the dynamic questioning engine, so as to match the subject-specific questioning strategy library and generate a set of questioning options; The output module is used to generate question answering content including a core answer layer and an associated knowledge expansion layer based on the user's selection operation of an option in the question option set.
Citation Information
Patent Citations
Intelligent classroom question answering interaction system and method
CN113225575A
Knowledge graph-based medicine supply chain question-answering method and system
CN117251543A
Energy data analysis method and system based on large model and proprietary knowledge base
CN120179869A
Water conservancy intelligent question and answer interaction platform and method driven by knowledge graph
CN120196707A
Architecture and processes for computer learning and understanding
US20170371861A1