Paper screen interaction method, device and system
By recognizing the video frame type and audio stream of paper pages, generating echo images or close-up images, and dynamically analyzing user intent, this solves the problems of fixed functions and insufficient understanding of user intent in paper-screen interaction, and achieves real-time, accurate intelligent feedback and improved learning efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CENTURY TAL EDUCATION TECH CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-17
AI Technical Summary
Existing paper-screen interaction methods rely on users to perform specific operations, resulting in limited and fixed functions. They are difficult to understand user intentions in real time and accurately, and to provide effective intelligent feedback.
By identifying the frame type of video frames transmitted from the mobile device, a feedback image or close-up image is generated. Combined with the audio stream analysis of user intent, the system can proactively perceive the user's learning status, dynamically generate targeted intent responses, and avoid additional user operations.
It enables real-time and accurate understanding of user intent and provides effective intelligent feedback without requiring user interaction, significantly improving the efficiency of assisted learning.
Smart Images

Figure CN121879628A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information human-computer interaction technology, and in particular to paper screen interaction methods, devices, and systems. Background Technology
[0002] With the development of intelligent education, paper-screen interaction has become an important demand. Paper-screen interaction refers to the process where users learn using paper materials (such as textbooks, workbooks, and test papers) and receive intelligent feedback assistance such as knowledge explanations and answer analysis from display terminals (such as mobile phones, tablets, and learning machines).
[0003] However, current paper-based screen interaction mainly relies on users manually clicking the screen, making pre-defined gestures, or issuing specific voice commands to achieve pre-set functions such as correction or analysis. This paper-based screen interaction method is too passive, depending on the user to perform specific operations, resulting in limited and fixed functions. It is difficult to understand the user's intentions in real time and accurately, and to provide effective intelligent feedback. Summary of the Invention
[0004] In view of this, embodiments of this application provide methods, apparatus, and systems related to paper screen interaction.
[0005] This application provides a paper screen interaction method, which is applied to a cloud server; the method includes: For each video frame in the video stream transmitted by the mobile terminal, the frame type of the video frame is identified; the frame type is either a first type or a second type; if the frame type of the video frame is a first type, it is indicated that the video frame is to be displayed on the mobile terminal; if the frame type of the video frame is a second type, it is indicated that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance. If the frame type of the video frame is the first type, then the corresponding echo image of the video frame is generated based on the video frame and returned to the mobile terminal for display, and structured data is generated based on the paper page content area in the echo image and stored in the index library. If the video frame is of type two, when the paper page in the video frame is identified as the same page as the paper page in the stored most recently recalled image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the most recently recalled image, the region containing the mapping position is found from the existing index library, a close-up image is generated based on the region, and returned to the mobile terminal for display; the close-up image contains at least knowledge content related to the region where the target identifier is located, generated based on the lecture model; Furthermore, the audio stream in the video stream is transcribed to obtain audio text, and the audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input content. Based on the multimodal input content, the user's intent on the paper page and the intent response are obtained, and the intent response is returned to the mobile terminal for display.
[0006] This application also provides a paper-screen interactive device, which is applied to a cloud server; the device includes: The identification module is used to identify the frame type of each video frame in the video stream transmitted by the mobile terminal; the frame type is a first type or a second type; if the frame type of the video frame is the first type, it indicates that the video frame is to be displayed on the mobile terminal; if the frame type of the video frame is the second type, it indicates that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance. The image processing module is used to generate a corresponding echo image based on the video frame and return it to the mobile terminal for display if the frame type of the video frame is the first type; and to generate structured data based on the paper page content area in the echo image and store it in the index library. If the video frame is of type two, when the paper page in the video frame is identified as the same page as the paper page in the stored most recently recalled image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the most recently recalled image, the region containing the mapping position is found from the existing index library, a close-up image is generated based on the region, and returned to the mobile terminal for display; the close-up image contains at least knowledge content related to the region where the target identifier is located, generated based on the lecture model; The intent parsing module is used to transcribe the audio stream in the video stream to obtain audio text, and use the audio text and video frames of type 1 or type 2 within a specified time period as multimodal input content. Based on the multimodal input content, the module obtains the user's intent on the paper page and the intent response, and returns the intent response to the mobile terminal for display.
[0007] This application also provides a paper-screen interactive system, which includes: a cloud server and a mobile terminal; A cloud server is used to identify the frame type of each video frame in the video stream transmitted from the mobile terminal; the frame type is a first type or a second type; if the frame type of the video frame is the first type, it indicates that the video frame is to be displayed on the mobile terminal; if the frame type of the video frame is the second type, it indicates that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance. If the frame type of the video frame is the first type, then a corresponding echo image is generated based on the video frame, the echo image and its display priority are returned to the mobile terminal, and structured data is generated based on the paper page content area in the echo image and stored in the index library. If the video frame is of type two, then when the paper page in the video frame is identified as the same page as the paper page in the stored most recently recalled image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the most recently recalled image, the region containing the mapping position is found from the existing index library, and a close-up image is generated based on the region. The close-up image and its display priority are returned to the mobile terminal. The close-up image contains at least knowledge content related to the region where the target identifier is located, generated based on the lecture model. In addition, the audio stream in the video stream is transcribed to obtain audio text, and the audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input content. Based on the multimodal input content, the user's intent on the paper page and the intent response are obtained, and the intent response and the display priority of the intent response are returned to the mobile terminal. The mobile device is used to display the echo image, the close-up image, and the intent response according to display priority.
[0008] This application also provides an electronic device, including: a processor and a computer-readable storage medium for storing computer program instructions, wherein the computer program instructions, when executed by the computer-readable storage medium, cause the processor to perform the steps of the above method.
[0009] This application also provides a machine-readable storage medium storing computer program instructions that, when executed, enable the implementation of the steps described above.
[0010] As can be seen from the above technical solution, in this embodiment, for each video frame of the video stream transmitted by the mobile terminal, the frame type of the video frame is identified; if the frame type of the video frame is the first type (indicating that the video frame is to be displayed on the mobile terminal), then a corresponding echo image is generated for the video frame and returned to the mobile terminal for display, and structured data is generated based on the paper page content area in the echo image and stored in the index library. If the frame type of the video frame is the second type (indicating that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance), then when the paper page in the video frame is identified as the same page as the paper page in the stored most recent echo image, a mapping position with a mapping relationship with the target identifier in the video frame is found from the most recent echo image, the area containing the mapping position is found from the existing index library, and a close-up image is generated based on the area and returned to the mobile terminal for display.
[0011] In this way, by identifying the frame type of the video frame, the system proactively triggers the corresponding processing flow (i.e., generating an echo image or a local feature image) and returns the processing result to the mobile device for display. This proactively senses the user's current state in the paper-screen interaction scenario (e.g., viewing the entire paper page or focusing on a specific question), thereby proactively triggering the display and analysis of relevant knowledge content. This proactive paper-screen interaction method eliminates the need for users to perform additional manual clicks, specific gestures, or voice commands, effectively avoiding the interaction burden caused by passive interaction. It provides effective feedback without any interaction burden on the user, improving the efficiency of assisted learning.
[0012] Furthermore, the audio stream in the video stream is transcribed to obtain audio text. The audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input. Based on this multimodal input, the user's intent on the printed page and the intent response are obtained, and the intent response is returned for display on the mobile device. This allows for deep analysis of user intent and dynamic generation of targeted intent responses based on the content of all printed pages or the area where the target identifier is located, matching the audio text. This solves the problem of fixed and singular feedback in related technologies, enabling real-time and accurate understanding of user intent, providing effective intelligent feedback, and significantly improving the efficiency of assisted learning. Attached Figure Description
[0013] Figure 1 This is a schematic diagram illustrating an application scenario provided in the embodiments of this application; Figure 2 A flowchart illustrating the method provided in the embodiments of this application; Figure 3 A schematic diagram illustrating the process of identifying the frame type of the video frame provided in an embodiment of this application; Figure 4 Another schematic flowchart illustrating the method provided in this application embodiment; Figure 5 This is a schematic diagram of the system structure provided in the embodiments of this application; Figure 6 This is a schematic diagram of the device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0014] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, and to make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions in the embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0015] Before introducing the method provided in the embodiments of this application, the existing technical problems will be described in detail: Existing paper-screen interaction methods are mainly implemented through the following categories: The first type of method involves scanning or taking a single photo of an image using an image acquisition device (such as a scanning pen or mobile phone), processing the captured image for text recognition, and then displaying the results. This type of method is based on static image recognition technology and can only achieve "single capture and static output." However, students will perform dynamic operations such as page turning, and this type of method cannot capture images continuously and dynamically, thus failing to meet the needs of dynamic interaction. Therefore, the second type of method has emerged.
[0016] The second type of method involves synchronizing a video stream of students using paper materials to an electronic screen using real-time projection tools. For example, this can be done via wireless screen mirroring from a mobile phone to a computer or television. However, while this method can dynamically capture the learning process, it only synchronizes the video stream to the electronic screen and cannot understand the student's intentions or enable dynamic interaction with the student. Therefore, research has shifted to the third type of method in recent years.
[0017] The third method: execute preset fixed functions through preset interaction methods (manual clicks, specific gestures, or voice commands).
[0018] For example, students can directly click or select content to trigger functions such as English and Chinese analysis of the selected word or phrase. However, students can only click on pre-defined fixed areas and can only achieve the functions corresponding to those fixed areas (in this case, word lookup). Alternatively, there's a learning device that comes with an Augmented Reality (AR) picture book. Users click on different areas on the picture book to display 3D images of the objects within those areas. This requires the picture book and the learning device to be completely customized and used together, with the correspondence between areas on each page of the picture book and the corresponding 3D images needing to be set in advance.
[0019] For example, users can use specific gestures, such as "pointing a question mark" or "making a fist," to invoke a grading model to solve the questions on the current paper page. However, only specific gestures can be preset and bound to a single function.
[0020] For example, a user issues a preset voice command to trigger the corresponding function. For instance, if the voice command is "Open the knowledge point for question 5," the system will retrieve the knowledge point for question 5 from the knowledge point index. However, this requires the student to utter a preset voice command that the system can recognize to trigger the response. It cannot analyze intent based on real-time, non-preset voice commands from the user. For example, if a student issues a voice command like "Provide the solution steps for this question," the learning machine cannot identify where "this" refers, nor can it identify which area of the printed material the "solution steps" correspond to. Or, if a student issues a voice command like "Explain this formula derivation again," the learning machine cannot know "what the derivation formula is," and therefore does not know how to execute the command.
[0021] As described above, these methods all require users to learn specific interaction methods in advance and interrupt the learning process to perform specific operations. Learning machines and other mobile devices passively execute fixed functions, resulting in single and rigid functions. They are unable to understand user intentions in real time and accurately, and provide effective intelligent feedback.
[0022] Based on this, embodiments of this application provide paper-screen interaction methods, devices, and systems to proactively sense the user's learning status, understand the user's intentions in real time and accurately, and provide effective intelligent feedback.
[0023] The methods provided in the embodiments of this application are described in detail below, first in conjunction with... Figure 1 The application scenarios of the embodiments of this application are illustrated with examples: like Figure 1 As shown, students learn using paper materials. After the mobile device is powered on, it captures video and audio streams and transmits them to a cloud server. Optionally, as an example, the mobile device can be a smartphone, a learning machine, etc. The camera on the mobile device captures the video stream in real time at a frame rate of 30 frames per second and uses an efficient video encoding algorithm, such as H.265 encoding, to compress the captured video stream to reduce the amount of data during transmission. Then, the compressed video frames are sent to the cloud server in real time via a wireless network. During transmission, adaptive bitrate adjustment technology is used to adjust the video encoding bitrate in real time according to network conditions to ensure the stability and smoothness of the video stream transmission. For example, when network bandwidth is sufficient, the bitrate is increased to improve video quality; when network congestion occurs, the bitrate is reduced to ensure smooth video playback. Simultaneously, the mobile device captures the audio stream, keeping the timestamps aligned with the video stream, and transmits it to the cloud. The mobile device performs preliminary noise reduction on the audio signal (e.g., Gaussian filtering) and echo cancellation on the audio to reduce the amount of invalid data transmission. The cloud server receives video and audio streams and executes the steps of the method provided in the embodiments of this application (see the embodiments below for details) to assist students in learning.
[0024] It should be noted that the method provided in this application can also be applied to other scenarios such as grading and answering questions on paper-based assignments or exams, and collaborative paper-screen guidance in vocational education training. For ease of description, this application describes the scenario as a student learning scenario.
[0025] The following is based on Figure 1 The illustrated scenario details the steps of the method provided in the embodiments of this application: See Figure 2 , Figure 2 This is a flowchart illustrating the method provided in an embodiment of this application. As one embodiment, the execution entity of this method can be a cloud server.
[0026] like Figure 2 As shown, the process may include the following steps: S201, for each video frame in the video stream transmitted by the mobile terminal, identify the frame type of the video frame.
[0027] In this embodiment, the frame type is either a first type or a second type. If the frame type of the video frame is the first type (which can also be called the type to be displayed), it indicates that the video frame is to be displayed on the mobile device. If the frame type of the video frame is the second type (which can also be called the target identifier stable type), it indicates that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance.
[0028] In this embodiment, the specific implementation method for identifying the frame type of the video frame will be illustrated with examples later, and will not be repeated here.
[0029] S202, if the frame type of the video frame is the first type, then generate the corresponding echo image based on the video frame and return it to the mobile terminal for display, and generate structured data based on the paper page content area in the echo image and store it in the index library.
[0030] In this embodiment, the specific implementation method for generating the echo image corresponding to the video frame will be illustrated later, and will not be repeated here.
[0031] S203, if the frame type of the video frame is the second type, when it is identified that the paper page in the video frame is the same as the paper page in the stored recent echo image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the recent echo image, the area containing the mapping position is found from the existing index library, a close-up image is generated based on the area, and returned to the mobile terminal for display.
[0032] In this embodiment, the close-up image contains at least knowledge content related to the area where the target identifier is located, generated based on the topic model.
[0033] The specific implementation of step S203 will be illustrated later, and will not be repeated here.
[0034] S204, transcribe the audio stream in the video stream to obtain audio text, use the audio text and video frames of type 1 or type 2 within a specified time period as multimodal input content, obtain the user's intent on the paper page and the intent response based on the multimodal input content, and return the intent response to the mobile terminal for display.
[0035] In this embodiment, the specific implementation of step S204 will be illustrated later, and will not be repeated here.
[0036] This concludes the process. Figure 2 The process is shown below.
[0037] pass Figure 2 As shown in the flowchart, in this embodiment, for each video frame of the video stream transmitted by the mobile terminal, the frame type of the video frame is identified. If the frame type of the video frame is a first type (indicating that the video frame is to be displayed on the mobile terminal), a corresponding echo image is generated and returned to the mobile terminal for display. Structured data is generated based on the paper page content area in the echo image and stored in the index library. If the frame type of the video frame is a second type (indicating that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance), when the paper page in the video frame is identified as the same page as the paper page in the stored most recent echo image, a mapping position with a mapping relationship with the target identifier in the video frame is found from the most recent echo image. The area containing the mapping position is found from the existing index library, and a close-up image is generated based on the area and returned to the mobile terminal for display.
[0038] In this way, by identifying the frame type of the video frame, the system proactively triggers the corresponding processing flow (i.e., generating an echo image or a local feature image) and returns the processing result to the mobile device for display. This proactively senses the user's current state in the paper-screen interaction scenario (e.g., viewing the entire paper page or focusing on a specific question), thereby proactively triggering the display and analysis of relevant knowledge content. This proactive paper-screen interaction method eliminates the need for users to perform additional manual clicks, specific gestures, or voice commands, effectively avoiding the interaction burden caused by passive interaction. It provides effective feedback without any interaction burden on the user, improving the efficiency of assisted learning.
[0039] Furthermore, the audio stream in the video stream is transcribed to obtain audio text. The audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input. Based on this multimodal input, the user's intent on the printed page and the intent response are obtained, and the intent response is returned for display on the mobile device. This allows for deep analysis of user intent and dynamic generation of targeted intent responses based on the content of all printed pages or the area where the target identifier is located, matching the audio text. This solves the problem of fixed and singular feedback in related technologies, enabling real-time and accurate understanding of user intent, providing effective intelligent feedback, and significantly improving the efficiency of assisted learning.
[0040] The following is a detailed explanation of how the frame type of the video frame was identified: See Figure 3 , Figure 3 This is a schematic diagram illustrating the process of identifying the frame type of the video frame provided in an embodiment of this application. For example... Figure 3 As shown, the process may include the following steps: S301, determine whether the video frame includes a complete paper page and whether a target identifier exists.
[0041] In this embodiment, the target identifier can be a fingertip or a pen tip, etc.
[0042] The specific implementation of determining whether a video frame includes a complete paper page can be as follows: predict the position of the paper page using the YOLOv8 model (the position can be represented by the coordinates of 4 vertices in a specified coordinate system). If the predicted position exceeds the boundary of the video frame (the boundary of the video frame is represented by the pixel boundary), then it is determined that the video frame does not contain a complete paper page. If the predicted position does not exceed the boundary of the video frame, then it is determined that the video frame contains a complete paper page.
[0043] The specific implementation of determining whether a target identifier exists in a video frame can be as follows: the fingertip position is detected by a CNN model, and the fingertip position can be represented by the fingertip coordinates in a specified coordinate system.
[0044] S302, if the video frame includes a complete paper page and the video frame does not have a target identifier, then when there is no previous video frame for the current video frame, or when the similarity between the content of the paper page in the current video frame and the previous video frame is greater than a first similarity threshold and there is no video frame currently identified as the first type, or when the similarity between the content of the paper page in the current video frame and the previous video frame is greater than the first similarity threshold and the similarity between the content of the paper page in the current video frame and the most recent video frame identified as the first type is less than a second similarity threshold, the frame type of the video frame is determined to be the first type.
[0045] In this embodiment, the first similarity threshold and the second similarity threshold can be the same or different, and can be set according to the specific application scenario.
[0046] For example, if the video frame contains a complete piece of paper (such as a whole page of a textbook) and does not contain a fingertip, then it includes the following three cases: Case 1: If there is no previous video frame for this video frame, then the frame type of this video frame is directly determined to be the first type.
[0047] Scenario 2: There is a previous video frame for this video frame, but there is no previously identified first-type video frame. In this case, as long as the similarity between the content of the paper page in this video frame and the previous video frame is greater than 0.8, the frame type of this video frame is determined to be the first type.
[0048] Scenario 3: If there is a previous video frame for the current video frame, and there was a previously identified first-type video frame, then the similarity between the content of the paper page in the current video frame and the previous video frame is greater than 0.8, and the similarity between the current video frame and the paper page in the most recent identified first-type video frame is less than 0.3, then the frame type of the current video frame is determined to be first-type.
[0049] S303, if the video frame has a target identifier, then when the distance between the target identifier and the position of the target identifier in the video frame and the previous video frame is less than a first preset distance, and there is currently no video frame that has been determined to be of the second type, or when the distance between the fingertip and the position of the fingertip in the video frame and the previous video frame is less than the first preset distance, and the distance between the fingertip and the position of the nearest video frame that has been determined to be of the second type is greater than a second preset distance, the frame type of the video frame is determined to be of the second type.
[0050] In this embodiment, the first set distance is less than the second set distance. The first set distance is set to determine that the target identifier is stable, and the second set distance is set to distinguish it from the previous second type of video frame.
[0051] For example, if the video frame contains a fingertip, then the following two cases apply: Scenario 1: There is a previous video frame for this video frame, but there is no previously identified second-type video frame. In this case, if the distance between the fingertip and the previous video frame is less than 5 pixels, the frame type of this video frame is determined to be second-type.
[0052] Scenario 2: There is a previous video frame for this video frame, and there was a previously identified second-type video frame. In this case, the frame type of this video frame is determined to be second-type only when the distance between the fingertip and the previous video frame is less than 5 pixels, and the distance between the fingertip and the nearest identified second-type video frame is greater than 40 pixels (0.3).
[0053] The above provides a detailed explanation of how to identify the frame type of the video frame.
[0054] The following section details the specific implementation of generating an echo image and generating structured data based on the paper page content area in the echo image when the frame type of the video frame is determined to be the first type: As an example, the specific implementation of generating the echo image corresponding to the video frame based on the video frame can be as follows: perform image enhancement processing on the paper page in the video frame to obtain the echo image.
[0055] Specifically, based on the obtained coordinates of the four vertices, a perspective transformation algorithm (e.g., using OpenCV's getPerspectiveTransform function) is used to correct the tilted or deformed paper page into a regular rectangular plane. Furthermore, the Otsu thresholding algorithm is used to separate the paper page from the background, removing the background and retaining the main body of the paper page.
[0056] The main body of the preserved paper pages undergoes contrast enhancement and noise optimization. For example, adaptive histogram equalization (CLAHE) is used to enhance the grayscale difference between the text and the paper, thereby improving contrast. Non-local mean filtering (NL-means) is combined to remove image noise and improve the clarity of the handwriting edges, thus obtaining the echo image. The echo is then stored in the index and associated with a timestamp.
[0057] As an example, the specific implementation of generating structured data based on the content areas of the paper page in the echo image can be as follows: Through text detection and layout analysis, identify each content area of the paper page in the echo image. For each content area, generate the corresponding structured data using OCR recognition. Optionally, the structured data can be in JSON format.
[0058] Each content area is a question area. The structured data corresponding to the question area includes: the location of the question area (which can be represented by a coordinate frame) and the question text of the question in the question area.
[0059] Each content area is a picture book illustration area. The structured data corresponding to the illustration area includes: the location of the illustration area and the text description information of the illustration in the illustration area.
[0060] For example, the LayoutLMv3 model is used to perform text detection and layout analysis on the displayed image, identifying each question area (including question number, question stem, options, or answer area), thereby outputting structured data. The output structured data includes at least: question number (e.g., "3"), coordinate frame (x1, y1, x2, y2), and question text. The obtained structured data is then stored in an index.
[0061] The above provides a detailed explanation of the specific implementation method for generating an echo image and generating structured data based on the paper page content area in the echo image.
[0062] The specific implementation method for generating feature maps is described in detail below: As an example, after determining that the frame type of the video frame is the second type, generating the feature map includes the following steps: First, identify whether the paper page in the video frame is the same as the paper page in the most recently stored echo image.
[0063] Specifically, using the SIFT feature point matching algorithm, the similarity of the video frame with the most recent (time-most recent) echo image (e.g., the timestamp difference between the two is ≤1 second) is calculated. If the similarity of the paper page content is greater than the third similarity threshold (e.g., 0.85), it is determined that the paper page in the video frame is the same as the paper page in the most recent echo image; otherwise, it is determined that the paper page in the video frame is not the same as the paper page in the most recent echo image.
[0064] Secondly, find the mapping position that has a mapping relationship with the target identifier in the video frame from the most recent echo map.
[0065] Specifically, based on the feature point matching results, a coordinate transformation matrix between the video frame and the echo image is established to map the target identifier position (such as the fingertip coordinates) to the mapped position in the echo image.
[0066] Finally, the region containing the mapped location is found from the existing index, and a close-up image is generated based on that region.
[0067] Specifically, the index is searched to find structured data containing the mapping location. A local image is cropped from the echo image according to the content region location in the structured data. Based on the knowledge content of the content region where the target identifier is located in the structured data, knowledge content related to the region where the target identifier is located is obtained. A feature map is generated based on the obtained local image and related indication content.
[0068] For example, after finding the mapped coordinates of the fingertip coordinates in the video frame in the stored recent echo image, the coordinate frame containing the mapped coordinates is found in the index. A partial image indicated by the coordinate frame is cropped in the echo image (the cropping is extended by a certain distance, such as 5 pixels) as the target question area. The target question text belonging to the same structured data as the found coordinate frame is input into the question model (such as the Nine Chapters Model) to obtain the analysis steps and answer analysis of the target question. The answer analysis is placed in the answer area of the target question area to obtain a close-up image. This close-up image is sent to the mobile device for highlighting or pop-up display.
[0069] The above provides a detailed explanation of the specific implementation methods for generating feature maps.
[0070] The following section elaborates on analyzing intent and obtaining intent responses: As an example, analyzing intent and obtaining intent responses includes the following steps: First, the audio stream in the video stream is transcribed to obtain the audio text.
[0071] Specifically, the received speech signal is converted into text using an Automatic Speech Recognition (ASR) module. When the duration of no received speech signal exceeds a set threshold, the current audio stream is considered to have ended. At the end of the current audio stream, the resulting text sequence is identified as the audio text corresponding to that stream, and the timestamp used to generate the audio text is associated with it.
[0072] Secondly, audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input content.
[0073] Specifically, if a second type of video frame exists within the specified time period, the last second type video frame and audio text within that time period are used as the multimodal input content. If no second type of video frame exists within the specified time period, but a first type of video frame exists, the last first type video frame and audio text within that time period are used as the multimodal input content.
[0074] For example, after obtaining the second type of video frame or the first type of video frame within 3 seconds before the timestamp of the audio text, the definitions of multiple setting functions, the second type of video frame or the first type of video frame, and the coordinates of the four vertices of the paper page or the fingertip coordinates are concatenated according to the set format to obtain the matching multimodal input content.
[0075] Finally, the user's intent on the paper page and the intent response are obtained based on the multimodal input content.
[0076] Specifically, the multimodal input content is fed into a pre-configured Visual-Language Model (VLM) to obtain the target function identifier and parameter values of each parameter in the target function, which are output by the VLM model. The parameter values of each parameter in the target function are the parameter values required to execute the target function and realize the user's intent. Based on the target function identifier, the corresponding target tool is invoked to obtain the parameters and generate the intended response.
[0077] For example, the functions include at least one of the following: read_question, search_knowledge, search_chinese_word, and search_english_word. The search_chinese_word function requires parameters including: word content, coordinates of the question area, and question identifier (Identity Document, ID).
[0078] The correspondence between target function identifiers and their corresponding tools includes at least the following: the tools corresponding to the `search_chinese_word` and `search_english_word` functions are dictionary databases. The tool corresponding to the `search_knowledge` function is a search question bank.
[0079] In related technologies, when a user points to the same printed text and says "look up the uncommon words in this text" or "analyze the rhetorical devices in this text," it is difficult to distinguish between the two different intentions and call the corresponding "word lookup tools" and "literary analysis tools." However, the above method can accurately distinguish the above intentions and dynamically adjust the intention response according to the user's real-time intentions.
[0080] By employing the aforementioned multimodal input methods, collaborative recognition of "paper location (video stream)" and user needs (voice stream) is achieved, upgrading the interaction from "single-modal response" to "multimodal linkage." This overcomes the limitations of related technologies that can only execute preset voice commands, improving the accuracy of intent recognition and thus enhancing the user experience. Ultimately, through multimodal fusion and intelligent tool invocation, a closed-loop interaction is achieved, enabling dynamic tracking and precise positioning of paper content and real-time response to user intent. This breaks through the static and single-modal limitations of traditional paper-screen interaction, achieving the goal of real-time and accurate paper-screen interaction.
[0081] It should be noted that the intended response can be packaged into JSON format and returned to the mobile device.
[0082] The above provides a detailed explanation of analyzing intent and obtaining intent responses.
[0083] As one embodiment, the method further includes: returning the display priority of the echo image to the mobile device; returning the display priority of the close-up image to the mobile device; returning the display priority of the intended answer to the mobile device; so that the mobile device displays the echo image, close-up image, and intended answer according to their respective priorities. Optionally, the intended answer has a higher priority than the close-up image, and the close-up image has a higher priority than the echo image.
[0084] After receiving the content returned by the cloud server, the mobile device displays it according to the display priority of the received content.
[0085] Furthermore, mobile devices may also include other tools, such as interactive cards and favorites buttons, with other tools having the lowest priority.
[0086] To illustrate the method provided in this application in more detail, the following will be combined with... Figure 4 The solution provided in this application will be described in more detail by way of specific embodiments.
[0087] In this embodiment, the mobile device is a learning machine, and the paper page includes multiple question areas. The first type is denoted as the type to be displayed, and the second type is denoted as the fingertip stable type.
[0088] The process includes the following steps: First, for each video frame in the video stream transmitted by the learning machine, identify the frame type of that video frame.
[0089] For detailed steps, please refer to [link / details]. Figure 3 The specific implementation methods described in the illustrated embodiments will not be repeated here.
[0090] Secondly, if the frame type of the video frame is a type to be displayed, the display link is triggered to process the image to generate a display image. The display image and its display priority are sent to the learning machine, and the structured data of the display image is stored in the index path.
[0091] Specifically, the specific execution steps for processing the echo link image are detailed in the implementation method described in the above embodiment for generating the echo image, and will not be repeated here.
[0092] Furthermore, if the frame type of the video frame is the fingertip stable type, the question-and-answer correction link is triggered to process the image to generate a feature map of the area pointed to by the fingertip. The feature map and its display priority are sent to the learning machine so that the learning machine can highlight or display a close-up image in a pop-up window. The close-up image includes the target question pointed to by the fingertip and the corresponding solution steps and answer analysis.
[0093] Specifically, the detailed execution steps of the image processing in the question-and-answer correction link are described in the specific implementation method in the above embodiment for generating close-up images, and will not be repeated here.
[0094] Finally, the audio stream in the video stream is transcribed to obtain audio text. The definitions of multiple set functions, video frames of the fingertip-stabilized type or the type to be displayed, and the coordinates of the four vertices of the paper page or the fingertip coordinates are concatenated according to the set format to obtain the matching multimodal input content. The multimodal input content is input into the VLM model to obtain the target function identifier and the parameter values of each parameter in the target function output by the VLM model. Based on the target function identifier, the corresponding target tool is called to obtain each parameter and generate the intent response.
[0095] Specifically, the detailed execution steps for parsing user intent and generating intent answers are described in the above embodiments of parsing user intent and generating intent answers, and will not be repeated here.
[0096] The methods provided in the embodiments of this application have been described above. The systems and apparatus provided in the embodiments of this application are described below: See Figure 5 , Figure 5 This is a system structure diagram provided for an embodiment of this application. (See diagram below.) Figure 5 As shown, the system includes: a cloud server 501 and a mobile terminal 502; The cloud server 501 is used to identify the frame type of each video frame in the video stream transmitted by the mobile terminal; the frame type is either a first type or a second type; if the frame type of the video frame is a first type, it indicates that the video frame should be displayed on the mobile terminal; if the frame type of the video frame is a second type, it indicates that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance. If the frame type of the video frame is the first type, then generate the corresponding echo image based on the video frame, return the echo image and the display priority of the echo image to the mobile terminal, and generate structured data based on the paper page content area in the echo image and store it in the index library. If the video frame is of type two, when the paper page in the video frame is identified as the same as the paper page in the stored most recently recalled image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the most recently recalled image, the region containing the mapping position is found from the existing index library, a close-up image is generated based on the region, and the close-up image and its display priority are returned to the mobile device; the close-up image contains at least the knowledge content related to the region where the target identifier is located, generated based on the lecture model; Additionally, the audio stream in the video stream is transcribed to obtain audio text, and the audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input content. Based on the multimodal input content, the user's intent on the paper page and the intent response are obtained, and the intent response and the display priority of the intent response are returned to the mobile device. Mobile device 502 is used to display echo images, close-up images, and intent responses according to display priority.
[0097] It should be noted that the steps of any embodiment of the method provided in this application are executed by the cloud server, and will not be described again here.
[0098] See Figure 6 , Figure 6 This is a structural diagram of the device provided in an embodiment of this application. Figure 6 As shown, the device is used on a cloud server, and the device 600 includes: a recognition module 601 and an image processing module 602.
[0099] The identification module 601 is used to identify the frame type of each video frame in the video stream transmitted by the mobile terminal; the frame type is either a first type or a second type; if the frame type of the video frame is the first type, it indicates that the video frame is to be displayed on the mobile terminal; if the frame type of the video frame is the second type, it indicates that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance. The image processing module 602 is used to generate a corresponding echo image based on the video frame and return it to the mobile terminal for display if the frame type of the video frame is the first type, and to generate structured data based on the paper page content area in the echo image and store it in the index library. If the video frame is of type two, when the paper page in the video frame is identified as the same page as the paper page in the stored most recently recalled image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the most recently recalled image, the region containing the mapping position is found from the existing index, a close-up image is generated based on the region, and returned to the mobile device for display; the close-up image contains at least the knowledge content related to the region where the target identifier is located, generated based on the lecture model; The intent parsing module 603 is used to transcribe the audio stream in the video stream to obtain audio text, and use the audio text and video frames of type 1 or type 2 within a specified time period as multimodal input content. Based on the multimodal input content, the module obtains the user's intent on the paper page and the intent response, and returns the intent response to the mobile device for display.
[0100] As an example, identifying the frame type of the video frame includes: If the video frame contains a complete paper page and the video frame does not contain a target identifier, then: If there is no previous video frame for the current video frame, or if the similarity between the content of the paper page in the current video frame and the previous video frame is greater than a first similarity threshold and there is no video frame currently identified as the first type, or if the similarity between the content of the paper page in the current video frame and the previous video frame is greater than a first similarity threshold and the similarity between the content of the paper page in the current video frame and the most recent video frame identified as the first type is less than a second similarity threshold, then the frame type of the current video frame is determined to be the first type. If the video frame contains a target identifier, then: When the distance between the target identifier in the current video frame and the previous video frame is less than a first preset distance, and there is currently no video frame that has been identified as the second type, or when the distance between the target identifier in the current video frame and the previous video frame is less than the first preset distance, and the distance between the target identifier in the current video frame and the nearest video frame that has been identified as the second type is greater than a second preset distance, the frame type of the video frame is determined to be the second type.
[0101] As one example, generating the echo image corresponding to the video frame based on the video frame includes: Image enhancement processing was performed on the paper page in the video frame to obtain the echo image.
[0102] As one example, generating structured data based on the paper page content area in the echo image includes: Identify the content areas of the paper page in the displayed image; For each content area, generate the corresponding structured data.
[0103] As an example, any content area is a question area; the structured data corresponding to the content area includes: the area location corresponding to the question area, and the question text of the question in the question area; Finding the region containing the mapped location from an existing index and generating a close-up image based on that region includes: Find the target region containing the mapped location from the structured data in the index; Based on the local image indicating the location of the target area in the echoed image, the target question area is obtained; Input the target question text, which belongs to the same structured data as the target region, into the question model to obtain the corresponding parsed text; A close-up image is generated based on the target question area and the parsed text.
[0104] As one example, transcribing the audio stream in a video stream to obtain audio text includes: The received speech signal is converted into text using an Automatic Speech Recognition (ASR) module. When the duration of no voice signal received exceeds a set duration threshold, the current audio stream is determined to end. At the end of the current audio stream, the resulting text sequence is identified as the audio text corresponding to the current audio stream.
[0105] As one example, using audio text and video frames of type 1 or type 2 within a specified time period as multimodal input content includes: If a second type of video frame exists within the specified time period, then the last second type of video frame and audio text within the specified time period are used as multimodal input content. If there are no second-type video frames in the specified time period, but there are first-type video frames, then the last first-type video frame and audio text in the specified time period will be used as multimodal input content.
[0106] As one embodiment, the recognition module is also used to: return the display priority of the echo image to the mobile device. The image processing module is also used to: return the display priority of close-up images to the mobile device; The intent parsing module is also used to: return the display priority of the intent response to the mobile device; so that the mobile device displays the echo image, close-up image, and intent response according to the time priority.
[0107] This concludes the process. Figure 6 Structural description of the device shown.
[0108] See Figure 7 , Figure 7 This is a structural diagram of an electronic device provided in an embodiment of this application. Figure 7 As shown, the hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.
[0109] Based on the same application concept as the above method, this application embodiment also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the method disclosed in the above examples of this application.
[0110] For example, the aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For instance, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0111] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A paper screen interaction method, characterized by, This method is applied to a cloud server; the method includes: For each video frame in the video stream transmitted by the mobile terminal, the frame type of the video frame is identified; the frame type is either a first type or a second type; if the frame type of the video frame is a first type, it is indicated that the video frame is to be displayed on the mobile terminal; if the frame type of the video frame is a second type, it is indicated that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance. If the frame type of the video frame is the first type, then the corresponding echo image of the video frame is generated based on the video frame and returned to the mobile terminal for display, and structured data is generated based on the paper page content area in the echo image and stored in the index library. If the video frame is of type two, when the paper page in the video frame is identified as the same page as the paper page in the stored most recently recalled image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the most recently recalled image, the region containing the mapping position is found from the existing index library, a close-up image is generated based on the region, and returned to the mobile terminal for display; the close-up image contains at least knowledge content related to the region where the target identifier is located, generated based on the lecture model; Furthermore, the audio stream in the video stream is transcribed to obtain audio text, and the audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input content. Based on the multimodal input content, the user's intent on the paper page and the intent response are obtained, and the intent response is returned to the mobile terminal for display.
2. The method according to claim 1, characterized in that, The frame type for identifying the video frame includes: If the video frame contains a complete paper page and the video frame does not contain a target identifier, then: If there is no previous video frame for the current video frame, or if the similarity between the content of the paper page in the current video frame and the previous video frame is greater than a first similarity threshold and there is no video frame currently identified as the first type, or if the similarity between the content of the paper page in the current video frame and the previous video frame is greater than a first similarity threshold and the similarity between the content of the paper page in the current video frame and the most recent video frame identified as the first type is less than a second similarity threshold, then the frame type of the current video frame is determined to be the first type. If the video frame contains a target identifier, then: When the distance between the target identifier in the video frame and the previous video frame is less than a first preset distance, and there is currently no video frame that has been determined to be of the second type, or when the distance between the target identifier in the video frame and the previous video frame is less than the first preset distance, and the distance between the target identifier in the video frame and the nearest video frame that has been determined to be of the second type is greater than a second preset distance, the frame type of the video frame is determined to be of the second type.
3. The method according to claim 1, characterized in that, The step of generating the echo image corresponding to the video frame based on the video frame includes: The paper page in the video frame is subjected to image enhancement processing to obtain the echo image.
4. The method according to claim 1, characterized in that, The generation of structured data based on the paper page content area in the echo image includes: Identify the content areas of the paper page in the displayed image; For each content area, generate the corresponding structured data.
5. The method according to claim 4, characterized in that, Each content area is a question area; the structured data corresponding to this content area includes: the location of the question area and the question text of the question in the question area; The step of finding the region containing the mapped location from an existing index and generating a close-up image based on that region includes: The target region containing the mapped location is located in the structured data of the index library; Based on the local image indicating the location of the target area in the echo image, the target question area is obtained; The target question text, which belongs to the same structured data as the target region, is input into the question model to obtain the corresponding parsed text; The close-up image is generated based on the target question area and the parsed text.
6. The method according to claim 1, characterized in that, The step of transcribing the audio stream in the video stream to obtain audio text includes: The received speech signal is converted into text using an Automatic Speech Recognition (ASR) module. When the duration of no voice signal received exceeds a set duration threshold, the current audio stream is determined to end. At the end of the current audio stream, the resulting text sequence is identified as the audio text corresponding to the current audio stream.
7. The method according to claim 1, characterized in that, The audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input content, including: If a second type of video frame exists within the specified time period, then the last second type of video frame within the specified time period and the audio text are used as the multimodal input content. If there are no second-type video frames in the specified time period, but there are first-type video frames, then the last first-type video frame in the specified time period and the audio text are used as the multimodal input content.
8. The method according to claim 1, characterized in that, The method further includes: Return the display priority of the echo image to the mobile device; Return the display priority of the close-up image to the mobile device; The display priority of the intended answer is returned to the mobile device so that the mobile device displays the echo image, the close-up image, and the intended answer according to the time priority.
9. A paper-screen interactive device, characterized in that, This device is used in a cloud server; the device includes: The identification module is used to identify the frame type of each video frame in the video stream transmitted by the mobile terminal; the frame type is a first type or a second type; if the frame type of the video frame is the first type, it indicates that the video frame is to be displayed on the mobile terminal; if the frame type of the video frame is the second type, it indicates that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance. The image processing module is used to generate a corresponding echo image based on the video frame and return it to the mobile terminal for display if the frame type of the video frame is the first type; and to generate structured data based on the paper page content area in the echo image and store it in the index library. If the video frame is of type two, when the paper page in the video frame is identified as the same page as the paper page in the stored most recently recalled image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the most recently recalled image, the region containing the mapping position is found from the existing index library, a close-up image is generated based on the region, and returned to the mobile terminal for display; the close-up image contains at least knowledge content related to the region where the target identifier is located, generated based on the lecture model; The intent parsing module is used to transcribe the audio stream in the video stream to obtain audio text, and use the audio text and video frames of type 1 or type 2 within a specified time period as multimodal input content. Based on the multimodal input content, the module obtains the user's intent on the paper page and the intent response, and returns the intent response to the mobile terminal for display.
10. A paper-screen interactive system, characterized in that, The system includes: a cloud server and a mobile terminal; A cloud server is used to identify the frame type of each video frame in the video stream transmitted from the mobile terminal; the frame type is a first type or a second type; if the frame type of the video frame is the first type, it indicates that the video frame is to be displayed on the mobile terminal; if the frame type of the video frame is the second type, it indicates that the position of the same target identifier in the video frame and the previous video frame satisfies a stable distance. If the frame type of the video frame is the first type, then a corresponding echo image is generated based on the video frame, the echo image and its display priority are returned to the mobile terminal, and structured data is generated based on the paper page content area in the echo image and stored in the index library. If the video frame is of type two, then when the paper page in the video frame is identified as the same page as the paper page in the stored most recently recalled image, the mapping position that has a mapping relationship with the target identifier in the video frame is found from the most recently recalled image, the region containing the mapping position is found from the existing index library, and a close-up image is generated based on the region. The close-up image and its display priority are returned to the mobile terminal. The close-up image contains at least knowledge content related to the region where the target identifier is located, generated based on the lecture model. In addition, the audio stream in the video stream is transcribed to obtain audio text, and the audio text and video frames of type 1 or type 2 within a specified time period are used as multimodal input content. Based on the multimodal input content, the user's intent on the paper page and the intent response are obtained, and the intent response and the display priority of the intent response are returned to the mobile terminal. The mobile device is used to display the echo image, the close-up image, and the intent response according to display priority.