Data processing method and device based on visual semantic analysis
Patent Information
- Application Number
- CN202611011893.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,传统桌面语音助手依赖语音文本进行意图识别,当用户指令中包含这个这里这道题等指代词时,无法确定其所指向的桌面对象;而现有的截图问答方案又将截图作为孤立图片处理,缺少对课程上下文和桌面应用状态的联动理解,容易出现解释对象错误、回答缺乏课件支撑或无法执行后续桌面控制动作等问题,严重影响授课场景下的交互效率和准确性
[0008]根据本说明书实施例的第四方面,提供了一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现上述数据处理方法的步骤。
Smart Images

Figure CN122819476A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of intelligent education technology, and in particular to a data processing method based on visual semantic analysis. One or more embodiments of this specification also relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] In teaching or demonstration scenarios, teachers often need to select and display specific parts of the courseware, formulas, charts, questions, or student report pages while lecturing, and issue instructions such as "explain this problem, analyze the cause of the error, and explain it based on this diagram" via voice.
[0003] However, traditional desktop voice assistants rely on voice and text for intent recognition. When user commands contain referential words such as "this," "here," or "this question," they cannot determine the desktop object they are referring to. Furthermore, existing screenshot question-and-answer solutions treat screenshots as isolated images, lacking a linked understanding of the course context and desktop application state. This can easily lead to problems such as incorrect interpretation of the object, answers lacking courseware support, or inability to execute subsequent desktop control actions, seriously affecting the efficiency and accuracy of interaction in teaching scenarios. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a data processing method based on visual semantic analysis. One or more embodiments of this specification also relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a data processing method based on visual semantic analysis is provided, comprising: Acquire image data corresponding to the target visual region, wherein the target visual region is determined by a screenshot window, and the screenshot window is created in response to a screenshot trigger operation of the target client; Upon obtaining voice input data, voice recognition is performed on the voice input data to obtain voice text, wherein the voice input data is used to ask follow-up questions about the image data and / or trigger desktop interaction commands; The application status information of the target client is obtained, and the image data, the voice text, and the application status information are fused to obtain multimodal fusion information; By using a collaborative mechanism of instruction matching and agent reasoning, the target interaction intent of the multimodal fusion information is determined, and the corresponding target action is executed according to the target interaction intent to obtain the action execution result.
[0006] According to a second aspect of the embodiments of this specification, a data processing apparatus is provided, comprising: The acquisition module is configured to acquire image data corresponding to a target visual region, wherein the target visual region is determined by a screenshot window, and the screenshot window is created in response to a screenshot trigger operation from the target client; The recognition module is configured to perform speech recognition on the speech input data when the speech input data is acquired, and obtain speech text, wherein the speech input data is used to ask follow-up questions to the image data and / or trigger desktop interaction commands; The fusion module is configured to acquire the application status information of the target client and fuse the image data, the voice text, and the application status information to obtain multimodal fusion information; The execution module is configured to determine the target interaction intent of the multimodal fusion information through a collaborative mechanism of instruction matching and agent reasoning, and to execute the corresponding target action according to the target interaction intent to obtain the action execution result.
[0007] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above-described data processing method.
[0008] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the data processing method described above.
[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.
[0010] This specification provides a data processing method based on visual semantic analysis in one embodiment. By responding to a screenshot trigger operation from the target client, a screenshot window is created to determine the target visual region and acquire the corresponding image data. This allows for precise capture of any visual region within the corresponding screenshot window of the target client, providing a visual foundation for subsequent multimodal interaction. When voice input data is acquired, it is recognized to obtain voice-text, enabling the user to directly ask follow-up questions based on the captured image data or issue desktop interaction commands. By acquiring the application state information of the target client and fusing the image data, voice-text, and application state information, multimodal fusion information is obtained. This information is then combined with a collaborative mechanism of command matching and agent reasoning to complete intent triage and action execution. This allows the system to place image data within a realistic current context, significantly improving the accuracy of semantic understanding, action execution, and result feedback. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating a data processing method based on visual semantic analysis, provided in one embodiment of this specification. Figure 2 This is a schematic diagram of the screenshot visual semantic context acquisition process in a data processing method provided in one embodiment of this specification; Figure 3 This is a schematic diagram of the intent identification and execution triage process in a data processing method provided in one embodiment of this specification; Figure 4 This is a schematic diagram of the data fusion process in a data processing method provided in one embodiment of this specification; Figure 5 This is a schematic flowchart of a dual-path processing method provided in one embodiment of this specification; Figure 6 This is a flowchart illustrating the interactive closed loop in a data processing method provided in one embodiment of this specification; Figure 7 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this specification; Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0012] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0013] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this specification. The singular forms "a," "say," and "that" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terminology used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0014] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, words used herein may be interpreted as when, when, or in response to a determination.
[0015] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0016] This specification provides a data processing method, and also relates to a data processing apparatus, a computing device, and a computer-readable storage medium, which will be described in detail in the following embodiments.
[0017] See Figure 1 , Figure 1 A flowchart of a data processing method based on visual semantic analysis provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0018] Step 102: Obtain image data corresponding to the target visual region, wherein the target visual region is determined by a screenshot window, and the screenshot window is created in response to a screenshot trigger operation from the target client.
[0019] The target visual area refers to a partial area of the screen selected by the user in the desktop environment through a screenshot, which the user wishes to use as an interactive object. This could be a question, a formula, or a chart in a presentation slide, or a piece of text or a graphic element on a student's presentation page. The image data refers to the image file and its related metadata generated by rendering the target visual area. In this embodiment, the image data includes information such as image address, width, and height. The screenshot window refers to a transparent, top-mounted window created in response to a screenshot trigger operation from the target client. This window covers all other windows on the desktop and is used to carry the screen background and receive interactive operations such as selection, dragging, scaling, full-screen selection, or drawing annotations by the user. The screenshot trigger operation refers to a screenshot Q&A initiation action initiated by the user.
[0020] Specifically, when the target client receives a screenshot command triggered by the user, it creates a screenshot window in response to the screenshot trigger operation. This screenshot window is transparent and on top, allowing the user to freely select any visual area in a complex desktop environment with multiple windows and multiple screens.
[0021] After users select or annotate the area on the screenshot window using a mouse or touch, the image data corresponding to the selected area (i.e. the target visual area) is obtained, which transforms the user's blurry screen pointer into structured visual content.
[0022] For example, in a classroom teaching scenario, a teacher can use their mouse to select a geometry proof problem on the current page while facing the presentation window. The system will then take a screenshot of the selected area and upload it, obtaining the corresponding image data. Similarly, in a remote collaborative presentation scenario, the presenter can use the annotation function in the screenshot window to circle key data points on a chart in the shared screen. The annotated selection will then be rendered as an image, obtaining the corresponding image data for subsequent targeted explanations.
[0023] By creating a screenshot window, it is possible to accurately capture any visual area specified by the user in a complex environment with multiple windows and multiple screens on the desktop, and convert the visual area into structured image data. This provides clear visual data for pronouns in subsequent voice input data, fundamentally solving the problem of lack of screen content reference in traditional voice interaction.
[0024] In one or more embodiments of this specification, acquiring the image data corresponding to the target visual region includes: The screenshot window receives user interaction with the desktop area and determines the target visual area based on the interaction results. The target visual region is rendered into an image file to obtain the image data, wherein the image data includes the image address, image width information, and image height information of the image file.
[0025] Interactive operations refer to actions such as selecting areas on the desktop, including dragging with a mouse, touch gestures, or using a keyboard, to define the target visual area that the user wants to interact with. Interactive results refer to the position, size, and shape information of the selected area determined by the above interactive operations, including the starting coordinates, ending coordinates, width, and height of the selected area in the screen coordinate system.
[0026] Rendering to an image file refers to the process of extracting the user-selected visual area from the screen pixels and drawing it onto the canvas, and then encoding the canvas data into image formats such as PNG and JPEG; the image address refers to the Uniform Resource Locator (URL) returned after the image file is uploaded to the object storage service, which is used to enable subsequent target agents to remotely access the image; the image size information refers to the actual pixel size of the image file, including the actual pixel width and height, which is used to preserve the spatial proportion and layout features of the visual content during multimodal fusion.
[0027] Specifically, when creating a screenshot window, users can drag and drop the mouse to draw a rectangular selection area on any desktop area, or switch to full-screen mode or preset ratio mode in the selection tool, or use the brush tool to circle and annotate specific content.
[0028] After the user completes the selection and confirms, the position information of the selected area in the screen coordinate system is mapped to the current display parameters to determine the range of the target visual area. Then, the screen pixels corresponding to the target visual area are drawn onto an off-screen canvas. The pixel data on the canvas is then encoded according to a standard image format and output as an image file. This image file is uploaded to an object storage service, and the returned image URL (i.e., image address) is received. Simultaneously, the actual pixel width and height of the image file are recorded as a structured component of the image data.
[0029] For example, during a lecture, a teacher can use the mouse to drag and select a rectangular area containing complex formulas on the presentation slideshow page. The screenshot window captures the coordinates of the starting and ending points of the selection, determines the selection area, renders the area as a PNG image, and uploads it. The system records the image address and its dimensions of 1920×1080.
[0030] The data processing method provided in the embodiments of this specification receives diverse desktop interaction operations from users through a screenshot window. Specifically, it can flexibly adapt to various visual region specification methods such as box selection, full screen, proportional selection, and pen annotation, and accurately convert the determined target visual region into structured image data containing image address and size information. This provides a unified, complete, and spatially characteristic visual input foundation for subsequent multimodal fusion, ensuring the flexibility and accuracy of visual region acquisition in different application scenarios.
[0031] In one or more embodiments of this specification, when image data is acquired, the image data can be explained or queried, and the corresponding processing path can be determined based on the different instructions received. Specific implementation methods are described below: After acquiring the image data corresponding to the target visual region, the process further includes: In response to the explanation instruction, the image data and explanation mode identification data are input into the target agent to obtain the explanation content of the target agent for the image data; In response to an inquiry command, the image data is saved and the voice input data and / or text input data are acquired.
[0032] Among them, the explanation command refers to the command triggered by the user through the target client after taking a screenshot, which requires explanation of the image data; the explanation mode identifier data is the prompt information used to guide the target agent to perform the explanation task, specifically used to identify a specific explanation mode, that is, in the explanation mode, the target agent will explain the knowledge based on the image data; the question command refers to the command triggered by the user after taking a screenshot, indicating that they want to interact further with the image data. The difference between the question command and the explanation command is that when the question command is received, the image data will not be processed immediately, but will enter a state of waiting for the user to input a question.
[0033] Specifically, upon obtaining image data, the system will determine the user-triggered scene command specific to that image data, and then execute different processing paths based on the different types of scene commands. In practice, if the received scene command is a narration instruction, without waiting for any additional input, the system directly encapsulates the image address from the image data into a Markdown image format, and sends the extended parameters carrying the knowledge narration identifier as narration mode identifier data to the target agent, enabling the target agent to generate narration content for that image data.
[0034] When the received scene command is an inquiry command, the image data is first saved, and the multimodal input channel is opened to enter a waiting state. At this time, the user can ask follow-up questions about the image data through voice or text, thereby obtaining specific questions about the image data.
[0035] For example, in a classroom teaching scenario, after the teacher selects a physics formula on the presentation slide, they click the "explain" button to trigger the explanation instruction. This encapsulates the formula image in Markdown format and sends it along with a knowledge explanation identifier to the course agent (i.e., the target agent), so that the course agent can generate an explanation of the derivation process of the physics formula.
[0036] For example, in a teaching and tutoring scenario, after the teacher selects a student's wrong question, clicks the "Ask" button to trigger the inquiry command. At this time, the question image is saved and voice input is enabled. The teacher waits for the teacher to ask specific follow-up questions such as "Why do students easily make mistakes on this question?" and then combines the question image with the voice input data to provide an answer.
[0037] The data processing method provided in the embodiments of this specification implements a dual-path screenshot question-and-answer strategy by responding to explanation commands and inquiry commands. On the one hand, it supports direct and rapid response to explanation commands, allowing users to obtain immediate visual content explanations without secondary input, thus improving operational efficiency in classroom demonstration scenarios. On the other hand, it supports voice input waiting under inquiry commands, enabling users to ask more specific follow-up questions about the image content. The combination of the two satisfies the need for immediate explanation while also taking into account the flexibility of in-depth interaction.
[0038] Step 104: If voice input data is obtained, perform voice recognition on the voice input data to obtain voice text, wherein the voice input data is used to ask follow-up questions to the image data and / or trigger desktop interaction commands.
[0039] In this context, voice input data refers to the voice signal collected by the microphone when the user triggers an inquiry command, which is used to ask questions about the selected target visual area, or the voice signal used to trigger desktop interaction commands in the context of existing image data. Voice recognition refers to the process of converting the collected voice signal into text content that can be processed by a computer; this process can be completed by the voice recognition module. Voice-to-text refers to the text content output by the voice recognition module after transcribing the voice input data; this voice-to-text is used to provide semantic supplementation or a basis for desktop action control in addition to image data. It should be noted that in this embodiment, the voice input data is associated with the previously acquired target visual area; when the user speaks the voice, it is assumed to be directed to the target visual area, or it triggers desktop interaction control within the context of the image data.
[0040] Specifically, when a user selects a target visual region through a screenshot window and triggers an inquiry command, the image data corresponding to the target visual region is saved as visual context, and the voice acquisition channel is activated, entering a state of waiting for voice input. Upon acquisition of voice input data, the voice recognition module processes the voice input data, converting it into text content to obtain voice-text.
[0041] By associating voice text with saved image data, the voice text can raise specific questions or provide supplementary explanations to the image data, thus enabling the originally static image data to gain dynamic semantic expansion.
[0042] For example, in a classroom error analysis scenario, a teacher selects a math problem in a presentation on the client's desktop, clicks the "Ask" button, and then says, "In which step are students likely to make mistakes on this problem?" The system recognizes this speech as text and uses it as follow-up questions related to the problem image. Similarly, in a software operation guidance scenario, a user selects a complex configuration panel on the interface and says, "How do I set these parameters?" The system recognizes the speech as text and associates it with a screenshot of the configuration panel to generate subsequent operation instructions for that panel.
[0043] In practical applications, when a query command is triggered, follow-up questions can also be asked about the image data by inputting text; this is not a limitation. In this embodiment, voice input enables efficient follow-up questions about the image data, making it more suitable for teaching or demonstration scenarios.
[0044] By using voice input data to follow up on image data, users can ask questions efficiently, significantly improving operational efficiency and naturalness of interaction in scenarios such as teaching and demonstration. At the same time, the separate acquisition of voice text and image data allows the system to flexibly handle questions from multiple different angles targeting the same visual area. Furthermore, by associating voice text and image data, the system can integrate multimodal data to achieve more accurate action execution.
[0045] Step 106: Obtain the application status information of the target client, and fuse the image data, the voice text, and the application status information to obtain multimodal fusion information.
[0046] The target client refers to the application that carries out teaching or demonstration interaction; the application status information refers to various context data of the target client's current running environment, including at least the window status of the current desktop (such as window stacking relationship, whether the window is in a collapsed state, etc.). If a presentation is displayed in the desktop area of the target client, the application status information may also include the presentation's playback status (such as the current playback page number, display title, etc.), as well as the intelligent agent or course configuration status bound to the current session.
[0047] Fusion refers to the process of combining and encapsulating information from different modalities and data sources according to a unified data structure. Through fusion, different types of input data can be processed by the target intelligent agent in the same request.
[0048] Specifically, when acquiring image data and voice text, the multimodal context building module is launched. This module reads the currently maintained application state information from the desktop context module, such as the presentation state, the current window state, and the agent configuration associated with the current session.
[0049] The multimodal context construction module combines these application state information with image data and voice text according to a preset multimodal request format. For example, it converts image data into Markdown image format, fills in voice text as user question fields, and constructs complete multimodal fusion information of visual, text and desktop context in combination with the application state information associated with the current session, for use in subsequent intent recognition and action execution stages.
[0050] For example, in a scenario where a teacher is giving a lesson via a presentation, the teacher selects a formula on the current slide and says, "Explain the meaning of this formula." After obtaining the formula screenshot and the spoken text, the system combines the current presentation status, the course agent bound to the current session, and the existing historical message context to construct multimodal fusion information.
[0051] By multimodal fusion of image data, voice text, and application state information of the target client, visual semantic analysis is no longer limited to the surface features of screenshot pixels. Instead, it can be associated with real teaching contexts, presentation content, course knowledge systems, and historical interaction records. This ensures that the answers generated by the intelligent agent are supported by courseware background and related to course knowledge, effectively solving the problem of answers being detached from actual application scenarios in ordinary screenshot question-and-answer sessions. It significantly improves the accuracy and context relevance of semantic understanding and answers.
[0052] In one or more embodiments of this specification, when course content is displayed on the desktop area, the application status information includes course context information of the course content. The step of fusing the image data, the voice text, and the application status information to obtain multimodal fusion information includes: The image address and image size information corresponding to the image data are fused with the voice text and the course context information to obtain the multimodal fusion information.
[0053] The course content refers to the courseware, textbooks, exercises, or knowledge materials currently being presented in the desktop area of the target client that are related to teaching or training. This course content can be displayed in the form of presentation slides, documents, web-based courseware, or teaching software interfaces, and there are no restrictions on its presentation. The course context information refers to the background knowledge data related to the current course content, which may include the current courseware title, the current slide page number, chapter identifiers, course configuration (such as the bound course knowledge base), the dialogue history accumulated in the current session, and the course's knowledge structure. This information is used to provide teaching context support for visual semantic analysis.
[0054] Specifically, when course content is displayed on the desktop area, the course context information is read, and the image address and size information corresponding to the image data are embedded into the multimodal request structure in Markdown image format. Simultaneously, the voice text is filled in as user input fields, and the course context information is appended to the multimodal request as extended parameters, forming unified multimodal fusion information. This fusion process establishes a connection between image data, voice text, and course context, enabling the intelligent agent to understand and respond specifically when processing visual content, in conjunction with the currently taught course knowledge system.
[0055] For example, in a higher mathematics teaching scenario, when a teacher selects a calculus formula on the current slide and explains how the formula is derived, the system integrates the formula screenshot address, size information, the courseware title "Higher Mathematics Chapter 2", the current page number "Page 15", the knowledge base of the course agent "Higher Mathematics Course Assistant", and previous dialogue history about the concept of limits into a multimodal request.
[0056] The data processing method provided in the embodiments of this specification incorporates course context information into multimodal fusion when course content is displayed on the desktop area. This prevents image data and voice issues from being processed in isolation, but rather deeply connects them with the knowledge system of the currently being taught course. This ensures that the agent's answers are based on specific courseware content and course background, effectively avoiding issues that are out of context and improving the professionalism and relevance of the answers.
[0057] Step 108: Determine the target interaction intent of the multimodal fusion information through the collaborative mechanism of instruction matching and agent reasoning, and execute the corresponding target action according to the target interaction intent to obtain the action execution result.
[0058] Among them, intelligent agent reasoning refers to the method of achieving intention reasoning through target intelligent agents. Target intelligent agents refer to intelligent models or intelligent agent systems deployed in the cloud or locally that have multimodal understanding capabilities. These intelligent agents can receive fused inputs containing multiple modal information such as images, text, and course context, and perform semantic analysis and reasoning output. Their specific implementation may include course intelligent agents, multimodal large models, image classification models, formula recognition models, or combinations thereof, which are not limited here.
[0059] The target interaction intent refers to the user interaction purpose determined by comprehensively judging multimodal fusion information. The interaction intent includes at least two types: visual content question and answer and desktop action control. Visual content question and answer means that the user wants to obtain explanation, analysis or answer for the captured visual area. Desktop action control means that the user wants to control the desktop application to perform operations such as turning pages, closing windows or opening functions through voice commands.
[0060] The target action refers to the specific execution behavior triggered according to the determined target interaction intent. When the intent is desktop action control, the target action can be to execute the corresponding desktop control through inter-process communication or event emitter. When the intent is visual content question and answer, the target action can be to call the agent to generate answer content for the target visual area. The action execution result refers to the specific output generated after the target action is executed and can be fed back to the user, including the explanatory text generated by the agent, the analysis answer, or the status change notification after the desktop action is executed, which is not limited here.
[0061] Specifically, when the system enters the intent triage stage, the target agent does not directly complete the desktop interaction command matching. Instead, the local voice command control module first preprocesses the voice text and matches it with the interaction command library. For data that does not match the local desktop interaction command but carries image data or course context, it is then sent to the target agent to perform visual semantic analysis and answer generation. That is, through the collaborative mechanism of command matching and agent reasoning, accurate intent recognition in multiple scenarios is achieved.
[0062] By inputting multimodal fusion information into the target agent and matching it with the instruction library, the unified routing of voice control and visual question answering in the same intent routing process is realized. This enables the system to accurately respond to desktop control commands such as page turning and closing, and to generate explanations and analyses with course background for user-specified visual content. The two interaction modes can be seamlessly switched within the same framework, effectively solving the problem of low interaction efficiency caused by separating voice commands and image question answering into two independent entry points in traditional solutions.
[0063] In one or more embodiments of this specification, intent recognition is based on speech-text. Specifically, the speech-text is compared with desktop interaction commands in the interaction command library for consistency. If the speech-text matches the desktop interaction commands in the interaction command library, the target interaction intent is determined to be desktop action control; otherwise, it is determined to be visual content question-and-answer. Specific implementation methods are as follows: The mechanism of determining the target interaction intent of the multimodal fusion information through instruction matching and agent reasoning includes: The voice text is matched with instructions. If the voice text matches the desktop interaction instructions in the interaction instruction library, the target interaction intent is determined to be desktop action control. If no matching desktop interaction command is found, the multimodal fusion information is input into the target agent to determine that the target interaction intent is visual content question answering.
[0064] Instruction matching refers to the process of comparing the voice text with the instructions in the preset interactive instruction library. In this embodiment, instruction matching is performed after the voice text is preprocessed. Preprocessing includes operations such as removing punctuation marks and normalizing pinyin. The matching method can be exact matching, fuzzy matching, or semantic similarity matching, etc., which are not limited here.
[0065] An interactive instruction library refers to a pre-built collection of various desktop interactive commands and their corresponding trigger words or phrases. This library includes at least page-turning commands (next page, previous page), window control commands (such as closing a window, opening a bullet screen), and wake-up commands, with each command bound to a specific desktop action. Desktop interactive commands refer to the command bars in the interactive instruction library that belong to the desktop action control category. Desktop action control refers to the type of interactive intent aimed at controlling the behavior of desktop applications. The target actions under this intent include, but are not limited to, slide turning, window closing, and function activation, and are executed through inter-process communication or event emitters. Visual content question answering refers to the type of interactive intent aimed at obtaining explanations, analyses, or answers for the captured target visual area. The target action under this intent is to call the target agent to perform visual semantic analysis on the multimodal fusion information and generate answer content.
[0066] Specifically, the voice command control module preprocesses the speech text in the multimodal fusion information, such as removing punctuation and performing pinyin normalization. The preprocessed text is then matched against desktop interaction commands in the interaction command library. When a match is successful, meaning the speech text matches one or more desktop interaction commands in the library, the system determines the current target interaction intent as desktop action control. At this point, the system does not enter the visual question-and-answer generation path of the target agent, but instead prepares to execute the corresponding desktop control action based on the matched command.
[0067] When a match fails, meaning the voice text fails to match any desktop interaction command in the interaction command library, the system retains the image data and course context in the multimodal fusion information and sends this multimodal fusion information to the target agent, which then performs visual semantic analysis and generates a response.
[0068] For example, during a presentation, if a teacher selects a chart and says "next page," the system removes punctuation and normalizes the pronunciation of the spoken text before matching it against the interaction command library. "Next page" matches a page-turning desktop interaction command, so the system determines the target interaction intent is desktop action control and then executes the slide-turning operation. Similarly, if a teacher selects a question and says "How do I do this question?", the spoken text "How do I do this question?" does not match any desktop interaction command in the interaction command library. The system determines the target interaction intent is visual content question-and-answer, and calls the target agent to generate a solution based on the question screenshot and the course context. Furthermore, in a multi-screen presentation scenario, if the presenter selects a data chart on the extended screen and says "full screen display," the spoken text matches a full-screen command, and the system determines the target interaction intent is desktop action control and executes full-screen switching. However, if the presenter says "What's wrong with this data?", the system determines the target interaction intent is visual content question-and-answer and performs analysis.
[0069] In practical applications, when users input voice but do not provide image data, the system can directly match the voice input data with commands. For example, if the voice input data is "close window", "close window" matches the desktop interaction command of window control, and the system determines that the target interaction intent is desktop action control and executes the window closure. However, when the user says "how to operate this interface", the voice text does not match the desktop command, and the system determines that the intent is visual content question and answer, and calls the intelligent agent to generate operation guidance.
[0070] The data processing method provided in the embodiments of this specification achieves intent separation between desktop action control and visual content question answering by setting a pre-configured instruction matching mechanism in the target intelligent agent. This eliminates the need for the target intelligent agent to perform complex visual semantic analysis on all voice inputs, ensuring both the response speed of desktop interaction instructions and that ambiguous colloquial questions can enter the visual content question answering path to obtain understanding and answers based on screen content. The separation logic of the two interaction modes is simple and clear, and can adapt to the actual needs of teachers frequently alternating between control instructions and content questions in teaching scenarios.
[0071] In one or more embodiments of this specification, different target actions are executed for different interaction intentions to obtain corresponding action execution results. The specific implementation methods are as follows: The step of executing the corresponding target action according to the target interaction intent and obtaining the action execution result includes: When the target interaction intent is desktop action control, the target desktop action is executed on the target client according to the desktop interaction instruction to obtain the desktop interaction result; When the target interaction intent is visual content question answering, the target agent generates an answer based on the multimodal fusion information to obtain the target solution result.
[0072] Among them, the target desktop action refers to the specific application layer operation corresponding to the matched desktop interaction command. This operation includes, but is not limited to, slide turning (next page, previous page), window management (closing window, switching windows), and function toggles (turning on bullet comments, turning off bullet comments, turning on name tag, error analysis), etc. Each desktop interaction command corresponds to one target desktop action.
[0073] Desktop interaction results refer to the execution status feedback obtained by the system after the target desktop action is completed, including the status indicator of whether the operation was successful or failed, as well as the change information of the desktop environment after the operation is completed, such as the current page number after turning the page, the status change notification after the window is closed, etc.
[0074] Answer generation refers to the process by which the target agent, after receiving multimodal fusion information, combines visual content from image data, the intent of the question in speech text, and contextual information of the course to perform comprehensive semantic understanding and reasoning, and finally outputs a response in natural language form. The target solution result refers to the specific content output by the target agent after the answer is generated. This content includes at least the explanatory text, analysis conclusions, or solutions to the user's question, but is not limited here.
[0075] Specifically, when the target interaction intent is desktop action control, the system determines the corresponding target desktop action based on the hit desktop interaction command. Then, it sends an action execution signal to the corresponding module of the target client via inter-process communication (IPC) or an event emitter. For example, it sends a page-turning command to the presentation plugin communication module or a window management command to the main process. After the target client executes the target desktop action, it can return the action execution status, thus providing the system with a desktop interaction result containing information on whether the operation was successful or not, as well as information on environmental changes.
[0076] When the target interaction intent is visual content question answering, the multimodal fusion information is completely transmitted to the target agent. The target agent performs semantic analysis on the visual region corresponding to the image data. If necessary, it can combine visual semantic capabilities such as text recognition, formula structure analysis, or chart feature extraction. At the same time, it combines the question intent expressed by the voice text and the course context information to perform comprehensive reasoning, generate specific answer content for the captured target visual region, and finally output the target answer result in structured or natural language form.
[0077] The data processing method provided in the embodiments of this specification, through the desktop action control path, can achieve rapid and accurate response to desktop interactive operations such as page turning and window management, ensuring the execution efficiency of control commands; through the visual content question-and-answer path, it can generate answer content through the target intelligent agent, making the answer content visually based and pedagogically targeted; the two execution paths are distinguished according to different intent types, which not only meets the immediacy requirements of high-frequency desktop control operations in teaching scenarios, but also ensures the quality of in-depth question-and-answer for visual content, achieving a unity of desktop control efficiency and content understanding depth.
[0078] In one or more embodiments of this specification, upon obtaining the action execution result, the action execution result is fed back to the target client. Specific implementation methods are described below: After executing the corresponding target action according to the target interaction intent and obtaining the action execution result, the method further includes: The results of the action execution are fed back to the target client in the form of message stream display, speech synthesis playback, and / or digital human broadcast.
[0079] Among them, message flow display refers to arranging and displaying the results of action execution in the message flow area in the form of text, Markdown text, or mixed text and images in chronological order, allowing users to view the interaction results by reading; speech synthesis playback refers to synthesizing the text content in the action execution results into natural language audio signals through text-to-speech technology, and playing them through audio output devices such as speakers or headphones, allowing users to obtain the result information by hearing without reading the screen; digital human broadcast refers to presenting the results of action execution through the voice broadcast of a virtual digital human figure. During the broadcast, the digital human coordinates lip movements and body movements to convey the result content to the user in a more human-like way.
[0080] Specifically, after obtaining the action execution result, the system can choose to use one or more feedback methods to execute the output based on the result. For the message stream display method, the system encapsulates the action execution result in a message stream format, presenting the content in text, Markdown format, and images. For the speech synthesis playback method, the system sends the text content from the action execution result to a text-to-speech module, which synthesizes the text into speech audio and plays it in real time. For the digital human broadcasting method, the synthesized speech audio can be synchronously driven with the digital human's lip-sync animation and body movements to complete the result broadcast in a human-like visual form. It should be noted that the three feedback methods can be used individually or in combination according to the needs of the scenario. For example, in a classroom setting, message stream display and speech synthesis playback can be enabled simultaneously, allowing teachers to view text content and receive information through hearing.
[0081] The data processing method provided in the embodiments of this specification allows for flexible configuration and combination of various feedback methods, including message flow display, speech synthesis playback, and digital human broadcasting. This enables the action execution results to be presented to the user in a form most suitable for the current interaction scenario and user needs. It not only preserves the integrity of the information displayed in the text but also provides the convenience of auditory broadcasting. Furthermore, it enhances the anthropomorphism and friendliness of the interaction through the digital human image. The selectivity and combinability of the three methods enable the system to adaptively adjust the feedback strategy in different scenarios such as classroom teaching, student self-study, and remote training, significantly improving the experience quality and scenario adaptability of desktop intelligent interaction.
[0082] This specification provides a data processing system based on desktop visual semantic analysis. The system includes modules such as a screenshot visual acquisition module, a visual context encapsulation module, a speech recognition module, a voice command control module, a desktop context module, a multimodal context construction module, an intent routing module, an agent generation module, and a feedback execution module.
[0083] The screenshot visual acquisition module is used to create multi-screen transparent screenshot windows, supporting selection by box, full screen, proportional selection, pen annotation, and selection size display, capturing the target visual area specified by the user in the desktop environment; the visual context encapsulation module is used to convert the selected content into an image file, upload the image URL, and record the image width, height, proportion, and command entry, encapsulating the visual area into structured image data; the speech recognition module is used to collect the user's speech and transcribe it into text, converting the speech input into speech text that can be processed by the system; the voice command control module is used to remove punctuation, normalize pinyin, and match the command library to the speech text, recognizing desktop interaction commands such as wake-up, page turning, closing, and error analysis.
[0084] The desktop context module maintains the presentation status, current courseware, current page, course configuration, and window status, providing status information for the application's runtime environment. The multimodal context construction module combines Markdown images, voice questions, text questions, course context, and extended parameters into unified multimodal fusion information for the agent to comprehensively understand. The intent routing module distinguishes different target interaction intents such as knowledge explanation, screenshot Q&A, voice control, contextual Q&A, error analysis, and ordinary chat. The agent generation module calls the course agent or multimodal model to perform visual semantic analysis, content explanation, question and answer generation, and inference output to generate answers to user questions. The feedback execution module executes answer display, voice broadcasting, and desktop actions through message streams, TTS (Text To Speech), digital humans, IPC (Inter-Process Communication), or emitters, feeding back the action execution results to the target client in various forms.
[0085] The following provides a detailed explanation of the data processing methods using the aforementioned data processing system. (See also...) Figure 2 , Figure 2 This diagram illustrates the process of obtaining the visual semantic context of a screenshot in a data processing method provided in one embodiment of this specification.
[0086] Specifically, when the user selects to screenshot the Q&A, a screenshot event is triggered and a screenshot window is displayed. The screenshot window is created based on the current display information and a screen background is loaded. This screenshot window is a transparent, top-mounted window.
[0087] When a screenshot layer is created, the user enters the "Select / Annotate Area" stage. Users can drag and drop the mouse to select, drag, zoom, select a full-screen area, or annotate with a brush. The screenshot window simultaneously displays the selection size and supports brush editing. After the user confirms the selection, the user enters the "Generate Image File" stage. The screenshot window draws the selection onto the canvas and converts the canvas data into an image file. Next, the user enters the "Upload to Obtain URL" stage, where the image file is uploaded to object storage to obtain the image URL. Finally, the user enters the "Return to Presenter" stage, where the screenshot window sends the image data to the presenter window (i.e., the server). The image data includes the source, command, image URL, image width, and image height. Specifically, the user-uploaded image address and size information can be transmitted to the presenter in the form of a "visual semantic carrier," preserving the selection range, proportions, and visual content. Alternatively, the input can be routed based on the command type at the screenshot scene entry point. If the command is "KnowledgeExplanation" (i.e., the explanation command in the above embodiment), the user enters the direct explanation path; if the command is "AskAI" (i.e., the inquiry command in the above embodiment), the user enters the voice / text input space.
[0088] Through this process, the system can respond to the screenshot operation triggered by the user on the target client, create a screenshot window, and receive the user's interactive operations on the desktop area through the screenshot window. Based on the interaction results, the system determines the target visual area, renders the target visual area as an image file, and obtains image data containing image address, image width information, and image height information. This enables accurate acquisition and encapsulation of the user-specified visual area in a multi-window, multi-screen environment on the desktop.
[0089] See Figure 3 , Figure 3 This diagram illustrates a flow chart of intent identification and execution triage in a data processing method provided in one embodiment of this specification.
[0090] Upon receiving voice input data, the system can obtain the voice recognition text (i.e., voice text) through real-time voice recognition. Then, it enters the instruction library matching stage. The specific voice instruction control module performs preprocessing operations such as removing punctuation, normalizing pinyin, and deduplicating general instructions on the voice text, and matches the processed text with the interactive instruction library.
[0091] After matching is completed, the system proceeds to the command judgment stage to determine whether the voice or text matches the control commands in the interaction command library (including desktop interaction commands such as wake-up, page turning, closing, and error analysis). The system also checks whether there is a screenshot context (i.e., whether the user has taken a screenshot and entered the AskAI or KnowledgeExplanation scenario), that is, whether it belongs to a multimodal question scenario with an image. If a screenshot context exists, the system constructs a multimodal question by combining the image URL and aspect ratio with the voice / text question (i.e., the multimodal information in the above embodiment).
[0092] When voice-text matches a control command, the system determines the target interaction intent as desktop action control and enters the desktop action execution path. Specifically, it performs desktop control operations such as page turning, window closing, and enabling subtitles via IPC or emitter, obtaining the desktop interaction result. Of course, if a semantic response needs to be generated within this desktop action path, it can also enter the semantic response path.
[0093] When the voice / text does not match any desktop interaction command, and there is a screenshot context that constitutes a multimodal question with an image, the system determines that the target interaction intent is a visual content question and answers, enters the semantic answer generation path, submits the multimodal question containing images and voice / text to the target agent, calls the target agent to perform visual semantic analysis, content explanation, question and answer generation and reasoning output, and finally returns the target answer result in the form of streaming text or TTS.
[0094] This routing process unifies voice command matching, screenshot context judgment, visual question answering, and desktop action execution into a single intent routing process, achieving unified intent routing for voice control and visual question answering, while ensuring the response speed of control commands and the depth and quality of visual question answering.
[0095] See Figure 4 , Figure 4 A schematic flowchart of data fusion in a data processing method provided in one embodiment of this specification is shown.
[0096] In practice, visual input, voice input, and desktop context are the starting points. Visual input includes screenshots and visual semantic information extracted from them, such as OCR text, graphics, formulas, tables, and pen-annotated areas (as visual cues); voice input includes teacher questions, voice commands, and real-time recognized text (as voice cues); and desktop context includes desktop application status information such as PPT presentation status, current courseware, current page, course agent, teacher conversation history, window status, and reporting entry point.
[0097] In practice, when a user selects a target visual area through the screenshot window, the system uses the image URL, aspect ratio, and selection command as the visual context. With voice input data available, this allows the system to associate pronouns in the speech with specific visual areas. It's important to note that after taking a screenshot, the user can first ask a question via voice, and then switch to text editing to add details. The system retains the image URL and aspect ratio, submitting the edited text along with the image to ensure no loss of visual context.
[0098] After the above three types of inputs enter the "context fusion and disambiguation" stage, the system will fuse the URL of the screenshot image with the aspect ratio, voice text and course context information. Taking the disambiguation of colloquial pronouns such as "this," "here," and "this question" as an example, by matching the pronouns in the voice with the visual area bound to the screenshot image, the ambiguous colloquial expression can obtain a clear screen content reference.
[0099] The fused multimodal information is then directed to three paths in the intent triage process: the first is the "visual question-and-answer" path, which is used to explain the screenshot content, analyze the chart formulas or questions; the second is the "contextual question-and-answer" path, which is used to generate a broadcastable answer by combining the current course context and historical conversations; and the third is the "voice control" path, which is used to execute control commands such as opening, closing, page turning, and error analysis.
[0100] By deeply integrating visual input, voice questions, and desktop context, rather than processing voice commands or screenshots in isolation, the system can use desktop visual semantics as the basis for intent understanding, achieving joint understanding and unified routing of visual regions, voice questions, and course background.
[0101] See Figure 5 , Figure 5 A schematic flowchart of dual-path processing in a data processing method provided in one embodiment of this specification is shown.
[0102] Users can select any area on the desktop, such as PPT, desktop, questions, or charts, and can annotate and edit the selected area with a pen. After the system obtains the image data corresponding to the target visual area, it splits the data into two processing paths based on the type of command triggered by the user in the screenshot window: knowledge explanation and AI question.
[0103] In the knowledge explanation path, the system responds to the explanation instruction (KnowledgeExplanation) by directly encapsulating image data containing image addresses and aspect ratios into Markdown image prompts, and sending them to the target agent along with the `knowledge_explain` extended parameter as explanation mode identifier data. The target agent then generates an automatic explanation of the visual content in the image data based on visual semantics. After generating the explanation content, it can be played back to the user via TTS (Text-to-Speech), while the screenshot window closes or restores the classroom window state. This path eliminates the need to wait for additional input, achieving an instant response from selection to explanation.
[0104] In the AI-driven approach (i.e., the inquiry path), the system responds to the AskAI command and saves the image data as visual context. It then initiates voice input, waiting for the user to speak their question, and can switch to text input while retaining the image data if needed. After the user inputs their question via voice or text, the system combines the image from the image data with the voice or text question to create a multimodal input with image-based follow-up questions. It generates a contextual response that integrates visual content and question semantics, outputs the target solution, and provides feedback to the user through message flow display, speech synthesis playback, or digital human narration.
[0105] The two paths cater to two different teaching and interaction needs: "instant explanation without follow-up questions" and "in-depth questions based on images." This approach improves operational efficiency in direct explanation scenarios while also ensuring flexibility in in-depth interaction scenarios.
[0106] See Figure 6 , Figure 6 This document illustrates a flowchart of an interactive closed loop in a data processing method provided in one embodiment of this specification.
[0107] In the initial stage, the presenter window is in standby or floating expanded state. When the user triggers a screenshot operation, it enters the "Screenshot Activation" state, and the system creates a transparent, top-mounted selection window and displays the screenshot interface. After the user completes the selection annotation and confirms, it enters the "Visual Context Ready" state. The system uploads the selected image and obtains the image URL and size information. The screenshot window sends the image back to the presenter window, and the visual semantic context is ready.
[0108] The process then proceeds to the "Voice Questioning / Control" state based on the user's interaction method. Users can ask questions or issue control commands via voice input or text editing. After voice input, the system enters the "Agent Generation" state, submitting multimodal fusion information and awaiting the agent's streaming response or deep thinking results. Once generated, the system enters the "Feedback Execution" state, presenting the response or execution result to the user through TTS voice playback, digital human narration, or desktop action execution, while simultaneously playing the TTS message. Finally, after feedback execution, the system enters the "Restore Desktop" state, closing the screenshot window or restoring the classroom window. The system simultaneously closes the screenshot, completing the full closed loop from screenshot activation to feedback execution and desktop restoration.
[0109] Throughout the entire process, the "anomaly and recovery control" mechanism is consistently implemented, supporting anomaly handling and status control operations such as closing the screenshot recovery window, stopping generation, deleting images, switching between text / voice input, stopping TTS playback, and reporting message status, ensuring the stability of multi-window interaction.
[0110] This closed-loop mechanism enables the screenshot window, the narrator window, and the main process to synchronize their states such as screenshot ready, screenshot closed, image return, voice recognition, stop generation, and TTS playback through IPC messages. This avoids problems such as loss of screenshot results, window restoration errors, or interaction interruption in multi-window interaction, and ultimately forms a complete interactive closed loop from screenshot, upload, return, Q&A, broadcasting to closing and restoration.
[0111] The data processing method provided in the embodiments of this specification uses the selected area of the desktop screenshot as the visual semantic context of voice interaction, enabling referential words such as "this" and "here" in the voice to be bound to specific screen content, significantly improving the accuracy of voice interaction. Teachers can directly select visual content and ask follow-up questions by voice, without having to manually describe complex charts, formulas, or questions, effectively reducing the cost of classroom operations. At the same time, the intelligent agent's answer can refer to the screenshot content, the current PPT, and the course context, making the answer highly relevant to the teaching context. On this basis, the capabilities of the desktop voice assistant are expanded, not only understanding voice commands but also understanding the desktop content specified by the user. Furthermore, the complete link formed by screenshotting, uploading, feedback, questioning and answering, broadcasting, and closing and restoring enhances the stability of multi-window desktop applications and ensures continuous and reliable interaction.
[0112] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 7 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 7 As shown, the device includes: The acquisition module 702 is configured to acquire image data corresponding to a target visual region, wherein the target visual region is determined by a screenshot window, and the screenshot window is created in response to a screenshot trigger operation of the target client; The recognition module 704 is configured to perform speech recognition on the speech input data to obtain speech text when the speech input data is acquired, wherein the speech input data is used to ask follow-up questions to the image data and / or trigger desktop interaction commands; The fusion module 706 is configured to acquire the application status information of the target client and fuse the image data, the voice text, and the application status information to obtain multimodal fusion information; The execution module 708 is configured to determine the target interaction intent of the multimodal fusion information through a collaborative mechanism of instruction matching and agent reasoning, and to execute the corresponding target action according to the target interaction intent to obtain the action execution result.
[0113] Optionally, the acquisition module 702 is further configured to: The screenshot window receives user interaction with the desktop area and determines the target visual area based on the interaction results. The target visual region is rendered into an image file to obtain the image data, wherein the image data includes the image address and image size information of the image file.
[0114] The device further includes: The splitting module is configured to, in response to a narration command, input the image data and narration mode identifier data into the target agent to obtain the narration content of the target agent for the image data; and in response to a query command, save the image data and obtain the voice input data and / or text input data.
[0115] Optionally, the fusion module 706 is further configured to: The image address and image size information corresponding to the image data are fused with the voice text and the course context information to obtain the multimodal fusion information.
[0116] Optionally, the execution module 708 is further configured to: The voice text is matched with instructions. If the voice text matches the desktop interaction instructions in the interaction instruction library, the target interaction intent is determined to be desktop action control. If no matching desktop interaction command is found, the multimodal fusion information is input into the target agent to determine that the target interaction intent is visual content question answering.
[0117] Optionally, the execution module 708 is further configured to: When the target interaction intent is desktop action control, the target desktop action is executed on the target client according to the desktop interaction instruction to obtain the desktop interaction result; When the target interaction intent is visual content question answering, the target agent generates an answer based on the multimodal fusion information to obtain the target solution result.
[0118] The device further includes: The feedback module is configured to feed back the action execution result to the target client in the form of message stream display, speech synthesis playback, and / or digital human broadcast.
[0119] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.
[0120] Figure 8 A structural block diagram of a computing device 800 according to one embodiment of this specification is shown. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.
[0121] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0122] In one embodiment of this specification, the above-described components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0123] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server.
[0124] The processor 820 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above data processing method.
[0125] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.
[0126] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.
[0127] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.
[0128] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.
[0129] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the data processing method described above.
[0130] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0131] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0132] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0133] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0134] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data processing method based on visual semantic analysis, characterized in that, include: Acquire image data corresponding to the target visual region, wherein the target visual region is determined by a screenshot window, and the screenshot window is created in response to a screenshot trigger operation of the target client; Upon obtaining voice input data, voice recognition is performed on the voice input data to obtain voice text, wherein the voice input data is used to ask follow-up questions about the image data and / or trigger desktop interaction commands; The application status information of the target client is obtained, and the image data, the voice text, and the application status information are fused to obtain multimodal fusion information; By using a collaborative mechanism of instruction matching and agent reasoning, the target interaction intent of the multimodal fusion information is determined, and the corresponding target action is executed according to the target interaction intent to obtain the action execution result.
2. The data processing method according to claim 1, characterized in that, The acquisition of image data corresponding to the target visual region includes: The screenshot window receives user interaction with the desktop area and determines the target visual area based on the interaction results. The target visual region is rendered into an image file to obtain the image data, wherein the image data includes the image address and image size information of the image file; When course content is displayed in the desktop area, the application status information includes the course context information of the course content. The process of fusing the image data, the voice text, and the application status information to obtain multimodal fusion information includes: The image address and image size information corresponding to the image data are fused with the voice text and the course context information to obtain the multimodal fusion information.
3. The data processing method according to claim 1, characterized in that, After acquiring the image data corresponding to the target visual region, the process further includes: In response to the explanation instruction, the image data and explanation mode identification data are input into the target agent to obtain the explanation content of the target agent for the image data; In response to an inquiry command, the image data is saved and the voice input data and / or text input data are acquired.
4. The data processing method according to claim 1, characterized in that, The mechanism of determining the target interaction intent of the multimodal fusion information through instruction matching and agent reasoning includes: The voice text is matched with instructions. If the voice text matches the desktop interaction instructions in the interaction instruction library, the target interaction intent is determined to be desktop action control. If no matching desktop interaction command is found, the multimodal fusion information is input into the target agent to determine that the target interaction intent is visual content question answering.
5. The data processing method according to claim 4, characterized in that, The step of executing the corresponding target action according to the target interaction intent and obtaining the action execution result includes: When the target interaction intent is desktop action control, the target desktop action is executed on the target client according to the desktop interaction instruction to obtain the desktop interaction result; When the target interaction intent is visual content question answering, the target agent generates an answer based on the multimodal fusion information to obtain the target solution result.
6. The data processing method according to any one of claims 1-5, characterized in that, After executing the corresponding target action according to the target interaction intent and obtaining the action execution result, the method further includes: The results of the action execution are fed back to the target client in the form of message stream display, speech synthesis playback, and / or digital human broadcast.
7. A data processing apparatus, characterized in that, include: The acquisition module is configured to acquire image data corresponding to a target visual region, wherein the target visual region is determined by a screenshot window, and the screenshot window is created in response to a screenshot trigger operation from the target client; The recognition module is configured to perform speech recognition on the speech input data to obtain speech text when the speech input data is acquired, wherein the speech input data is used to ask follow-up questions to the image data and / or trigger desktop interaction commands; The fusion module is configured to acquire the application status information of the target client and fuse the image data, the voice text, and the application status information to obtain multimodal fusion information; The execution module is configured to determine the target interaction intent of the multimodal fusion information through a collaborative mechanism of instruction matching and agent reasoning, and to execute the corresponding target action according to the target interaction intent to obtain the action execution result.
8. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, characterized in that, when the computer programs / instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 1 to 6.