Processing method and interaction system
By dynamically analyzing and recognizing labels on the display interface, the system responds to handwriting pen touch operations in real time and dynamically loads the interactive interface, solving the problem of disconnect between functions and interface content in handwriting pen interaction and achieving a natural interactive experience where you can write as soon as you pick up the pen.
Patent Information
- Application Number
- CN202511029894.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-21
Smart Images

Figure CN120994083A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technology, and includes, but is not limited to, a processing method and an interactive system. Background Technology
[0002] With the technological advancements in mobile computing devices, users are placing higher demands on the human-computer interaction of touch-screen devices such as tablets, leading to the widespread adoption of styluses as a natural and efficient input tool. Existing technologies mostly trigger stylus functions through preset gestures or multiple clicks, but lack dynamic perception of interface content and the ability to automatically adjust function responses. This results in long interaction paths, cumbersome function triggering, and negatively impacts interaction efficiency and naturalness. Therefore, a method to address these technical issues is urgently needed to improve the overall user experience. Summary of the Invention
[0003] In view of this, embodiments of this application provide a processing method and an interactive system.
[0004] The technical solution of this application embodiment is implemented as follows:
[0005] In a first aspect, embodiments of this application provide a processing method, including:
[0006] The system obtains a touch operation from the operator on the display interface; in response to the touch operation being located in the first display content area, it displays a first interactive interface based on a first identification tag; in response to the touch operation being located in the second display content area, it displays a second interactive interface based on a second identification tag; wherein the first identification tag and the second identification tag are different, and the first identification tag and the second identification tag are based on the identification and marking of the currently displayed content of the display interface by the electronic device system, and the first interactive interface and the second interactive interface are different.
[0007] Secondly, embodiments of this application provide a processing apparatus, including:
[0008] The first acquisition module is used to acquire the touch operation of the operator on the display interface;
[0009] The first display module is configured to respond to a touch operation located in the first display content area and display a first interactive interface based on a first identification label.
[0010] The second display module is used to respond to a touch operation in the second display content area and display a second interactive interface based on a second identification tag; wherein the first identification tag and the second identification tag are different, and the first identification tag and the second identification tag are based on the electronic device system to identify and mark the currently displayed content of the display interface, and the first interactive interface and the second interactive interface are different.
[0011] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it enables the acquisition of a touch operation by an operator on a display interface. In response to a touch operation located in a first display content area, a first interactive interface is displayed based on a first identification tag. In response to a touch operation located in a second display content area, a second interactive interface is displayed based on a second identification tag. The first identification tag and the second identification tag are different, and the first identification tag and the second identification tag are based on the electronic device system's identification and marking of the currently displayed content of the display interface. The first interactive interface and the second interactive interface are different.
[0012] Fourthly, embodiments of this application provide a storage medium storing executable instructions, which, when executed by a processor, enable the acquisition of a touch operation by an operator on a display interface; in response to a touch operation located in a first display content area, display a first interactive interface based on a first identification tag; in response to a touch operation located in a second display content area, display a second interactive interface based on a second identification tag; wherein the first identification tag and the second identification tag are different, and the first identification tag and the second identification tag are identified and marked by the electronic device system based on the current display content of the display interface, and the first interactive interface and the second interactive interface are different.
[0013] Fifthly, embodiments of this application provide a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, they enable the acquisition of a touch operation by an operator on a display interface. In response to a touch operation located in a first display content area, a first interactive interface is displayed based on a first identification tag. In response to a touch operation located in a second display content area, a second interactive interface is displayed based on a second identification tag. The first identification tag and the second identification tag are different, and the first identification tag and the second identification tag are identified and marked by the electronic device system based on the current display content of the display interface. The first interactive interface and the second interactive interface are different. Attached Figure Description
[0014] Figure 1 A schematic diagram illustrating the implementation flow of a processing method provided in an embodiment of this application;
[0015] Figure 2 A schematic diagram illustrating the implementation process of loading a target model, provided in an embodiment of this application;
[0016] Figure 3A A schematic diagram illustrating the implementation flow of the widget icon processing method provided in this application embodiment;
[0017] Figure 3B This is a schematic diagram illustrating the implementation process of loading different models for different regions, as provided in the embodiments of this application.
[0018] Figure 4 A schematic diagram of a three-level architecture of a handwriting pen interaction system provided in this application embodiment;
[0019] Figure 5 This is a schematic diagram of the composition structure of a processing device provided in an embodiment of this application;
[0020] Figure 6 This is a schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of the embodiments will be further described in detail below with reference to the accompanying drawings. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0022] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0023] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0025] In related technologies, the triggering of handwriting pen interaction functions often relies on fixed UI controls or preset gestures, resulting in a disconnect between function and interface content. Users need to go through a lengthy interaction path to complete the operation, which is costly to learn and lacks flexibility. For example, users need to click or swipe the edge multiple times to activate the annotation mode, and the handwriting pen function cannot be intelligently adapted to the current interface content, resulting in a poor user experience.
[0026] To address the aforementioned issues, this application proposes an active pen-on-write method. By combining automatic recognition of the currently displayed content with real-time response to stylus touch operations, it achieves a natural interactive experience for writing instantly upon lifting the pen. Specifically, the system dynamically analyzes the display interface to identify different functional areas and generates corresponding recognition tags for each area. When a user approaches or touches a certain area with a stylus, the system can load the corresponding interactive interface and target model based on the recognition tags, thereby achieving a highly personalized interactive effect that closely matches the scenario. This method not only shortens the user's operation path but also enhances the intelligence and naturalness of the interaction, effectively solving the problem of function-scenario mismatch in traditional stylus interaction.
[0027] The interface display methods provided in the embodiments of this application can be executed by an electronic device, which can be a mobile terminal with touch display function such as a tablet computer or a smartphone.
[0028] Figure 1 This application provides a flowchart illustrating a processing method, as shown below. Figure 1 As shown, the method includes:
[0029] Step S110: Obtain the touch operation of the operating body on the display interface;
[0030] Here, the operating object can be a stylus. The operating object, through its built-in sensors or a communication interface with the electronic device system, enables the electronic device system to perceive the position and movement of the operating object. In another embodiment, the operating object can be a finger or palm; for example, the position and movement of the operating object can be perceived by the electronic device system using sensors worn by the user or the electronic device's own sensors. When the user touches the screen with the operating object, the electronic device system can collect the position information of the operating object in real time and determine whether the operating object has triggered a specific display content area. For example, when the user taps a certain area on the screen with the operating object, the electronic device system can recognize this touch operation and record the coordinate information of the operating object. Touch operations include, but are not limited to, tapping, swiping, pressing, and hovering actions, which are recognized and converted into input signals by the touchscreen sensors. The system analyzes the position, trajectory, and time information of the touch operation to determine the user's current operating intention and interaction area. For example, when the user approaches the screen with a stylus but has not yet touched it, the system can perceive the proximity of the operating object by the change in the distance between the stylus tip and the screen and prepare subsequent response logic. For example, when a user brings their finger close to the screen but hasn't yet made contact, the system can use the proximity sensor on the screen to sense the proximity of the object and prepare subsequent response logic. The behavior pattern of the object in different scenarios will affect the system's response strategy. For example, in a text editing area, the user may only need to tap to insert the cursor; while in an image display area, the user may need to swipe or select a specific area to trigger image processing functions. Therefore, touch operation is not only the starting point of interaction but also an important basis for the system to determine the next action.
[0031] During implementation, the accuracy and response speed of touch operation recognition directly affect the user experience. Through highly sensitive touch sensors and real-time data processing algorithms, the system can capture and analyze touch events within milliseconds, thus providing a reliable data foundation for subsequent identification tag generation and interactive interface loading.
[0032] Step S120: In response to the touch operation being located in the first display content area, display the first interactive interface based on the first identification tag;
[0033] Here, the display content area can be different areas on the electronic device screen divided according to function, such as a text editing area, an image display area, and a comment / interaction area. Each display content area has different interaction requirements and functional characteristics, and the electronic device system can automatically identify and generate corresponding identification tags based on the currently displayed content. Identification tags, such as structured identification information generated based on visual semantic analysis, are used to distinguish different functional areas and are used for subsequent loading of corresponding interactive interfaces or functional modules.
[0034] The first identification tag is a structured identifier generated by an electronic device for the first display content area. The first identification tag can be configured to include the functional attributes, interaction methods, and associated target model information of the first display content area. For example, in a text editing area, the identification tag might include descriptions of functions such as text input, cursor positioning, and text conversion; in an image display area, the identification tag might include descriptions of functions such as image enhancement, filter application, and annotation addition.
[0035] Upon receiving a touch operation, the electronic device system first determines whether the touch operation falls within a specific display content area. If the touch operation is confirmed to be within a first display content area, the electronic device system will retrieve the identification tag corresponding to the first display content area and load the corresponding interactive interface based on the information from the identification tag. For example, in a text editing area, the system might pop up an interactive interface containing a text input box, a font selector, and a voice input button to allow the user to quickly perform text input operations.
[0036] During implementation, electronic device systems can accurately identify standard or non-standard control areas (such as temporary annotation areas drawn by users or custom views of third-party applications) through dynamic interface semantic segmentation technology, thereby ensuring the applicability and accuracy of the identification labels. Simultaneously, by pre-loading relevant model resources, the system can respond quickly and provide a complete interactive experience when touch operations occur.
[0037] Step S130: In response to the touch operation being located in the second display content area, display the second interactive interface based on the second identification tag;
[0038] The first identification tag and the second identification tag are different. The first identification tag and the second identification tag are based on the electronic device system's identification and marking of the currently displayed content of the display interface. The first interactive interface is different from the second interactive interface.
[0039] Here, the second identification tag is similar to the first identification tag; it is also a structured identification information automatically generated by the electronic device system based on the currently displayed content. The second identification tag corresponds to the second display content area. Because the first and second identification tags are different, the electronic device system can load completely different interactive interfaces based on the differences between these two types of identification tags to meet the interactive needs of different display content areas.
[0040] For example, when the user switches from the text editing area to the image display area, the electronic device system loads a completely new interactive interface based on the second identification tag. This interactive interface may include image enhancement options, drawing tools, annotation functions, etc. A comparison reveals that the second interactive interface differs significantly from the first interactive interface in both content and form. This difference demonstrates that this embodiment achieves dynamic loading of the interactive interface based on the identification tag.
[0041] During implementation, electronic device systems, by integrating visual language models with context-aware technology, can achieve intelligent switching and seamless transitions across display content areas. For example, when a user's writing tool slides from the text editing area to the image display area, the system can not only switch the interactive interface but also preload relevant image processing models to ensure that the user receives instant feedback the moment they put pen to paper.
[0042] The interface display method provided in this application embodiment enables the electronic device system to generate different identification tags in different display content areas and dynamically load corresponding interactive interfaces based on these identification tags, thereby achieving a highly scene-appropriate personalized interactive effect. This not only shortens the user's operation path but also improves the intelligence and naturalness of the interaction, effectively solving the problem of function and scene mismatch in the interaction between the operator and the screen.
[0043] In some embodiments, the step S130 above, in which "the first identification tag and the second identification tag are identified and marked based on the electronic device system's identification of the currently displayed content of the display interface," can be achieved through the following steps:
[0044] Step S131: Obtain the screen data stream of the display interface;
[0045] Here, screen data stream refers to the display interface image information that electronic devices capture and transmit in real time during operation. Screen data streams are typically output continuously in the form of video frames, containing all graphic elements, text content, and user interface components currently displayed on the screen. By capturing screen data streams, the system can understand the content area the user is interacting with in real time, thus providing basic data support for subsequent area segmentation and function adaptation.
[0046] During implementation, screen data streams can be obtained through the operating system's Application Programming Interface (API) or the underlying graphics rendering engine. For example, in Android, applications can use mechanisms such as SurfaceFlinger or Virtio-GPU to obtain screen image streams; in iOS, applications can achieve similar functionality through the Core Graphics or Metal frameworks. Screen data streams can be raw pixel data or partially processed structured image information.
[0047] By acquiring screen data streams, the system can perceive the current state of the displayed interface in real time without relying on manual user triggering, laying the foundation for subsequent intelligent interaction functions.
[0048] Step S132: Divide the screen data stream into regions based on a preset model to obtain multiple regions with identification tags;
[0049] Different identification tags represent different interactive functions for corresponding areas, and the identification tags include the first identification tag and the second identification tag.
[0050] Here, the pre-trained model refers to a pre-trained machine learning model used for semantic segmentation and function recognition of the screen data stream. Based on a joint visual-language model, the pre-trained model combines image recognition and natural language processing capabilities to divide various areas of the display interface into blocks with different functional attributes. For example, the system may use text editing areas, image display areas, and comment / interaction areas as segmentation objects, assigning a unique identification label to each area.
[0051] The region segmentation process includes multiple stages such as image input, feature extraction, and classification prediction. The system first inputs the screen data stream into a preset model, which extracts image features using structures such as convolutional neural networks (CNN). The system then uses semantic segmentation algorithms to perform pixel-level classification of the image, and finally generates a functional block map with recognition labels.
[0052] By dividing the region based on a preset model, the system can break through the limitations of traditional user interface (UI) control recognition methods. It can not only recognize standard controls (such as buttons and text boxes), but also non-standard areas (such as hand-drawn annotation areas and custom views of third-party applications), significantly improving scene adaptability.
[0053] Identification tags are metadata used to identify the functional attributes of display interface areas. Each identification tag represents a specific interactive behavior or operation mode of an operator (such as a stylus). For example, the first identification tag in the text editing area may be associated with functions such as inserting a cursor and text conversion; the second identification tag in the image display area may be associated with functions such as artificial intelligence (AI) image enhancement and annotation; and the identification tags in the comment interaction area may be associated with functions such as automatic comment suggestions and emoticon recommendations.
[0054] The first and second identification tags represent two different types of interactive functions, which the system dynamically switches between based on the user's current intention and the context information of the displayed interface. The first interactive interface is displayed to interact with the first target model, and the second interactive interface is displayed to interact with the second model. For example, when the user's stylus approaches the text editing area, the system can load model resources related to text input; when the user's stylus approaches the image display area, the system can load model resources related to image processing.
[0055] By differentiating different identification tags, the system achieves dynamic function adaptation in multiple scenarios. The system avoids redundant interaction paths such as traditional fixed gestures or multiple clicks, making stylus operation more intuitive and efficient.
[0056] During implementation, there is a close relationship between the screen data stream, the preset model, and the recognition tags. The screen data stream serves as the input source, providing real-time image information to the preset model; the preset model is responsible for analyzing and semantically understanding the images, outputting functional areas with recognition tags; the recognition tags further guide the system on how to respond to user interaction requests, thus forming a closed-loop feedback intelligent interaction process.
[0057] In this embodiment, the system acquires screen data streams and divides the screen area based on a preset model to generate functional blocks with identification labels. This allows for dynamic adaptation of the stylus's interactive functions to different areas. It effectively solves the problem of disconnect between function triggering and the scenario in existing solutions, thereby achieving a natural interactive experience where users can write while using the stylus, further improving user efficiency and interaction satisfaction during operation.
[0058] In some embodiments, the first interactive interface is adapted to the currently executable task of the first display content area based on a first target model, and the second interactive interface is adapted to the currently executable task of the second display content area based on a second target model, wherein the first target model and the second target model are different.
[0059] Here, the first target model refers to a pre-configured computational model for handling interactive tasks in the first display content area, based on its functional attributes. For example, in a text editing area, the first target model could be a handwriting-to-text model; in an image display area, the first target model could be an image enhancement or annotation model. Each first target model is optimized for a specific user intent, enabling rapid reasoning and output of context-appropriate results.
[0060] There is a mapping relationship between the first target model and the first display content area. The system identifies the area type where the current touch operation is located by recognizing the label, and selects the first target model corresponding to the area type recognized by the system for loading and inference. The above dynamic adaptation mechanism enables the stylus function to intelligently switch according to different usage scenarios without requiring the user to manually switch operation modes or adjust relevant parameters.
[0061] Similar to the first target model, the second target model can be another set of functional models adapted for different display content areas. Since the first and second display content areas correspond to different functional scenarios, the target models adapted to these two display content areas are different to meet the needs of diverse users.
[0062] In this embodiment, dynamic switching and intelligent adaptation of touch operation functions are achieved by adapting different target models to different display content areas. By distinguishing different display content areas during operation and using target models that match each area for adaptation processing, intelligent adaptation of different functions of the operator can be achieved in various application scenarios. This adaptation method can improve the interaction flexibility between the system and the user, ultimately meeting the diverse needs of users in different usage scenarios.
[0063] This application also provides a method for loading a target model, such as... Figure 2 As shown, this can be achieved through the following steps:
[0064] Step S210: In response to the operator approaching but not touching the target display content area, preload at least part of the target model data of the target interactive interface corresponding to the target display content area;
[0065] The operator can be an input tool used by the user to interact with the electronic device, such as a finger or a stylus. In some embodiments, the operator is communicatively connected to the electronic device. This communication connection can be achieved by establishing a data transmission channel between the operator and the electronic device via wired or wireless means. This allows the operator to provide real-time feedback of the user's intentions (such as movement trajectory, pressure level, etc.) to the electronic device and receive control commands or status information returned by the communication connection. Communication connection methods include, but are not limited to, Bluetooth, Universal Serial Bus (USB), and Near Field Communication (NFC) connections. By establishing a stable communication connection, the system can perceive user intentions in real time and respond quickly, thereby improving interaction efficiency and user experience.
[0066] The target display area can be a specific functional block on the screen that the user is currently focusing on or about to interact with, such as a text editing area, image display area, or comment / interaction area. When the user's input object approaches the target display area but has not yet actually touched it, the system performs semantic segmentation of the interface based on a visual language model. The system identifies the functional attributes of the target display area and preloads some of the model data required for the interactive interface related to the target display area. This preloading mechanism can significantly shorten the actual response time of subsequent operations because all resources are preloaded into memory, avoiding the time delay of waiting for the model to fully load.
[0067] When the operating object approaches the target display content area, the system preloads some data of the target model. For example, when a stylus is hovering over the screen of an electronic device, the system can prepare relevant resources in advance. When the user actually triggers the operation, it can achieve a faster response, thereby optimizing the interactive experience, reducing user waiting time, and improving overall efficiency.
[0068] Step S220: In response to the operator touching the target display content area, load all the data of the target model to display the target interactive interface, wherein the target display content area includes the first display content area and the second display content area, and the target interactive interface includes the first interactive interface and the second interactive interface.
[0069] When the user finally touches the target display area, the system switches from intent prediction to dynamic execution, loading all data from the target model and displaying the corresponding interactive interface. This target model's data includes complete functional modules, an AI inference engine, and relevant parameter configurations, ensuring the interactive interface has full functional support. The target display area encompasses multiple types of functional blocks; for example, text editing areas and image display areas correspond to different interactive interfaces. Each interactive interface can be customized for its specific functional blocks to provide the most suitable user experience.
[0070] By loading all the target model's data and displaying the target's interactive interface after the operator touches the target's display area, the system achieves a seamless transition from prediction to execution, enhancing the naturalness and intelligence of the interaction. The design of different interactive interfaces corresponding to different display areas further enhances the system's scene adaptability, enabling users to obtain a more accurate and efficient interactive experience.
[0071] In this embodiment, after the operator establishes a communication connection with the electronic device, the system can preload some data of the target model when the operator approaches the target display content area. Then, when the operator touches the target display content area, the system will quickly load all the data of the target model and display the target interactive interface. This significantly shortens the response time of user operations, thereby improving interaction efficiency, such as enabling a natural interactive experience where users can simply pick up a pen and use it.
[0072] In some embodiments, the preloading includes loading the static resources of the target model, including at least one of the following: model metadata, model weights, and dependency libraries; the loading includes loading the dynamic objects of the target model, including at least one of the following: model architecture, model instance, and runtime context.
[0073] Here, the target model can be an algorithm model prepared to support different functions, such as a text recognition model, an image processing model, and a speech synthesis model. Loading the static resources and dynamic objects of the target model can be done in advance before the actual operation occurs, loading the static resources (such as model metadata, model weights, and dependent libraries) and dynamic objects (such as model architecture, model instances, and runtime context) of the target model that may be needed, in order to reduce response latency.
[0074] During implementation, there is a logical relationship between coordinate data and the type or functional attributes of functional blocks. When the pen tip approaches the screen, the system locates the specific functional block using coordinate data, and combined with the type or functional attributes of the functional block, the system determines which static resources and dynamic objects of the target model need to be loaded, thus achieving efficient preloading.
[0075] In this embodiment, by predicting the type or functional attributes of the functional block and loading the static resources and dynamic objects of the target model during the approach phase to the screen, the system can respond quickly at the moment the pen touches the screen. This method avoids the delay problem caused by model loading in traditional solutions, thereby improving the smoothness of user-system interaction.
[0076] In some embodiments, the display interface is the main screen interface, and the target display content area is a widget icon display area. This application also provides a processing method based on widget icons, such as... Figure 3A As shown, this can be achieved through the following steps:
[0077] Step S301: Block the touch swipe event triggered by the operator touching the widget icon display area;
[0078] Here, the home screen is the default user interface displayed after the device starts up, typically used for quick access to frequently used applications, functional modules, and personalized settings. The home screen supports user-customizable layouts and allows the addition of various widgets to enhance functionality. For example, in Android, the home screen can include weather widgets, calendar widgets, music control widgets, etc., allowing users to directly view key information or perform quick operations.
[0079] The widget icon display area is a dedicated area on the home screen for showcasing dynamic widgets. Unlike static icons (such as application icons), widgets in this area can update their content in real time, providing richer interactive options. The widget icon display area is typically located on the left, right, or top of the screen, arranged according to user settings. By limiting actions to the widget icon display area, accidental activation of other functional modules can be effectively avoided, improving interaction precision and user experience.
[0080] Touch swipe events can be a series of action events triggered by a user's continuous movement of their finger or stylus on the screen, typically used for switching pages, scrolling content, or dragging objects. However, in some scenarios, such as when a user wants to trigger only the function corresponding to a certain widget without initiating a swipe operation, the system triggers a swipe operation based on the touch area of the widget icon, which may interfere with the user's actual intention.
[0081] The target model can be an interaction model generated based on the semantic analysis of the current interface and user behavior prediction. Its role is to determine whether the current touch behavior should be responded to or blocked. When a user swipes on the widget icon display area, the system can use the target model to evaluate whether the swipe triggered by the touch on the widget icon display area conforms to the expected behavior in the current context. If it does not conform, the system automatically blocks the swipe event triggered by the touch on the widget icon display area to prevent accidental operations. For example, when a user attempts to input text using a search engine widget and performs a swipe, the system can use the target model to identify the swipe triggered by the touch on the widget icon display area as an accidental touch and prevent page switching or content scrolling.
[0082] Step S302: Configure the widget icon display area to recognize the trajectory input information of the touch object.
[0083] Trajectory input information can be path data left by a user during continuous touch operations on the screen, and can be used to recognize gestures, draw graphics, input text, or execute specific commands. In this embodiment, the system intelligently configures the widget icon display area through a target model, enabling the widget icon display area to recognize and respond to the user's trajectory input. For example, a user can input text on the widget icon display area.
[0084] During implementation, the above steps are closely interconnected: First, by limiting the operation to the widget icon display area on the main screen, the target location of the user's operation is clearly defined; second, the target model is used to shield swipe events triggered by the touch of the widget icon display area, preventing accidental operations from affecting the user experience; finally, by recognizing trajectory input information, the interaction forms between the user and the widget are further enriched. Limiting the operation to the widget icon display area on the main screen, shielding swipe events triggered by the touch of the widget icon display area using the target model, and recognizing trajectory input information form a complete touch operation optimization mechanism. This mechanism can improve the intelligence level and user-friendliness of the interaction while ensuring operational accuracy.
[0085] In this embodiment, by introducing a target model into the widget icon display area of the main screen interface, the sliding events triggered by the touch of the widget icon display area are shielded and the trajectory input information is recognized, thereby significantly improving the accuracy of user operation and the smoothness of the interactive experience.
[0086] In some embodiments, step S302 above, "configuring the widget icon display area to recognize the trajectory input information of the touch body", can be achieved through the following steps:
[0087] Step 3021: Determine that the widget image display area includes text input functionality;
[0088] Widget icon display areas can be small interactive modules on a user interface used to showcase specific functions, such as function entry points presented as icons on the main interface of an application or in a floating panel. Text input functionality can be achieved by enabling the widget icon display area to convert handwriting into editable text content via stylus or touch input. Text input functionality typically relies on character recognition algorithms to instantly generate standard text after the user finishes writing.
[0089] In this application, the system uses a scene-aware mechanism to determine whether the current widget icon display area has text input functionality. When it detects that the user is about to write in the widget icon display area, the system can activate the relevant model and preload resources, thereby enabling an interactive experience where the user can write as soon as they pick up a pen. This approach avoids the redundant operation of manually clicking to enter the input mode in traditional solutions, thus improving user efficiency and user experience.
[0090] Step 3022: Call the trajectory recognition model corresponding to the text input function so that the widget icon display area can recognize the trajectory input information of the touch object.
[0091] Here, the trajectory recognition model is an artificial intelligence-based handwriting recognition algorithm that can convert the trajectory data of a user's pen into digital text in real time. The trajectory recognition model typically includes multi-dimensional input parameters such as stroke order, pressure changes, and speed curves. After training, it can accurately recognize text content under different font styles and writing habits.
[0092] During implementation, once the system confirms that the current widget icon display area has text input functionality, it will automatically call the trajectory recognition model that matches the widget icon display area. The process of automatically calling the trajectory recognition model that matches the widget icon display area starts when the user approaches the screen, thus ensuring that recognition and output of results can be completed the instant the pen touches the screen, thereby achieving a zero-latency handwriting input experience.
[0093] In this embodiment, the intelligent recognition capability of the widget icon display area is further enhanced. Especially in text input scenarios, the system can not only accurately recognize handwriting, but also dynamically adapt model parameters according to the context, improving recognition accuracy and response speed. This allows users to quickly start writing in any widget icon display area that supports text input, greatly optimizing the user experience of the stylus.
[0094] In some embodiments, the target display content area includes at least one of the following: a text editing area, an image display area, and a comment interaction area. This application embodiment uses a method for matching and loading different models for different areas, such as... Figure 3B As shown, this can be achieved through the following steps:
[0095] Step S311: In response to the target display content area being a text editing area, load the third target model and display the third interactive interface to achieve intelligent text editing in the text editing area;
[0096] Here, the text editing area can be the interface part used to input or modify text content, such as a chat window or a document editing box.
[0097] During implementation, when the system detects that the operator is about to operate in the text editing area, it can load a third target model specifically for text processing and display the corresponding third interactive interface. The system achieves intelligent text editing function by loading and displaying these contents.
[0098] The third target model can be a Natural Language Processing (NLP) model. The third target model has functions such as character recognition, automatic error correction, grammar checking, and style transfer. The third target model can convert handwritten handwriting into standard text in real time, and it can also perform intelligent optimization.
[0099] The third interactive interface typically includes auxiliary elements such as cursor positioning, character prediction, and suggested word pop-ups to complete text input and editing tasks more efficiently.
[0100] This step combines the pen trajectory with the reasoning capabilities of the AI model to achieve an intelligent writing experience that allows users to write as soon as they pick up the pen, without the need for additional clicks or switching tools, significantly reducing the learning cost and operational complexity for users.
[0101] Step S312: In response to the target display content area being an image display area, load the fourth target model and display the fourth interactive interface to achieve intelligent image editing of the image display area;
[0102] Here, the image display area can be a part of the interface used to view images, such as an album view or an image preview area.
[0103] During implementation, when the system detects that the operator is about to operate in the image display area, it can load a fourth target model specifically for image processing, and the system can display a fourth interactive interface corresponding to the model, so that the system can realize intelligent image editing functions.
[0104] The fourth target model can be a computer vision model, which has functions such as image enhancement, filter application, object recognition, and cropping suggestions. It can intelligently adjust parameters such as composition, color, and contrast of the image based on the user's gestures or pen strokes.
[0105] The fourth interactive interface typically includes graphical controls such as brush tools, erasers, filter options, and layer management, making it convenient for users to perform detailed image editing operations.
[0106] This step provides intelligent editing support in the image display area, allowing users to directly use a stylus to perform image retouching, annotation, and marking operations, thus improving the flexibility and convenience of image processing.
[0107] Step S313: In response to the target display content area being a comment interaction area, load the fifth target model and display the fifth interactive interface to perform at least one of the following intelligent functions on the comment information input by the operator through the fifth interactive interface: comment polishing, comment error correction, and comment expansion.
[0108] Here, the comment interaction area can be a part of the interface used for posting comments, messages, likes, etc. Comment interaction areas are commonly found on social media platforms, video playback pages, etc.
[0109] During implementation, when a user is detected operating in the comment interaction area, the system can load a fifth target model specifically for comment processing and display the corresponding fifth interactive interface to achieve intelligent comment assistance functions.
[0110] The fifth objective model can be a Natural Language Processing (NLG) model. This model possesses functions such as text summarization, grammar correction, sentiment analysis, and content expansion. It can intelligently optimize and supplement user-input comments. For example, the fifth objective model can analyze text, establish the context of the current conversation, and quickly generate commentary text.
[0111] The fifth interactive interface typically includes interactive components such as a comment suggestion bar, error message box, and content expansion button. These interactive components help users improve their comment content and enhance the quality of their expression.
[0112] This step introduces an intelligent assistance mechanism into the comment interaction area, enabling users not only to quickly write comments but also to receive advanced services such as grammar correction, tone optimization, and content expansion, thereby enhancing the expressive effect and interactive value of comments.
[0113] In this embodiment, by loading the corresponding target model and displaying the interactive interface according to the different types of the target display content area, the scenario-based adaptation of the operation body function can be achieved, thereby improving the operation efficiency and thus improving the problem of long triggering paths and high learning costs of operation body functions in traditional tablet devices.
[0114] In some embodiments, the target display content area is the image display area;
[0115] In step S210 above, "in response to the operator approaching but not touching the target display content area, preloading at least part of the target model data of the target interactive interface corresponding to the target display content area" can be achieved through the following process:
[0116] In response to the operator approaching but not touching the target display content area, the image editing toolchain is preloaded;
[0117] Here, the toolchain serves as the front-end carrier for user interaction, encapsulating the AI model into operable tools and defining its execution logic. For example, preloading the toolchain means preloading the basic functional modules required for image editing (such as cropping, filters, color correction, etc.), but not initializing the complete model yet.
[0118] The preloading mechanism for the image editing target model can be dynamically determined based on the content type of the current interface. For example, when entering the image browsing interface, the system identifies a certain area in the image browsing interface as the image display area based on interface semantic analysis, and preloads the toolchain related to image editing into the background for execution.
[0119] During implementation, frequently used tools (such as user history records) can be preloaded according to priority to reduce memory usage.
[0120] In step S220 above, "in response to the operator touching the target display content area, loading all the data of the target model to display the target interactive interface" can be achieved through the following steps:
[0121] In response to the touch of the target display content area by the operating body, the image editing model corresponding to the toolchain and the fourth interactive interface are loaded, and the fourth interactive interface displays the image editing controls provided by the image editing model.
[0122] Here, touch refers to the action of an object contacting the screen and generating an input signal; the image editing model is an algorithm or AI model used to support specific image editing tasks, such as image recognition models, style transfer models, etc.; the target interactive interface is a temporary pop-up or overlaid graphical interface used to display image editing-related control elements, such as filter options, text boxes, brush style selectors, etc.; graphical user interface elements can be UI elements that users can directly manipulate, such as sliders, drop-down menus, icon buttons, etc., used to adjust image properties or launch specific editing functions.
[0123] Loading a complete image editing model provides computationally intensive functions such as portrait beautification and background replacement. The fourth interactive interface can display editing controls (such as sliders, buttons, and menus), aligned with the pre-loaded toolchain.
[0124] During implementation, when the user actually touches the image display area, the system immediately loads the image editing algorithm model that matches the pre-loaded target model and simultaneously opens the target interactive interface. Users can then directly interact using the graphical user interface elements. This workflow achieves a seamless transition from intent perception to function execution. The system design significantly reduces the user's learning curve while enhancing the intuitiveness and convenience of human-computer interaction.
[0125] In this embodiment, when the system detects that a stylus is near the image display area, it automatically preloads the toolchain of the image editing model. When the stylus interacts with the screen, it loads the complete image editing model and the fourth interactive interface. These technical means can significantly optimize the response speed and interaction efficiency during stylus operation, ultimately achieving a more natural and smooth operating experience, further improving user satisfaction and ease of use when editing images on a tablet device.
[0126] This application provides an interactive system, including a stylus and a touch-screen electronic device, wherein...
[0127] The touch electronic device is used to obtain the touch operation of the stylus on the display interface;
[0128] Here, the touch-screen electronic device can be a mobile terminal with touch display capabilities, such as a tablet computer or a smartphone.
[0129] Touch operations include, but are not limited to, tapping, swiping, pressing, and hovering. These actions are recognized by the touchscreen sensor and converted into input signals. The system analyzes the location, trajectory, and timing of the touch operations to determine the user's current intention and interaction area.
[0130] The touch-sensitive electronic device is further configured to, in response to the touch operation being located in a first display content area, display a first interactive interface based on a first identification tag; and in response to the touch operation being located in a second display content area, display a second interactive interface based on a second identification tag;
[0131] The first identification tag and the second identification tag are different. The first identification tag and the second identification tag are based on the identification and marking of the currently displayed content of the display interface by the touch electronic device system. The first interactive interface is different from the second interactive interface.
[0132] Here, the second identification tag is similar to the first identification tag; it is structured identification information automatically generated by the electronic device system based on the currently displayed content. The second identification tag corresponds to the second display content area. Because the first and second identification tags are different, the electronic device system can load completely different interactive interfaces based on the differences between these two types of identification tags to meet the interactive needs of different display content areas.
[0133] In this embodiment, the touch-screen electronic device system generates different identification tags in different display content areas and dynamically loads the corresponding interactive interface based on these identification tags, achieving a highly scenario-appropriate personalized interactive effect. This not only shortens the user's operation path using the stylus but also improves the intelligence and naturalness of the interaction, effectively solving the problem of function and scenario mismatch in traditional stylus interaction.
[0134] Figure 4 This application provides a schematic diagram of a three-level architecture for a stylus interaction system, as shown in the embodiments below. Figure 4 As shown, the schematic diagram includes:
[0135] Level 1: Scene Awareness 41;
[0136] During implementation, based on the visual language model, the system uses dynamic interface semantic segmentation to divide the screen into various functional blocks (such as "text editing area", "image display area" and "comment interaction area"), and generates structured labels for each block.
[0137] For example, once a user opens any application, the system can automatically capture the screen data stream, then call the interface analysis engine to analyze and process it, dividing the interface into several functional blocks and generating a series of structured label data.
[0138] Based on scene awareness, compared to traditional solutions that rely on static user interface (UI) control recognition (such as buttons and text boxes), this solution integrates visual and textual semantics to recognize standard and non-standard control areas (such as temporary annotation areas drawn by users or custom views of third-party applications), making it applicable to a wider range of scenarios.
[0139] Level 2: Intentional Prediction 42;
[0140] During the pen input process, the system matches the scene type of the current pre-pen area based on real-time coordinates and preloads the corresponding functional modules.
[0141] During implementation, users click on a specific area of the application using a stylus. The process is divided into approaching the screen and touching the screen. During the approach phase, the system determines the current area's attributes based on coordinate data from the pen movement and label data from scene awareness, and preloads relevant model resources. These model resources can be divided into two parts: AI model resources provided by the system itself when the stylus is lifted (such as handwriting-to-text conversion, image erasure, and text enhancement); and external AI model resources invoked by the system depending on the area being processed.
[0142] For example, during the pen stroke, if the pen is 1 to 3 centimeters away from the screen, the screen can receive hover events transmitted by the pen. These events contain the pen's position coordinates relative to the screen. Based on these position coordinates, the attributes of the area where the pen will be placed can be predicted by combining the functional areas defined by scene awareness.
[0143] Here, the model preloading solutions that can be adopted at present include, but are not limited to:
[0144] 1. After system initialization, a complete inference path is executed once through a warm-up process to warm up the relevant parameters and prepare for subsequent model loading.
[0145] 2. Memory mapping (mmap) can map model files such as .tflite and .pth to the process address space, avoiding disk read and write (IO) and realizing model preloading;
[0146] 3. System service-based caching solution: By delegating the model loading logic to an independent, persistent Service process, and using AIDL to enable cross-model and cross-process calls, model preloading is achieved.
[0147] Compared to the traditional response after the stroke, this method predicts the intention during the stroke stage, preloads resources that may be needed for the next operation, speeds up the model loading process, and quickly responds to subsequent stroke operations.
[0148] Level 3: Dynamic execution 43;
[0149] Based on context-aware dynamic mode switching, the system automatically switches the interaction mode according to the block type the moment the pen touches the paper. For example, in the text editing area, the handwriting is converted into standard text in real time when the cursor is inserted; in the image display area, the pen strokes are automatically associated with the functions provided by the relevant artificial intelligence model for image processing; and in the comment interaction area, extended comments and suggestions are automatically generated.
[0150] For example, when the stylus touches the screen, the dynamic execution engine loads the relevant model, performs inference and outputs results, and dynamically displays the current model output (such as image enhancement, search, and generating enhanced text based on context).
[0151] In this way, it breaks through the traditional single input mode of "pen touch trigger - response" and realizes the mode of "pen touch prediction - pen touch use" that allows for immediate use and multiple uses of a single pen, greatly shortening the interaction path.
[0152] The following are examples illustrating three different types of pen stroke scenarios:
[0153] I. Implementing intelligent annotation of images and text in social media applications;
[0154] When users browse social media application interfaces, most videos and images are not editable. This solution intelligently analyzes the current page, automatically identifying image and text areas, marking them, and recording their type and coordinates. As the user picks up the stylus and touches the screen (pen input), the system predicts the user's touch point based on the stylus's electrical signal. Combining this with the previously marked information, it pre-loads AI drawing and editing tools. The relevant AI models are then hot-launched when the user puts pen to the screen, allowing for rapid model loading and operations such as image enhancement, erasing, and background removal; or text analysis to establish the context of the current conversation and quickly generate commentary text.
[0155] 2. Annotate text in non-editable areas of Portable Document Format (PDF) documents;
[0156] When users work with scanned PDF documents, the main body of the document is primarily composed of images, making it impossible to directly select text or perform related text operations. Based on this system's intelligent scene awareness function, it can quickly highlight areas for main text, charts, and blank annotations in the scanned document and pre-process the text information from the images using OCR. As the user writes, the system combines the previously obtained text information with the corresponding language model to provide annotation suggestions for the document.
[0157] 3. The Home screen allows for easy drawing;
[0158] The main screen contains UI elements such as application launch icons and application widgets. Widgets, as customizable main screen elements, offer preview and simple interaction capabilities. Users cannot write on the main screen using a stylus; all writing actions are consumed as swipe events. Based on the system's intelligent sensing capabilities, the main screen can be pre-divided into swipe function areas, widget interaction areas, etc., and different operation effects can be assigned to users depending on the division of the function area. For example, in the widget interaction area, swipe events are no longer responded to when the user uses the stylus; the user can trigger handwriting interaction without clicking or other operations, and a language model conversation context can be established based on the written content to perform text processing such as expansion and polishing.
[0159] Based on the foregoing embodiments, this application provides a processing device, which includes various modules, each module including sub-modules, and each sub-module including units. It can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0160] Figure 5 This is a schematic diagram of the composition structure of the processing device provided in the embodiments of this application, such as... Figure 5 As shown, the device 500 includes:
[0161] The first acquisition module 510 acquires the touch operation of the operating body on the display interface;
[0162] The first display module 520 is configured to display a first interactive interface based on a first identification tag in response to the touch operation being located in the first display content area.
[0163] The second display module 530 is used to display a second interactive interface based on a second identification tag in response to the touch operation being located in the second display content area; wherein the first identification tag and the second identification tag are different, and the first identification tag and the second identification tag are based on the electronic device system's identification and marking of the currently displayed content of the display interface, and the first interactive interface is different from the second interactive interface.
[0164] In some embodiments, the above processing device further includes a second acquisition module and a division module, wherein the second acquisition module is used to acquire the screen data stream of the display interface; the division module is used to divide the screen data stream into regions with identification tags based on a preset model; wherein different identification tags indicate that the corresponding regions implement different interactive functions, and the identification tags include the first identification tag and the second identification tag.
[0165] In some embodiments, the first interactive interface is adapted to the currently executable task of the first display content area based on a first target model, and the second interactive interface is adapted to the currently executable task of the second display content area based on a second target model, wherein the first target model and the second target model are different.
[0166] In some embodiments, the processing device further includes a preloading module and a first loading module, wherein the preloading module is configured to preload at least a portion of the target model data of the target interactive interface corresponding to the target display content area in response to the operator approaching but not touching the target display content area; the first loading module is configured to load all the data of the target model to display the target interactive interface in response to the operator touching the target display content area, wherein the target display content area includes the first display content area and the second display content area, and the target interactive interface includes the first interactive interface and the second interactive interface.
[0167] In some embodiments, the preloading includes loading the static resources of the target model, including at least one of the following: model metadata, model weights, and dependency libraries; the loading includes loading the dynamic objects of the target model, including at least one of the following: model architecture, model instance, and runtime context.
[0168] In some embodiments, the display interface is the main screen interface, the target display content area is a widget icon display area, and the processing device further includes a shielding module and a configuration module. The shielding module is used to shield touch swipe events triggered by the operator touching the widget icon display area. The configuration module is used to configure the widget icon display area to recognize the trajectory input information of the operator's touch.
[0169] In some embodiments, the configuration module includes a determining submodule and a calling submodule, wherein the determining submodule is used to determine that the widget image display area includes a text input function; and the calling submodule is used to call the trajectory recognition model corresponding to the text input function so that the widget icon display area can recognize the trajectory input information of the touch object.
[0170] In some embodiments, the target display content area includes at least one of the following: a text editing area, an image display area, and a comment interaction area. The processing device includes a second loading module, a third loading module, and a fourth loading module. The second loading module is configured to load a third target model and display a third interactive interface in response to the target display content area being a text editing area, to enable intelligent text editing in the text editing area. The third loading module is configured to load a fourth target model and display a fourth interactive interface in response to the target display content area being an image display area, to enable intelligent image editing in the image display area. The fourth loading module is configured to load a fifth target model and display a fifth interactive interface in response to the target display content area being a comment interaction area, to perform at least one of the following intelligent functions on the comment information input by the operator through the fifth interactive interface: comment polishing, comment error correction, and comment expansion.
[0171] In some embodiments, the target display content area is the image display area; the preloading module is used to preload the image editing toolchain in response to the operator approaching but not touching the target display content area; the first loading module is used to load the image editing model corresponding to the toolchain and the fourth interactive interface in response to the operator touching the target display content area, wherein the fourth interactive interface displays the image editing controls provided by the image editing model.
[0172] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0173] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause electronic devices (such as mobile phones, tablets, laptops, desktop computers, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0174] Correspondingly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the processing method provided in the above embodiments.
[0175] Correspondingly, embodiments of this application provide an electronic device, Figure 6 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, the hardware entity of the device 600 includes a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the program, it implements the steps in the processing method provided in the above embodiments.
[0176] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data and video communication data) of the processor 602 and various modules in the electronic device 600, and can be implemented by flash memory or random access memory (RAM).
[0177] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0178] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0179] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0180] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0181] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0182] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0183] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0184] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a mobile phone, tablet computer, laptop computer, desktop computer, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.
[0185] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0186] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0187] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0188] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A processing method, the method comprising: Obtain the touch operation of the operator on the display interface; In response to the touch operation being located in the first display content area, a first interactive interface is displayed based on the first identification tag; In response to the touch operation being located in the second display content area, a second interactive interface is displayed based on the second identification tag; The first identification tag and the second identification tag are different. The first identification tag and the second identification tag are based on the electronic device system's identification and marking of the currently displayed content of the display interface. The first interactive interface is different from the second interactive interface.
2. The method as described in claim 1, wherein the first identification tag and the second identification tag are based on the identification and marking of the currently displayed content of the display interface by the electronic device system, comprising: Obtain the screen data stream of the display interface; The screen data stream is divided into regions based on a preset model to obtain multiple regions with identification tags; Different identification tags represent different interactive functions for corresponding areas, and the identification tags include the first identification tag and the second identification tag.
3. The method as described in claim 1, wherein the first interactive interface adapts to the currently executable task of the first display content area based on the first target model, and the second interactive interface adapts to the currently executable task of the second display content area based on the second target model, wherein the first target model and the second target model are different.
4. The method of claim 3, further comprising: In response to the operator approaching but not touching the target display content area, at least part of the target model data of the target interactive interface corresponding to the target display content area is preloaded; In response to the operator touching the target display content area, all data of the target model is loaded to display the target interactive interface, wherein the target display content area includes the first display content area and the second display content area, and the target interactive interface includes the first interactive interface and the second interactive interface.
5. The method of claim 4, wherein the preloading includes loading the static resources of the target model, including at least one of the following: model metadata, model weights, and dependency libraries; and the loading includes loading the dynamic objects of the target model, including at least one of the following: model architecture, model instance, and runtime context.
6. The method as described in claim 4, wherein the display interface is the main screen interface, the target display content area is the widget icon display area, and the method further includes: The touch swipe event triggered by the operator touching the widget icon display area is blocked; The widget icon display area is configured to recognize the trajectory input information of the touch object.
7. The method of claim 6, wherein configuring the widget icon display area to recognize the trajectory input information of the touch object includes: The image display area of the widget is determined to include text input functionality; The trajectory recognition model corresponding to the text input function is invoked so that the widget icon display area can recognize the trajectory input information of the touch object.
8. The method of claim 4, wherein the target display content area includes at least one of the following: a text editing area, an image display area, and a comment interaction area, and the method further includes: In response to the target display content area being a text editing area, a third target model is loaded and a third interactive interface is displayed to enable intelligent text editing of the text editing area; In response to the target display content area being an image display area, a fourth target model is loaded and a fourth interactive interface is displayed to enable intelligent image editing of the image display area; In response to the target display content area being a comment interaction area, a fifth target model is loaded and a fifth interactive interface is displayed to perform at least one of the following intelligent functions on the comment information input by the operator through the fifth interactive interface: comment polishing, comment error correction, and comment expansion.
9. The method as described in claim 8, wherein the target display content area is the image display area; In response to the operator approaching but not touching the target display content area, at least a portion of the target model data of the target interactive interface corresponding to the target display content area is preloaded, including: In response to the operator approaching but not touching the target display content area, the image editing toolchain is preloaded; Correspondingly, the step of loading all data of the target model to display the target interactive interface in response to the touch of the operating body on the target display content area includes: In response to the touch of the target display content area by the operating body, the image editing model corresponding to the toolchain and the fourth interactive interface are loaded, and the fourth interactive interface displays the image editing controls provided by the image editing model.
10. An interactive system comprising a stylus and a touch-sensitive electronic device, wherein, The touch electronic device is used to obtain the touch operation of the stylus on the display interface; The touch-sensitive electronic device is further configured to, in response to the touch operation being located in a first display content area, display a first interactive interface based on a first identification tag; and in response to the touch operation being located in a second display content area, display a second interactive interface based on a second identification tag; The first identification tag and the second identification tag are different. The first identification tag and the second identification tag are based on the identification and marking of the currently displayed content of the display interface by the touch electronic device system. The first interactive interface is different from the second interactive interface.