An information processing method, apparatus, device, and storage medium

By acquiring interface information and interaction focus from the graphical user interface, and generating response information based on spatial distribution information, the problem of inaccurate understanding of vague instructions is solved, resulting in more accurate responses and a better user experience.

CN122111270APending Publication Date: 2026-05-29BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-02-12
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, ambiguous instructions in graphical user interfaces, such as "change this to blue", cannot be accurately understood by the model, resulting in inaccurate response information and poor user experience.

Method used

By acquiring interface information from the graphical user interface, including interactive objects and interactive focus, determining the context of the request information based on spatial distribution information, and generating response information using a natural language processing model, the request intent of ambiguous commands can be accurately understood.

Benefits of technology

This improved the accuracy of response information and enhanced the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122111270A_ABST
    Figure CN122111270A_ABST
Patent Text Reader

Abstract

Provided in one scenario are an information processing method, device, equipment, and storage medium. The method comprises: in response to a preset event trigger, obtaining interface information of a graphical interaction interface, wherein the interface information comprises a first interaction object and an interaction focus, and the interaction focus represents an interaction position in the graphical interaction interface; based on spatial distribution information of the first interaction object relative to the interaction focus, obtaining context information of request information; and based on the context information and the request information, generating response information corresponding to the request information. The technical solution in one scenario accurately understands the request intention corresponding to a fuzzy instruction, solves the problem in the related art that the response information is inaccurate due to an inability to accurately understand the request intention, improves the accuracy of the response information, and improves the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In one case, computer technology is involved, and more particularly, an information processing method, apparatus, device, and storage medium are involved. Background Technology

[0002] With the development of computer technology, the capabilities of models have gradually increased. Models can not only process natural language text but also be applied to graphical user interfaces (GUIs) to handle request information, thus improving the intelligence of interactions. However, in related technologies, request information is often expressed in highly colloquial or contextual terms, such as "change this part to blue." "This part" is a vague instruction that the model cannot understand, making it difficult for the model to correctly interpret the request intent. This often results in returning generic or invalid responses, leading to a poor user experience. Summary of the Invention

[0003] In one scenario, an information processing method, apparatus, device, and storage medium are provided that can improve the accuracy of response information and enhance the user experience.

[0004] Firstly, in one scenario, an information processing method is provided, including:

[0005] In response to a preset event, the interface information of the graphical interactive interface is obtained, wherein the interface information includes a first interactive object and an interactive focus, and the interactive focus represents the interactive position in the graphical interactive interface;

[0006] Based on the spatial distribution information of the first interactive object relative to the interactive focus, the contextual information of the request information is obtained.

[0007] Based on the context information and the request information, response information corresponding to the request information is generated.

[0008] Secondly, in one scenario, an information processing apparatus is also provided, the apparatus comprising:

[0009] The first acquisition module is used to acquire interface information of the graphical interactive interface in response to a preset event trigger, wherein the interface information includes a first interactive object and an interactive focus, and the interactive focus represents the interactive position in the graphical interactive interface.

[0010] The second obtaining module is used to obtain contextual information of the request information based on the spatial distribution information of the first interactive object relative to the interactive focus;

[0011] The information generation module is used to generate response information corresponding to the request information based on the context information and the request information.

[0012] Thirdly, in one scenario, an electronic device is also provided, the electronic device comprising:

[0013] One or more processors;

[0014] Storage device for storing one or more programs.

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the information processing method as described in any embodiment of this disclosure.

[0016] Fourthly, in one scenario, a storage medium containing computer-executable instructions is also provided, which, when executed by a computer processor, are used to perform the information processing method as described in any embodiment of this disclosure.

[0017] In one scenario, an information processing method is provided. Responding to a preset event trigger, it obtains interface information from a graphical user interface (GUI). This interface information includes a first interactive object and an interactive focus, where the interactive focus represents the interactive position within the GUI. Based on the spatial distribution information of the first interactive object relative to the interactive focus, contextual information of the requested information is obtained. The interactive operation includes the request information. This allows for accurate identification of a second interactive object near the interactive focus within the GUI, generating the context of the request information. Then, based on the contextual information, the request information is understood, and corresponding response information is generated. This achieves accurate understanding of the request intent corresponding to ambiguous instructions, solving the problem of inaccurate response information caused by the inability to accurately understand the request intent in related technologies, thus improving the accuracy of the response information and enhancing the user experience. Attached Figure Description

[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0019] Figure 1 This is a diagram illustrating an application scenario of an information processing method provided under a particular condition.

[0020] Figure 2 This is a flowchart illustrating an information processing method provided in one scenario.

[0021] Figure 3 This is a flowchart illustrating another information processing method provided in one scenario.

[0022] Figure 4 This is a schematic diagram of the structure of an information processing device provided in one scenario;

[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided in one scenario. Detailed Implementation

[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0030] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0031] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0032] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0033] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0034] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0035] Figure 1 This is a diagram illustrating an application scenario of an information processing method provided under certain conditions. It should be noted that... Figure 1 This is merely an example of an embodiment under one scenario of this disclosure, intended to help those skilled in the art understand the technical content of such an embodiment, and does not imply that the embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios.

[0036] like Figure 1 As shown, the system may include a server 110, a network 120, several terminal devices 130, and a database 140. The terminal devices include at least one of mobile terminals and PCs.

[0037] Server 110 can be a physical server containing an independent host, or it can be a virtual server hosted in a host cluster. During operation, server 110 can run server-side programs for human-computer interaction applications, thus acting as the server-side component of these applications. Examples of human-computer interaction applications include media content design tools (such as animation storyboard design tools), whiteboard tools, 3D modeling software, game engines, virtual reality applications, and metaverse applications. The user interface of these applications is a graphical interface. The human-computer interaction application is combined with a first model, which acts as an intelligent assistant to enhance the application's intelligence. In one scenario, the intelligent assistant can be a computer program module or a hardware / software integrated system built based on machine learning or large model technology, possessing autonomous perception, decision-making, and execution capabilities. Optionally, the intelligent assistant obtains real-time or non-real-time data from the external environment (including data or physical environments) through APIs (Application Programming Interfaces) or data input ports. Relying on pre-trained models, reinforcement learning algorithms, or rule engines, it analyzes and infers from the perceived data to generate task plans that meet preset goals. Based on task planning, pre-set external tools are invoked to perform inference tasks, ultimately obtaining response information for the user's input request information.

[0038] exist Figure 1 In the system configuration described, server 110 may include one or more components that implement the functions performed by server 110. These components may include at least one of software components and hardware components executable by one or more processors. A user operating terminal device 130 can interact with server 110 using a client to utilize the services provided by these components. It is understood that different system configurations may exist. Figure 1 This is a schematic diagram of an application scenario for implementing an information processing method under certain conditions.

[0039] Terminal device 130 connects to server 110 via network 120. Typically, network 120 can be any type of network, and it can use any of the various available protocols (including TCP / IP (Transmission Control Protocol / Internet Protocol) and IPX (Internet Packet Exchange)) to support data communication. Network 120 can include various connection types, such as wired communication links, wireless communication links, or fiber optic cables.

[0040] Database 140 can be used to store design drafts of media content, object information of the first interactive object, object information of the second interactive object, relevance, interaction focus, and viewport information, etc. Database 140 can reside in various locations. For example, database 120 used by server 110 can be local to server 110, or it can be located away from server 110 and can communicate with server 110 via network 120 or a dedicated connection. Database 140 can be of different types. In some cases, database 140 used by server 110 can be a relational database. One or more of these databases 140 can store, update, and retrieve database 140 and data from database 140 in response to commands.

[0041] In some cases, one or more of the databases 140 may also be used by applications to store application data. The databases 140 used by applications may be of different types, such as key-value stores, object stores, or regular storage types supported by the file system.

[0042] It should be noted that, Figure 1 The system can be configured and operated in various ways to achieve information processing and devices in various situations.

[0043] Figure 2 This is a flowchart illustrating an information processing method provided in one scenario, applicable to situations where a first model is assisted in understanding requested information. The method can be executed by an information processing device, which can be implemented in the form of software and / or hardware, or optionally, by an electronic device, such as a mobile terminal, a PC, or a server.

[0044] like Figure 2 As shown, the method includes:

[0045] S210. In response to a preset event trigger, obtain interface information of the graphical interactive interface, wherein the interface information includes a first interactive object and an interactive focus, and the interactive focus represents the interactive position in the graphical interactive interface.

[0046] The triggering condition of the preset event includes any one of the following: the preset event is triggered in response to a triggering operation on a preset control in the graphical interactive interface; the preset event is triggered in response to an interactive operation on the canvas area in the graphical interactive interface.

[0047] The preset controls include interactive controls in the graphical user interface. For example, preset interactive controls include confirmation buttons or send buttons.

[0048] In one scenario, a dialog box for the intelligent assistant is displayed in a preset area of ​​the graphical user interface (GUI). Correspondingly, the request operation for the GUI can be an input operation for the dialog box. This input operation includes text input or voice input, etc. In response to the input operation for the dialog box, the request information is displayed within the dialog box. The dialog box also displays the activation state of a preset control. The activation state refers to the interactive state of the preset control. In response to a trigger operation for the preset control, a preset event is triggered.

[0049] In another scenario, a preset event is triggered in response to a preset interactive operation on the canvas area of ​​the graphical user interface. For example, the preset interactive operation may include a drag operation on a first interactive object. Alternatively, the preset interactive operation may include a zoom operation on the canvas area. This disclosure does not specifically limit the type of preset interactive operation.

[0050] Optionally, in one scenario, preset events can be triggered periodically or at set intervals to promptly capture changes in the spatial layout of the graphical user interface. One embodiment provided in this disclosure not only triggers preset events upon user request but also periodically or at set intervals during user input, thus preparing contextual information related to the request before the user's request and reducing request response latency.

[0051] Graphical user interfaces (GUIs) are the human-computer interaction interfaces of human-computer interaction applications. Examples of GUIs include those for animation storyboard design tools, online whiteboard tools, and virtual reality tools.

[0052] Interface information represents content related to the graphical user interface (GUI). For example, interface information includes the first interactive object and the interaction focus. The first interactive object is the interactive object included in the GUI. The interaction focus represents the user's position in the GUI. For example, in the GUI of an animation storyboard design tool, multiple storyboard blocks and text annotations are displayed as interactive objects. In a 2D scene, the interaction focus typically includes the cursor position or the viewport center point. The viewport is the visible area of ​​the canvas. In a 3D scene, the interaction focus typically includes the user's virtual position, the intersection of the user's line of sight and the 3D scene, or a virtual focus in front of the camera. Optionally, in a 3D scene, the viewport is the visible range determined by the view frustum.

[0053] S220. Based on the spatial distribution information of the first interactive object relative to the interactive focus, obtain the contextual information of the request information.

[0054] Spatial distribution information represents the spatial relationship between the first interactive object and the interaction focus. Optionally, spatial distribution information includes the distance and orientation of the first interactive object relative to the interaction focus. Contextual information represents the contextual information of the request information. Contextual information is used to assist in understanding the request information. The request information is the operation request described by the user in natural language. For example, if the user hovers the cursor over a certain area of ​​the canvas, the request information may include the user inputting "optimize this layout" into the intelligent assistant.

[0055] In one scenario, based on the spatial distribution information of the first interactive object relative to the interactive focus, a second interactive object and the relevance between the second interactive object and the interactive focus are obtained; the contextual information is then generated based on the second interactive object and the relevance.

[0056] The second interactive object is the first interactive object whose spatial distribution information with the interactive focus meets preset conditions. Relevance is used to quantify the degree of association between the second interactive object and the user's current interactive focus.

[0057] In one scenario, the relevance between the second interactive object and the interactive focus is determined based on multiple spatial perception dimensions. For example, these multiple spatial perception dimensions may include a distance dimension. Optionally, the multiple spatial perception dimensions may also include at least one of a zoom dimension, a visibility dimension, and an offset dimension. Camera orientation, zoom level, and viewpoint data are determined based on interface information from the graphical user interface.

[0058] Optionally, factors influencing the relevance of the distance dimension may include Euclidean distance, etc.

[0059] Optionally, factors influencing the scaling dimension's impact on relevance may include a scaling weight function. The scaling weight function dynamically adjusts the distance penalty effect based on the current viewport's scaling level. It converts a fixed distance into a dynamic, visually perceived distance. In response to a zoom-in operation, the scaling level increases, and the scaling weight function increases the penalty effect on relevance due to the increase in distance. This aligns with the intuitive feeling of focusing on nearby details from a microscopic perspective. In response to a zoom-out operation, the scaling level decreases, and the scaling weight function weakens the penalty effect on relevance due to the increase in distance. This aligns with the ability to consider more distant interactive objects from a macroscopic perspective.

[0060] Optionally, factors influencing the relevance of the visibility dimension may include visibility weights. The view frustum is determined based on the camera orientation. The occlusion area of ​​the first interactive object is determined based on the view frustum, and the visibility weight is determined based on the occlusion area. Alternatively, the pixel-level visibility weight of the first interactive object can be obtained through GPU hardware occlusion query. Alternatively, the visibility weight can be calculated through depth buffer analysis. Alternatively, the visibility weight can be calculated using bounding box intersection testing.

[0061] Optionally, the bias dimension's influence on relevance may include a bias parameter. The bias parameter is related to the distance between the center of the second interactive object and the viewport center. If the second interactive object is located close to the center of the current viewport, it can be considered that the second interactive object is likely to be the object of user attention, or that the second interactive object will influence the object of user attention.

[0062] In another scenario, a second model is used to identify the interface information of the graphical user interface, obtain the spatial distribution information of the first interactive object relative to the interactive focus, and determine the contextual information of the request information based on the spatial distribution information. The step of generating response information corresponding to the request information based on the contextual information and the request information includes: understanding the request information based on the contextual information using the second model, and generating response information corresponding to the request information.

[0063] The second model can be a model used to implement an intelligent assistant. It can be a model used to process two or more modalities of information (e.g., text, images, audio, video, or point cloud data). In one scenario, the second model can refer to a pre-trained model with hundreds of billions of parameters, capable of simultaneously receiving, understanding, and generating two or more heterogeneous modalities of information. In another scenario, the second model employs a network architecture adapted to multiple modalities of information, converting the raw data into feature vectors of different modalities. Through attention mechanisms, feature concatenation, or gating fusion algorithms, it maps the feature vectors of different modalities to a unified semantic space, achieving the association and interaction of heterogeneous modal information, and outputting single or multiple modal result data according to the task objective.

[0064] In one scenario, the interface information can be a snapshot or screenshot of the graphical user interface. The snapshot records the running state of the graphical user interface. The snapshot includes at least one of the following: visual interface information, application memory data, process state, and configuration parameters. By recognizing the snapshot, the second model can obtain the spatial relationship between the first interactive object and the interaction focus, thereby aiding in understanding the request intent and generating the corresponding response information.

[0065] S230. Based on the context information and the request information, generate response information corresponding to the request information.

[0066] The first model can be a model used to implement an intelligent assistant. For example, the first model can include a natural language processing model. Optionally, the first model can be a deep learning model that has undergone self-supervised learning on tens of billions of text data points, possesses a parameter scale of tens of billions, and is capable of understanding, generating, and processing natural language. Optionally, the first model can include a large language model.

[0067] In one scenario, the contextual information can be structured data adapted to the input data format of the first model. Optionally, at least a portion of the object information of the second interactive object is converted into structured data conforming to the input data format of the first model. This object information includes object identifier, object type, object coordinates, object size, text content, distance and relevance between the second interactive object and the interaction focus, etc. A prompt text is generated based on the request information and contextual information. The prompt text is input into the first model, which, combined with the contextual information, resolves ambiguous references in the request information, understands the request intent, and generates response information that matches the request intent.

[0068] One technical solution involves obtaining interface information of a graphical user interface (GUI) in response to a preset event. This interface information includes a first interactive object and an interactive focus, where the interactive focus represents the interactive position within the GUI. Based on the spatial distribution information of the first interactive object relative to the interactive focus, contextual information of the requested information is obtained. The interactive operation includes the request information. This allows for the accurate identification of a second interactive object near the interactive focus within the GUI, thereby generating the context of the request information. Then, a first model understands the request information based on the contextual information and generates corresponding response information. This accurately understands the request intent corresponding to ambiguous commands, solving the problem of inaccurate response information caused by the inability to accurately understand the request intent in related technologies. This improves the accuracy of the response information and enhances the user experience.

[0069] Figure 3 This is a flowchart illustrating another information processing method provided in one scenario. Based on the above embodiments, it specifically defines obtaining a second interactive object and the relevance between the second interactive object and the interactive focus based on the spatial distribution information of the first interactive object relative to the interactive focus; and generating the contextual information based on the second interactive object and the relevance.

[0070] like Figure 3 As shown, the method includes:

[0071] S310. In response to a preset event trigger, obtain interface information of the graphical interactive interface, wherein the interface information includes a first interactive object and an interactive focus, and the interactive focus represents the interactive position in the graphical interactive interface.

[0072] For example, in response to a preset event, a snapshot of the graphical user interface is obtained.

[0073] S320. Obtain the interaction focus and viewport information based on the interface information.

[0074] Viewport information is used to characterize the viewport state. Viewport information includes the zoom level, viewpoint, and occlusion information corresponding to the graphical user interface. In an infinite canvas or 3D space, users switch between a macroscopic global view and a microscopic local viewpoint through zooming and panning operations. Two primary interactive objects may have very close physical coordinates at the data level, but when the user zooms in to the maximum, one of the primary interactive objects may have already been moved off the screen (or viewport). Conversely, when the user zooms out to the minimum, two secondary interactive objects on the screen may appear visually adjacent, but their physical coordinates may be far apart. Using viewport information, the distance between two primary interactive objects can be combined with the user's viewport state to simulate the perceptual patterns of the human eye at different visual scales, dynamically adapting to the user's switching between macroscopic and microscopic viewpoints, providing a foundation for the accurate quantification of relevance.

[0075] The scaling level represents the ratio of the current view to its initial size. For example, the scaling level includes the scaling factor.

[0076] For example, the interaction focus and viewport information are determined based on a snapshot of the graphical user interface.

[0077] S330. Based on the distance between the first interactive object and the interactive focus, and the viewport information, obtain the second interactive object and the correlation between the second interactive object and the interactive focus.

[0078] The distance between the second interactive object and the interactive focus is less than or equal to the search radius. The search radius is related to the scaling level. For example, the search radius and scaling level are inversely proportional. The higher the scaling level, the smaller the search radius. The lower the scaling level, the larger the search radius.

[0079] For example, the retrieval radius is determined based on the scaling level, wherein the scaling level represents the view scaling ratio of the graphical user interface; the first interactive object is retrieved based on the interaction focus and the retrieval radius to obtain the second interactive object, wherein the distance between the second interactive object and the interaction focus is less than or equal to the retrieval radius; the relevance between the second interactive object and the interaction focus is determined based on the distance and the viewport information.

[0080] Based on a snapshot of the graphical user interface, a structure tree of the first interactive objects is constructed. Each node in the structure tree represents a first interactive object. The visible first interactive objects in the structure tree are determined based on viewport information. For example, first interactive objects located off-screen are filtered out based on viewport information. Alternatively, first interactive objects located outside the view frustum are filtered out based on viewport information.

[0081] For the first interactive object visible in the tree structure, calculate the distance between the center of the first interactive object and the interaction focus. If the distance is less than or equal to the retrieval radius, then the first interactive object is identified as the second interactive object. It should be noted that the tree structure-based retrieval method described in this case is only an example; other index structures that utilize spatial partitioning to accelerate proximity search can also be used.

[0082] For example, the relevance between the second interactive object and the interactive focus is determined based on the distance between the second interactive object and the interactive focus, the scaling level, the occlusion information of the second interactive object, and the offset parameter, wherein the offset parameter is related to the distance between the center of the second interactive object and the center of the viewport.

[0083] In one scenario, the Euclidean distance between the center of the second interactive object and the interaction focus is calculated. Optionally, to eliminate resolution differences, the current viewport size needs to be used as the normalization reference. For example, the diagonal length of the current viewport can be used as the normalization reference for the Euclidean distance.

[0084] In one scenario, for any second interactive object, the relevance between the current second interactive object and the interaction focus is determined based on the Euclidean distance, scaling weight function, visibility weight, and bias parameter. It can be understood that a higher relevance indicates a greater relevance between the second interactive object and the request intent.

[0085] The scaling weight function is an inverse proportional function of the scaling level. The larger the scaling level, the smaller the scaling weight, and the faster the Euclidean distance decays. Optionally, the scaling weight function is determined based on the scaling level and a preset coefficient. The preset coefficient adjusts the degree of influence of scaling on the Euclidean distance.

[0086] One approach is to determine the visibility weight of the second interactive object using GPU-based hardware occlusion queries. Alternatively, an intersection test can be performed based on the bounding box of the second interactive object to calculate the occlusion area ratio, and the visibility weight can be determined based on this ratio. It should be noted that the above methods for calculating visibility weight are merely examples and not limitations; ray tracing techniques can also be used to determine the visibility weight of the second interactive object, among other methods.

[0087] Specifically, the offset distance of the second interactive object is calculated based on its center coordinates and the viewport center coordinates. The offset parameter is then determined based on the offset distance and the viewport width. The closer the second interactive object is to the viewport center, the higher the offset parameter.

[0088] S340. Determine the sorting result of the second interactive object based on the relevance.

[0089] S350. Based on the sorting result and the object information of the second interactive object, generate the context information.

[0090] In one scenario, a third interactive object is obtained based on a preset threshold of the first model and the number of information units of the object information. The preset threshold represents the threshold of the number of information units processed by the first model in a single interaction. The first model is used to generate the response information based on the context information and the request information. The context information is determined based on the object information of the third interactive object.

[0091] Here, an information unit is the smallest semantic unit obtained after word segmentation of the input text of the first model. The information unit is the basic unit by which the first model understands request information and generates response information. To prevent contextual information from exceeding the preset threshold of the first model, second interaction objects need to be selected one by one according to the sorting results. For each selected second interaction object, the number of information units it occupies after converting its information into structured text is determined. The number of information units occupied by the selected second interaction objects is accumulated; when the difference between the total number of information units and the preset threshold is less than a preset value, the selection of second interaction objects stops. The selected second interaction object becomes the third interaction object.

[0092] In another scenario, the second interaction objects are sorted in descending order based on relevance, and the top n second interaction objects are selected as the third interaction objects. Here, n is a preset positive integer.

[0093] S360. Based on the context information and the request information, generate response information corresponding to the request information.

[0094] For example, the object information of the third interactive object is transformed into structured data conforming to the input data format of the first model. This object information includes object identifier, object type, object coordinates, object size, text content, distance and relevance between the third interactive object and the interaction focus, etc. Prompt text is generated based on the request information and contextual information. Inputting the prompt text into the first model provides rich decision-making support. By combining the first model with contextual information, ambiguous references in the request information are resolved to understand the request intent and generate response information that matches that intent.

[0095] In one scenario, a user enters "Make the character's expression in this storyboard more surprised" in the intelligent assistant dialog box of the animation storyboard design tool and clicks the send button. In response to the click of the send button, a snapshot of the current graphical interface of the animation storyboard design tool is obtained. The cursor position and viewport information are obtained based on the snapshot. The search radius is determined based on the zoom level included in the viewport information. The structure tree of the first interactive object can be a quadtree. The first interactive object in the quadtree is traversed, and the distance between the first interactive object and the cursor position is calculated. If the distance is less than or equal to the search radius, the first interactive object is used as the second interactive object. Assume three second interactive objects are retrieved: storyboard tile A (position coordinates (518, 479)), storyboard tile B (position coordinates (580, 482)), and text annotation C (content is xxx, position coordinates (522, 550)). It is understood that more second interactive objects may be retrieved in actual applications; the three second interactive objects selected in this disclosure are merely examples and not limited. Based on the distance and viewport information between the second interactive object and the interactive focus, the relevance between the second interactive object and the interactive focus is determined. Assume the result of the relevance ranking is [Storyboard Patch A, Text Annotation C, Storyboard Patch B]. The third interactive object is determined based on the ranking result. The object information of the third interactive object is converted into contextual information in XML (eXtensible Markup Language) format, conforming to the input data format of the first model. The XML-formatted contextual information is embedded into a predefined template to form a complete spatially aware cue word. The user's input request information is combined with this spatially aware cue word to obtain the cue text for the first model. The cue text is sent to the first model for inference. The first model, through the cue text, clarifies that "this storyboard" refers to storyboard patch A, and generates an adjustment instruction that meets the requirements of the request information as the response information.

[0096] One technical solution involves determining the retrieval radius based on the scaling level, enabling dynamic adjustment of the object retrieval range according to the scaling level. The system tree of the first interactive object is traversed based on the interaction focus and the retrieval radius to obtain the second interactive object, achieving efficient spatial data retrieval. Furthermore, based on the distance between the second interactive object and the interaction focus, and viewport information, the relevance between the second interactive object and the interaction focus is obtained. Contextual information is generated based on the second interactive object and the relevance, comprehensively considering physical distance, scaling information, and occlusion information to calculate the relevance between the second interactive object and the interaction focus, achieving precise quantification of the relevance. By understanding the request information based on contextual information through the first model, a deep understanding of ambiguous instructions in the request information can be achieved, generating highly scene-relevant and accurate response information.

[0097] Figure 4This is a schematic diagram of the structure of an information processing device provided in one scenario. The device can be implemented in the form of software and / or hardware, and optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0098] like Figure 4 As shown, the device includes: a first acquisition module 410, a second acquisition module 420, and an information generation module 430.

[0099] The first obtaining module 410 is used to obtain interface information of the graphical interactive interface in response to a preset event trigger, wherein the interface information includes a first interactive object and an interactive focus, and the interactive focus represents the interactive position in the graphical interactive interface.

[0100] The second obtaining module 420 is used to obtain contextual information of the request information based on the spatial distribution information of the first interactive object relative to the interactive focus;

[0101] The information generation module 430 is used to generate response information corresponding to the request information based on the context information and the request information.

[0102] Optionally, the triggering condition for the preset event includes any one of the following:

[0103] In response to a triggering operation on a preset control in the graphical user interface, the preset event is triggered;

[0104] The preset event is triggered in response to an interactive operation on the canvas area of ​​the graphical user interface.

[0105] Optionally, the second acquisition module 420 is specifically used for:

[0106] Based on the spatial distribution information of the first interactive object relative to the interactive focus, a second interactive object and the correlation between the second interactive object and the interactive focus are obtained;

[0107] The contextual information is generated based on the second interaction object and the relevance.

[0108] Optionally, obtaining the second interactive object and the relevance between the second interactive object and the interactive focus based on the spatial distribution information of the first interactive object relative to the interactive focus includes:

[0109] The interaction focus and viewport information are obtained based on the interface information;

[0110] Based on the distance between the first interactive object and the interactive focus, and the viewport information, the second interactive object and the correlation between the second interactive object and the interactive focus are obtained.

[0111] Optionally, obtaining the second interactive object and its relevance to the interactive focus based on the distance between the first interactive object and the interactive focus, and the viewport information, includes:

[0112] The retrieval radius is determined based on the scaling level, wherein the scaling level represents the view scaling ratio of the graphical user interface.

[0113] The first interactive object is retrieved based on the interaction focus and the retrieval radius to obtain the second interactive object, wherein the distance between the second interactive object and the interaction focus is less than or equal to the retrieval radius;

[0114] Based on the distance and the viewport information, the relevance between the second interactive object and the interactive focus is determined.

[0115] Optionally, generating the contextual information based on the second interaction object and the relevance includes:

[0116] The ranking result of the second interactive object is determined based on the relevance;

[0117] The contextual information is generated based on the sorting result and the object information of the second interactive object.

[0118] Optionally, generating the contextual information based on the sorting result and the object information of the second interactive object includes:

[0119] Based on the preset threshold of the first model and the number of information units of the object information, a third interactive object is obtained, wherein the preset threshold represents the threshold of the number of information units processed by the first model in a single interaction, and the first model is used to generate the response information based on the context information and the request information;

[0120] The context information is determined based on the object information of the third interactive object.

[0121] Optionally, obtaining contextual information about the request based on the spatial distribution information of the first interactive object relative to the interactive focus includes:

[0122] The second model is used to identify the interface information of the graphical user interface, obtain the spatial distribution information of the first interactive object relative to the interactive focus, and determine the contextual information of the request information based on the spatial distribution information.

[0123] The step of generating response information corresponding to the request information based on the context information and the request information includes:

[0124] The second model understands the request information based on the contextual information and generates response information corresponding to the request information.

[0125] In one scenario, the provided information processing apparatus can execute the information processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0126] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0127] Figure 5 This is a schematic diagram of the structure of an electronic device provided in one scenario. See below for reference. Figure 5 It shows an electronic device suitable for implementing a particular situation (e.g.) Figure 5 The diagram shows the structure of the terminal device or server 500. In one case, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0128] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.

[0129] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0130] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 508, or installed from ROM 502. When the computer program is executed by processing device 501, it performs the functions defined in one scenario of the method.

[0131] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0132] In one scenario, the electronic device provided is based on the same inventive concept as the information processing method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0133] In one scenario, a computer storage medium is provided that stores a computer program, which, when executed by a processor, implements the information processing method provided in the above embodiments.

[0134] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0135] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0136] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0137] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0138] In response to a preset event, the interface information of the graphical interactive interface is obtained, wherein the interface information includes a first interactive object and an interactive focus, and the interactive focus represents the interactive position in the graphical interactive interface;

[0139] Based on the spatial distribution information of the first interactive object relative to the interactive focus, the contextual information of the request information is obtained.

[0140] Based on the context information and the request information, response information corresponding to the request information is generated.

[0141] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0143] The units described in a particular scenario can be implemented in software or hardware. In some cases, the name of a unit does not constitute a limitation on the unit itself.

[0144] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0146] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0147] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0148] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An information processing method, comprising: In response to a preset event, the interface information of the graphical interactive interface is obtained, wherein the interface information includes a first interactive object and an interactive focus, and the interactive focus represents the interactive position in the graphical interactive interface; Based on the spatial distribution information of the first interactive object relative to the interactive focus, the contextual information of the request information is obtained. Based on the context information and the request information, response information corresponding to the request information is generated.

2. The method according to claim 1, wherein the triggering condition of the preset event includes any one of the following: In response to a triggering operation on a preset control in the graphical user interface, the preset event is triggered; The preset event is triggered in response to an interactive operation on the canvas area of ​​the graphical user interface.

3. The method according to claim 1, wherein obtaining contextual information of the request information based on the spatial distribution information of the first interactive object relative to the interactive focus includes: Based on the spatial distribution information of the first interactive object relative to the interactive focus, a second interactive object and the correlation between the second interactive object and the interactive focus are obtained; The contextual information is generated based on the second interaction object and the relevance.

4. The method according to claim 3, wherein obtaining the second interactive object and the correlation between the second interactive object and the interactive focus based on the spatial distribution information of the first interactive object relative to the interactive focus comprises: The interaction focus and viewport information are obtained based on the interface information; Based on the distance between the first interactive object and the interactive focus, and the viewport information, the second interactive object and the correlation between the second interactive object and the interactive focus are obtained.

5. The method according to claim 4, wherein obtaining the second interactive object and the relevance between the second interactive object and the interactive focus based on the distance between the first interactive object and the interactive focus and the viewport information includes: The retrieval radius is determined based on the scaling level, wherein the scaling level represents the view scaling ratio of the graphical user interface. The first interactive object is retrieved based on the interaction focus and the retrieval radius to obtain the second interactive object, wherein the distance between the second interactive object and the interaction focus is less than or equal to the retrieval radius; Based on the distance and the viewport information, the relevance between the second interactive object and the interactive focus is determined.

6. The method according to claim 3, wherein generating the contextual information based on the second interaction object and the relevance includes: The ranking result of the second interactive object is determined based on the relevance; The contextual information is generated based on the sorting result and the object information of the second interactive object.

7. The method according to claim 6, wherein generating the contextual information based on the sorting result and the object information of the second interactive object includes: Based on the preset threshold of the first model and the number of information units of the object information, a third interactive object is obtained, wherein the preset threshold represents the threshold of the number of information units processed by the first model in a single interaction, and the first model is used to generate the response information based on the context information and the request information; The context information is determined based on the object information of the third interactive object.

8. The method according to claim 1, wherein obtaining contextual information of the request information based on the spatial distribution information of the first interactive object relative to the interactive focus includes: The second model is used to identify the interface information of the graphical user interface, obtain the spatial distribution information of the first interactive object relative to the interactive focus, and determine the contextual information of the request information based on the spatial distribution information. The step of generating response information corresponding to the request information based on the context information and the request information includes: The second model understands the request information based on the contextual information and generates response information corresponding to the request information.

9. An information processing apparatus, comprising: The first acquisition module is used to acquire interface information of the graphical interactive interface in response to a preset event trigger, wherein the interface information includes a first interactive object and an interactive focus, and the interactive focus represents the interactive position in the graphical interactive interface. The second obtaining module is used to obtain contextual information of the request information based on the spatial distribution information of the first interactive object relative to the interactive focus; The information generation module is used to generate response information corresponding to the request information based on the context information and the request information.

10. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the information processing method as described in any one of claims 1-9.

11. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the information processing method as described in any one of claims 1-9.