Method and device for positioning image elements by intelligent agent based on large language model
Through the agent method based on the large language model, the problem of low positioning efficiency and accuracy of image elements in the prior art is solved, and efficient positioning of rotating, scaling and deforming image elements is achieved.
Patent Information
- Application Number
- CN202510083162.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, template matching algorithms have poor effects when positioning rotated, scaled, and deformed image elements, and have low positioning efficiency for large images and large templates.
Using an agent method based on a large language model, the user's element marking instructions are obtained, image semantic analysis is performed, basic semantic results are generated, and language description results are generated through syntax template filling processing, and the position of image elements is finally determined in the target display image.
It improves the positioning efficiency and accuracy of image elements, can effectively handle rotated, scaled and deformed image elements, and is suitable for positioning large images and large templates.
Smart Images

Figure CN119992549A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method and device for locating image elements by an intelligent agent based on a large language model. Background Art
[0002] With the rapid development of image processing technology, Robotic Process Automation (RPA) software has been gradually applied to various large-scale or highly repetitive business scenarios, so as to simulate human control of various system software or automatically execute business processes through RPA software. Therefore, in the process of using RPA software, it is very important to locate the image elements in RPA software.
[0003] In the prior art, a template matching algorithm is usually used to locate image elements in RPA software. A sub-image element is slid in the image to be detected, and the similarity between the sub-image element and each image element in the image to be detected is calculated, so as to find the area that best matches the sub-image element.
[0004] However, although the template matching algorithm works well when the image to be detected has a fixed appearance and size, it works poorly for images to be detected that are rotated, scaled, and deformed, and is less efficient for large images and large templates. Therefore, how to perform image semantic analysis on captured image elements, realize image element positioning, and improve the efficiency and accuracy of image element positioning is a problem that needs to be solved urgently. Summary of the invention
[0005] The present invention provides a method and device for locating image elements by an intelligent agent based on a large language model, which can solve the problem of low efficiency and accuracy in locating image elements in RPA software.
[0006] According to one aspect of the present invention, a method for locating image elements by an agent based on a large language model is provided. The method is applied to an agent integrated with a large semantic model. The method comprises:
[0007] Acquire an element marking instruction input by a user based on a preset template image, and determine a target image element corresponding to the preset template image based on the element marking instruction;
[0008] Performing image semantic analysis on the target image element based on a preset semantic large model to generate a basic semantic result corresponding to the target image element;
[0009] Fill and process the basic semantic result based on a preset grammar template set to generate a basic language description result corresponding to the target image element;
[0010] Based on the preset semantic large model and the basic language description result, a target language description result corresponding to the target image element is determined in the target display image.
[0011] According to another aspect of the present invention, a device for locating image elements based on a large language model is provided, the device is applied to an agent integrated with a large semantic model, and the device comprises:
[0012] An element marking module, used to obtain an element marking instruction input by a user based on a preset template image, and determine a target image element corresponding to the preset template image based on the element marking instruction;
[0013] A semantic analysis module, used to perform image semantic analysis on the target image element based on a preset semantic large model, and generate a basic semantic result corresponding to the target image element;
[0014] A data filling module, used for filling and processing the basic semantic results based on a preset grammar template set to generate a basic language description result corresponding to the target image element;
[0015] A result generation module is used to determine the target language description result corresponding to the target image element in the target display image based on the preset semantic large model and the basic language description result.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for locating image elements based on a large language model by an intelligent agent as described in any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for locating image elements by an intelligent agent based on a large language model as described in any embodiment of the present invention when executed.
[0021] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the method for locating image elements by an intelligent agent based on a large language model as described in any embodiment of the present invention.
[0022] The technical solution of the embodiment of the present invention determines the target image element corresponding to the preset template image through the element marking instruction input by the user based on the preset template image. Then, the target image element is subjected to image semantic analysis based on the preset semantic big model, and the basic semantic result corresponding to the target image element is generated. Furthermore, the basic semantic result is filled and processed based on the preset grammar template set, and the basic language description result corresponding to the target image element is generated. Finally, the target language description result corresponding to the target image element is determined in the target display image based on the preset semantic big model and the basic language description result. Since the image elements are positioned by using semantic relationships, the problem of low positioning efficiency and accuracy of image elements in RPA software is solved, and image semantic analysis can be performed on the captured image elements to realize image element positioning, thereby improving the positioning efficiency and accuracy of image elements.
[0023] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 is a flow chart of a method for locating image elements based on a large language model provided by an intelligent agent according to Embodiment 1 of the present invention;
[0026] Figure 2 is a flow chart of a method for locating image elements based on a large language model provided by an intelligent agent according to a second embodiment of the present invention;
[0027] Figure 3 is a flowchart of an optional method for locating image elements by an intelligent agent based on a large language model according to Embodiment 2 of the present invention;
[0028] Figure 4 is a schematic diagram of a process of locating image elements by an intelligent agent based on a large language model according to Embodiment 2 of the present invention;
[0029] Figure 5 is a structural schematic diagram of a device for locating image elements based on a large language model provided by an intelligent agent according to a third embodiment of the present invention;
[0030] Figure 6It is a structural schematic diagram of an electronic device for implementing the method for locating image elements by an intelligent agent based on a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "target", "original", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0033] Embodiment 1
[0034] Figure 1 The flowchart of a method for locating image elements by an agent based on a large language model provided in the first embodiment of the present invention is applicable to the situation where an agent is used to locate image elements captured in RPA software. The method can be executed by a device for locating image elements by an agent based on a large language model. The device for locating image elements by an agent based on a large language model can be implemented in the form of hardware and / or software. The device for locating image elements by an agent based on a large language model can be configured in an electronic device. Figure 1 As shown, the method includes:
[0035] S110: Acquire an element marking instruction input by a user based on a preset template image, and determine a target image element corresponding to the preset template image based on the element marking instruction.
[0036] The preset template image may refer to a pre-selected user interface (UI) containing multiple image elements. Exemplarily, the preset template image may be each UI in the RPA software. Generally, the preset template image may include buttons, text boxes, icons, etc. The target image element may refer to an image element selected for marking in the preset template image. Exemplarily, the target image element may be a button, a text box, etc.
[0037] Among them, the element marking instruction may refer to an instruction for marking an image element in a preset template image. Generally, the element marking instruction includes a marking method, a marking name, and the corresponding image element to be marked, that is, the target image element. The marking method may represent a method of defining the outline of the image element. Exemplarily, if the image element to be marked is an image element with a clear boundary, the marking method may be a rectangular frame, a circle, or an ellipse; if the image element to be marked is an image element of a complex shape, the marking method may be a polygon; if the image element to be marked is a special image element, such as a kitten or a puppy, the marking method may be a key point. The marking name may represent the user's naming of the image element to be marked.
[0038] It is worth noting that in the embodiment of the present invention, the marking method can be selected according to user needs, and the embodiment of the present invention does not specifically limit this.
[0039] S120: Perform image semantic analysis on the target image element based on a preset semantic large model to generate a basic semantic result corresponding to the target image element.
[0040] Among them, the preset semantic big model may refer to a pre-set big model for semantic analysis of image elements. Exemplarily, the preset semantic big model may be a pre-trained big language model. Generally, the big language model has a strong natural language understanding and generation capability, and can convert the user's natural language input into computer-understandable instructions. At the same time, the big language model is one of the core technologies for realizing the intelligence of the intelligent agent. The intelligent agent uses the big language model to parse complex queries, understand the context, and provide more accurate responses.
[0041] Among them, image semantic analysis may refer to an operation of analyzing and determining the hierarchical relationship or contextual relationship of image elements. Basic semantic results may refer to analysis results generated after image semantic analysis of the target image element. For example, if the target image element is a button, the basic semantic result may be that the button is located in the upper right corner of the screen, etc.
[0042] S130: Fill in and process the basic semantic result based on a preset grammar template set to generate a basic language description result corresponding to the target image element.
[0043] Among them, the preset grammar template may refer to a preset template for formatting basic semantic results. Exemplarily, the preset grammar template may be: <element type>: <visual attribute>, <location description>, <context description>, <parent form description>. Usually, different types of basic semantic results correspond to different preset grammar templates. The preset grammar template set may refer to a set composed of various preset grammar templates. The basic language description result may refer to the natural language description result obtained after filling in the basic semantic result.
[0044] In an optional embodiment, before the basic semantic result is filled and processed based on the preset grammar template set to generate the basic language description result corresponding to the target image element, it also includes: obtaining a preset basic image element information set; wherein the preset basic image element information set includes the basic image element name and the basic attribute information corresponding to the basic image element; combining and processing the basic image element name and the basic attribute information corresponding to the basic image element to generate a preset attribute dictionary; based on the preset attribute dictionary and the preset template generation rules, generating a preset grammar template set corresponding to the preset attribute dictionary.
[0045] Among them, the basic image element may refer to each image element with a high probability of appearing in the RPA software. Exemplarily, the basic image element may be a button or an icon. The basic image element information may refer to the relevant information corresponding to the basic image element. Exemplarily, the basic image element information may include basic attribute information and the basic image element name. The basic attribute information may refer to the relevant attribute information corresponding to the basic image element. Exemplarily, the basic attribute information may include color, shape, position, contextual relationship, etc. The preset basic image element may refer to a pre-set basic image element. The preset basic image element information set may refer to a set composed of each preset basic image element information.
[0046] Among them, the preset attribute dictionary may refer to a pre-set tool for collecting relevant information of the basic image element for subsequent query. Generally, the preset attribute dictionary may include attribute description, position description, context description and parent window description. The attribute description may include visual attributes such as color, shape and size. Exemplarily, the attribute description may be a red round button. The position description may be used to describe the relative position of the basic image element in the image. Exemplarily, the position description may be located in the lower right corner of the screen. The context description may be used to describe the relationship between the basic image element and the surrounding elements. Exemplarily, the context description may be a blue icon next to it. The parent window description may be used to describe the parent window where the basic image element is located. Exemplarily, the parent window description may be a login window.
[0047] Among them, the preset template generation rule may refer to a pre-set rule for performing template conversion on a description set in a preset attribute dictionary. Usually, a preset attribute dictionary contains multiple description sets. A description set contains relevant description information of a basic image element, that is, a description set contains attribute description, position description, context description and parent window description corresponding to a basic image element. Exemplarily, the preset template generation rule may be a rule for arranging element type, attribute description, position description, context description and parent window description in order of priority. It is worth noting that in the embodiment of the present invention, the preset template generation rule may be selected according to actual needs, and the embodiment of the present invention does not specifically limit this.
[0048] Specifically, before filling and processing the basic semantic results based on the preset grammar template set and generating the basic language description results corresponding to the target image elements, the basic image elements with a high probability of appearing in the RPA software can be obtained in advance to construct a preset basic image element information set. Then, the basic image element name and basic attribute information corresponding to the same basic image element in the preset basic image element information set are combined and processed to generate a preset attribute dictionary. Finally, each description set in the preset attribute dictionary is converted into a template according to the preset template generation rules. Exemplarily, if the description set includes element type, attribute description, position description, context description and parent form description, the template after the description set is converted can be: <element type>: <visual attribute>, <position description>, <context description>, <parent form description>. Thus, a preset grammar template set corresponding to the preset attribute dictionary is generated, providing an effective basis for the subsequent generation of basic language description results.
[0049] S140. Determine a target language description result corresponding to the target image element in the target display image based on the preset semantic large model and the basic language description result.
[0050] The target display image may refer to a UI having a similar architecture to the preset template image. Generally, the target display image and the preset template image may be displayed on different terminals. The target display image may contain the same image elements as the preset template image, but the position, color and size of the image elements may be different.
[0051] The target language description result may refer to the natural language description result of the target image element in the target display image. Exemplarily, the target language description result may include the position information and size information of the target image element in the target display image. Usually, the position information may be coordinates, and the size information may be width and height.
[0052] The technical solution of the embodiment of the present invention determines the target image element corresponding to the preset template image through the element marking instruction input by the user based on the preset template image. Then, the target image element is subjected to image semantic analysis based on the preset semantic big model, and the basic semantic result corresponding to the target image element is generated. Furthermore, the basic semantic result is filled and processed based on the preset grammar template set, and the basic language description result corresponding to the target image element is generated. Finally, the target language description result corresponding to the target image element is determined in the target display image based on the preset semantic big model and the basic language description result. Since the image elements are positioned by using semantic relationships, the problem of low positioning efficiency and accuracy of image elements in RPA software is solved, and image semantic analysis can be performed on the captured image elements to realize image element positioning, thereby improving the positioning efficiency and accuracy of image elements.
[0053] Embodiment 2
[0054] Figure 2 A flowchart of a method for locating image elements by an intelligent agent based on a large language model is provided in the second embodiment of the present invention. This embodiment is refined based on the above-mentioned embodiment. In this embodiment, the operation of performing image semantic analysis on the target image element based on the preset semantic large model to generate a basic semantic result corresponding to the target image element is specifically refined. Specifically, it may include: performing full-dimensional feature extraction on the target image element based on the preset feature extraction algorithm in the preset semantic large model to obtain a basic visual feature set corresponding to the target image element; performing image semantic analysis on the target image element based on the semantic relationship between each basic visual feature in the basic visual feature set to determine the basic semantic result corresponding to the target image element. Figure 2 As shown, the method includes:
[0055] S210: Acquire an element marking instruction input by a user based on a preset template image, and determine a target image element corresponding to the preset template image based on the element marking instruction.
[0056] Specifically, when the user inputs an element marking instruction in the preset template image, the target image element corresponding to the preset template image can be determined based on the element marking instruction. For example, if the user marks a red round button with a rectangular frame in the preset template image and names it "Button 1", the button can be used as the target image element.
[0057] It is worth noting that in the embodiment of the present invention, the number of target image elements can be one or more. If there are more than one target image elements, the processing flow of each target image element is implemented in parallel, and the embodiment of the present invention does not specifically limit this.
[0058] S220. Perform full-dimensional feature extraction on the target image element based on a preset feature extraction algorithm in a preset semantic large model to obtain a basic visual feature set corresponding to the target image element.
[0059] Feature extraction may refer to the operation of converting pixel information of image elements into more representative and interpretable feature vectors. Full-dimensional feature extraction may refer to feature extraction in all dimensional ranges. Exemplarily, full-dimensional feature extraction may include color feature extraction, shape feature extraction, texture feature extraction, and parent window feature extraction.
[0060] Among them, the preset feature extraction algorithm may refer to a preset algorithm for extracting features from image elements. Generally, the preset feature extraction algorithms corresponding to features of different dimensions are different. Exemplarily, if the dimensional feature is color feature extraction, the preset feature extraction algorithm may be a color histogram, a color moment, and a preset clustering algorithm. The color histogram is used to count the frequency of occurrence of different colors in the image element, which can effectively represent the overall color distribution of the image element. The color moment is used to calculate the mean, variance, and skewness of the color distribution, which can more compactly represent the color feature. The preset clustering algorithm is used to extract the most important colors in the image element, thereby ignoring the interference caused by the secondary colors and realizing the extraction of the main color. If the dimensional feature is a shape feature, the preset feature extraction algorithm may be an edge detection algorithm, a contour description algorithm, and a shape detection algorithm. The edge detection algorithm is used to extract the contour information of the image element. The contour description algorithm is used to describe the extracted contour information to obtain a shape feature vector. The shape detection algorithm is used to detect geometric shapes such as straight lines or circles in the image element to identify elements of a specific shape. If the dimensional feature is a texture feature, the preset feature extraction algorithm may be a texture feature statistics algorithm, a texture feature analysis algorithm, and a filtering algorithm. The texture feature statistical algorithm is used to describe the spatial relationship between pixels in the grayscale space of image elements, and can extract features such as texture roughness and directionality. The texture feature analysis algorithm is used to describe the grayscale relationship between a pixel and its surrounding pixels, and can effectively represent local texture information. The filtering algorithm is used to extract texture features of different directions and scales. If the dimensional feature is a parent window feature, the preset feature extraction algorithm can be a window detection algorithm and a hierarchical relationship recognition algorithm. The window detection algorithm is used to identify the window area in the image element. Exemplarily, the window detection algorithm can be a target detection algorithm. The hierarchical relationship recognition algorithm is used to analyze the position, size and inclusion relationship between window areas and construct a hierarchical structure of the window. Exemplarily, the hierarchical relationship recognition algorithm can be a deep learning model or a rule-based method. For example, if one window is completely contained in another window, the former can be considered to be a child window of the latter.
[0061] Among them, the basic visual features may refer to the feature extraction results obtained after feature extraction of the target image element. Exemplarily, the basic visual features may include color features, shape features, texture features, and parent window features. The basic visual feature set may refer to the set composed of each basic visual feature under the full dimension corresponding to the same target image element.
[0062] S230: Perform image semantic analysis on the target image element based on the semantic relationship between each basic visual feature in the basic visual feature set to determine a basic semantic result corresponding to the target image element.
[0063] Specifically, after obtaining the basic visual feature set corresponding to the target image element, the semantic relationship between each basic visual feature can be used to understand the context of the target image element to obtain the basic semantic result corresponding to the target image element, for example, the button is located in the upper right corner of the screen, or the blue circular mark is located in the login form, etc. This provides an effective basis for subsequent operations.
[0064] S240: Perform an index operation on the preset grammar template set based on the target image element to obtain a target grammar template corresponding to the target image element.
[0065] The target grammar template may refer to a grammar template in a preset grammar template set that matches the target image element. Exemplarily, the image element name of the target image element may be used to index in the preset grammar template set to obtain a grammar template that is consistent with the image element name as the target grammar template.
[0066] S250: Fill in and process the basic semantic result based on the target grammar template to generate an original language description result corresponding to the target image element.
[0067] The original language description result may refer to a preliminarily generated unprocessed language description result. Exemplarily, the original language description result may be a language description result obtained by directly performing a filling process using a target grammar template.
[0068] S260: Perform data verification on the original language description result based on preset data processing rules to obtain a basic language description result that meets the preset data processing rules.
[0069] Among them, the preset data processing rules may refer to the pre-set rules for data verification of the original language description results. In an optional embodiment, the preset data processing rules may include: preset semantic refining rules and preset consistency checking rules. The preset semantic refining rules may refer to the pre-set method for removing redundant information in the original language description results. Exemplarily, the preset semantic refining rules may be a method based on an encoder-decoder, a method based on an extractive summary, a method based on an abstract summary, and a fine-tuning method based on a pre-trained language model. In general, the preset semantic refining rules can be used to remove stop words that contribute less to semantics, replace multiple words that express the same meaning with a more representative word, simplify descriptions according to predefined rules, and compress lengthy descriptions into more concise expressions. The preset consistency checking rules may refer to the pre-set method for consistency and matchability verification of the original language description results. Exemplarily, the preset consistency checking rules may be a cross-image verification method, an adversarial sample testing method, and a split testing method.
[0070] Specifically, after determining the basic semantic result corresponding to the target image element, the image element name of the target image element can be used to perform an index operation in the preset grammar template set to obtain a grammar template consistent with the image element name as the target grammar template corresponding to the target image element. Afterwards, the basic semantic result is filled using the target grammar template to generate the original language description result corresponding to the target image element. Finally, the original language description result is verified in turn using the preset semantic refinement rules and preset consistency check rules in the preset data processing rules. In this way, a basic language description result that meets the preset data processing rules is obtained, which provides an effective basis for subsequent image element positioning and ensures the accuracy of the basic language description result.
[0071] S270: Parse the basic language description result based on a preset language parsing algorithm in the preset semantic big model to obtain a structured query result corresponding to the basic language description result.
[0072] The preset language parsing algorithm may refer to a preset algorithm for performing structural parsing on a language description result. The structured query result may refer to a parsing result obtained after parsing based on the preset language parsing algorithm.
[0073] S280. Perform semantic analysis on the target display image based on the preset semantic analysis algorithm in the preset semantic big model to obtain a set of candidate semantic results corresponding to the target display image, and match and process the set of candidate semantic results based on the structured query result to obtain a target semantic result that matches the structured query result as a target language description result corresponding to the target image element.
[0074] Among them, the preset semantic analysis algorithm may refer to a pre-set algorithm for performing semantic analysis on a display image. The selected semantic result may refer to a semantic analysis result obtained after performing semantic analysis based on the preset semantic analysis algorithm. The set of selected semantic results may refer to a set of selected semantic results corresponding to the same target display image. The target semantic result may refer to a semantic result in the set of selected semantic results that matches the structured query result.
[0075] Specifically, after obtaining the basic language description result that meets the preset data processing rules, the preset language parsing algorithm in the preset semantic big model can be used to parse the basic language description result to obtain the structured query result corresponding to the basic language description result. At the same time, the target display image is semantically analyzed based on the preset semantic analysis algorithm in the preset semantic big model to obtain a set of candidate semantic results corresponding to the target display image. Then, the structured query result is used to perform data matching in the set of candidate semantic results to obtain the target semantic result that matches the structured query result as the target language description result corresponding to the target image element. Therefore, by using semantic relationships to locate image elements, the efficiency and accuracy of locating image elements can be improved.
[0076] S290: Acquire an image element adjustment instruction corresponding to the target image element, and implement the image element adjustment instruction in a target display image based on the target language description result.
[0077] The image element adjustment instruction may refer to an instruction input by a user to adjust a target image element. Exemplarily, the image element adjustment instruction may be a position movement instruction or a size adjustment instruction. Usually, the user may initiate the image element adjustment instruction by clicking a mouse or inputting a keyboard.
[0078] Specifically, after generating the target language description result corresponding to the target image element, the user's image element adjustment instruction for the target image element can be obtained, and then the image element adjustment instruction is implemented in the target display image based on the target language description result. Therefore, by combining with intelligent agents and large language models, RPA can realize more complex automation scenarios, for example, automatically executing corresponding operations after parsing user requests to complete dynamic adjustment of target image elements.
[0079] The technical solution of the embodiment of the present invention determines the target image element corresponding to the preset template image through the element marking instruction input by the user based on the preset template image. Then, based on the preset feature extraction algorithm in the preset semantic large model, the target image element is subjected to full-dimensional feature extraction to obtain the basic visual feature set corresponding to the target image element, and the target image element is subjected to image semantic analysis based on the semantic relationship between each basic visual feature in the basic visual feature set to determine the basic semantic result corresponding to the target image element. Furthermore, based on the target image element, the preset grammar template set is indexed to obtain the target grammar template corresponding to the target image element, and the basic semantic result is filled and processed based on the target grammar template to generate the original language description result corresponding to the target image element, and the original language description result is subjected to data verification based on the preset data processing rules to obtain the basic language description result that meets the preset data processing rules. Furthermore, the basic language description result is parsed based on the preset language parsing algorithm in the preset semantic big model to obtain the structured query result corresponding to the basic language description result, and the target display image is semantically analyzed based on the preset semantic analysis algorithm in the preset semantic big model to obtain the candidate semantic result set corresponding to the target display image, and the candidate semantic result set is matched and processed based on the structured query result to obtain the target semantic result matching the structured query result as the target language description result corresponding to the target image element. Finally, the image element adjustment instruction corresponding to the target image element is obtained, and the image element adjustment instruction is implemented in the target display image based on the target language description result. Since the image elements are positioned by using semantic relationships, the problem of low positioning efficiency and accuracy of image elements in RPA software is solved, and image semantic analysis can be performed on the captured image elements to realize image element positioning, thereby improving the positioning efficiency and accuracy of image elements.
[0080] Figure 3The flowchart of an optional method for locating image elements by an intelligent agent based on a large language model provided by an embodiment of the present invention is shown. Specifically, first, an element marking instruction input by a user based on a preset template image is obtained, and a target image element corresponding to the preset template image is determined based on the element marking instruction. Then, an image semantic analysis is performed on the target image element based on the preset semantic large model to generate a basic semantic result corresponding to the target image element, an index operation is performed on a preset grammar template set based on the target image element to obtain a target grammar template corresponding to the target image element, a basic semantic result is filled and processed based on the target grammar template to generate an original language description result corresponding to the target image element, and a data verification is performed on the original language description result based on a preset data processing rule to obtain a basic language description result that meets the preset data processing rule. Further, the basic language description result is parsed and processed based on a preset language parsing algorithm in the preset semantic large model to obtain a structured query result corresponding to the basic language description result, a semantic analysis is performed on the target display image based on a preset semantic analysis algorithm in the preset semantic large model to obtain a candidate semantic result set corresponding to the target display image, and the candidate semantic result set is processed based on the structured query result matching to obtain a target semantic result matching the structured query result as the target language description result corresponding to the target image element. Finally, the image element adjustment instruction corresponding to the target image element is obtained, and the image element adjustment instruction is implemented in the target display image based on the target language description result. Thus, a method for locating image elements by an intelligent agent based on a large language model is implemented.
[0081] Figure 4The figure shows a schematic diagram of a process of locating image elements by an intelligent agent based on a large language model provided by an embodiment of the present invention. Specifically, the process of locating image elements by an intelligent agent based on a large language model includes three parts: a preset semantic large model, a source-side RPA designer, and a target-side RPA designer. Among them, the source-side RPA includes a preset template image. The target-side RPA includes a target display image. The source-side RPA designer is used to trigger an element marking instruction input by a user based on a preset template image, and determine the target image element corresponding to the preset template image based on the element marking instruction. The preset semantic large model is used to perform image semantic analysis on the target image element generated by the source-side RPA designer, generate a basic semantic result corresponding to the target image element, and fill in the basic semantic result based on a preset grammar template set to generate a basic language description result corresponding to the target image element. The target-side RPA designer is used to display the target display image. The preset semantic large model is used to locate the target image element in the target display image based on the basic language description result, determine the target language description result corresponding to the target image element, and return the target language description result to the target-side RPA designer. It is worth noting that in an embodiment of the present invention, in order to avoid data deviation caused by the different resolutions of each device, after the preset semantic big model generates the basic language description result corresponding to the target image element, the basic language description result can be returned to the source-end RPA designer through the preset semantic big model to revise the target image element based on the basic language description result.
[0082] Embodiment 3
[0083] Figure 5 The structure diagram of a device for locating image elements based on a large language model provided by the third embodiment of the present invention is shown in FIG. The device is applied to an agent integrating a large semantic model, such as Figure 5 As shown, the device includes: an element marking module 310, a semantic analysis module 320, a data filling module 330 and a result generation module 340;
[0084] The element marking module 310 is used to obtain an element marking instruction input by a user based on a preset template image, and determine a target image element corresponding to the preset template image based on the element marking instruction;
[0085] A semantic analysis module 320, configured to perform image semantic analysis on the target image element based on a preset semantic macro model, and generate a basic semantic result corresponding to the target image element;
[0086] A data filling module 330, configured to fill in and process the basic semantic result based on a preset grammar template set to generate a basic language description result corresponding to the target image element;
[0087] The result generation module 340 is used to determine the target language description result corresponding to the target image element in the target display image based on the preset semantic large model and the basic language description result.
[0088] The technical solution of the embodiment of the present invention determines the target image element corresponding to the preset template image through the element marking instruction input by the user based on the preset template image. Then, the target image element is subjected to image semantic analysis based on the preset semantic big model, and the basic semantic result corresponding to the target image element is generated. Furthermore, the basic semantic result is filled and processed based on the preset grammar template set, and the basic language description result corresponding to the target image element is generated. Finally, the target language description result corresponding to the target image element is determined in the target display image based on the preset semantic big model and the basic language description result. Since the image elements are positioned by using semantic relationships, the problem of low positioning efficiency and accuracy of image elements in RPA software is solved, and image semantic analysis can be performed on the captured image elements to realize image element positioning, thereby improving the positioning efficiency and accuracy of image elements.
[0089] Optionally, the semantic analysis module 320 may be specifically used for:
[0090] Performing full-dimensional feature extraction on the target image element based on a preset feature extraction algorithm in a preset semantic large model to obtain a basic visual feature set corresponding to the target image element;
[0091] Image semantic analysis is performed on the target image element based on the semantic relationship between each basic visual feature in the basic visual feature set to determine a basic semantic result corresponding to the target image element.
[0092] Optionally, the device for locating image elements based on an intelligent agent of a large language model may also include: a template set construction module, which is used to obtain a preset basic image element information set before filling and processing the basic semantic results based on the preset grammar template set to generate a basic language description result corresponding to the target image element; wherein the preset basic image element information set includes the basic image element name and the basic attribute information corresponding to the basic image element; combine and process the basic image element name and the basic attribute information corresponding to the basic image element to generate a preset attribute dictionary; based on the preset attribute dictionary and preset template generation rules, generate a preset grammar template set corresponding to the preset attribute dictionary.
[0093] Optionally, the data filling module 330 may be specifically used for:
[0094] Performing an index operation on the preset grammar template set based on the target image element to obtain a target grammar template corresponding to the target image element;
[0095] Fill in and process the basic semantic result based on the target grammar template to generate an original language description result corresponding to the target image element;
[0096] The original language description result is verified based on the preset data processing rules to obtain a basic language description result that meets the preset data processing rules.
[0097] Optionally, the preset data processing rules may include: preset semantic refining rules and preset consistency checking rules.
[0098] Optionally, the result generation module 340 may be specifically used for:
[0099] Parsing the basic language description result based on a preset language parsing algorithm in the preset semantic big model to obtain a structured query result corresponding to the basic language description result;
[0100] Based on the preset semantic analysis algorithm in the preset semantic big model, a semantic analysis is performed on the target display image to obtain a set of candidate semantic results corresponding to the target display image, and the set of candidate semantic results is matched and processed based on the structured query result to obtain a target semantic result that matches the structured query result as a target language description result corresponding to the target image element.
[0101] Optionally, the device for locating image elements by an intelligent agent based on a large language model may also include: a post-processing module for obtaining image element adjustment instructions corresponding to the target image element after determining the target language description result corresponding to the target image element in the target display image based on the preset semantic large model and the basic language description result, and implementing the image element adjustment instructions in the target display image based on the target language description result.
[0102] The device for locating image elements based on a large language model provided by an embodiment of the present invention can execute the method for locating image elements based on a large language model provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0103] Embodiment 4
[0104] Figure 6A schematic diagram of an electronic device 410 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0105] like Figure 6 As shown, the electronic device 410 includes at least one processor 420, and a memory connected to the at least one processor 420 in communication, such as a read-only memory (ROM) 430, a random access memory (RAM) 440, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 420 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 430 or the computer program loaded from the storage unit 490 to the random access memory (RAM) 440. In the RAM 440, various programs and data required for the operation of the electronic device 410 can also be stored. The processor 420, the ROM 430, and the RAM 440 are connected to each other via a bus 450. An input / output (I / O) interface 460 is also connected to the bus 450.
[0106] Multiple components in the electronic device 410 are connected to the I / O interface 460, including: an input unit 470, such as a keyboard, a mouse, etc.; an output unit 480, such as various types of displays, speakers, etc.; a storage unit 490, such as a disk, an optical disk, etc.; and a communication unit 4100, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 4100 allows the electronic device 410 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0107] The processor 420 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 420 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 420 performs the various methods and processes described above, such as a method for locating image elements based on a large language model-based agent.
[0108] The method includes:
[0109] Acquire an element marking instruction input by a user based on a preset template image, and determine a target image element corresponding to the preset template image based on the element marking instruction;
[0110] Performing image semantic analysis on the target image element based on a preset semantic large model to generate a basic semantic result corresponding to the target image element;
[0111] Fill and process the basic semantic result based on a preset grammar template set to generate a basic language description result corresponding to the target image element;
[0112] Based on the preset semantic large model and the basic language description result, a target language description result corresponding to the target image element is determined in the target display image.
[0113] In some embodiments, the method for locating image elements based on an agent of a large language model may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 490. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 410 via the ROM 430 and / or the communication unit 4100. When the computer program is loaded into the RAM 440 and executed by the processor 420, one or more steps of the method for locating image elements based on an agent of a large language model described above may be performed. Alternatively, in other embodiments, the processor 420 may be configured to execute the method for locating image elements based on an agent of a large language model by any other appropriate means (e.g., by means of firmware).
[0114] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0115] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0116] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0117] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0118] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0119] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0120] The present application also discloses a computer program product, which includes a computer program. When the computer program is executed by a processor, the method for locating image elements by an intelligent agent based on a large language model provided in any embodiment of the present application is implemented. The program product and the method for locating image elements by an intelligent agent based on a large language model disclosed in each embodiment of the present application belong to the same inventive concept, and therefore will not be described in detail here.
[0121] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0122] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for locating image elements by an intelligent agent based on a large language model, characterized in that: The method is applied to an intelligent agent integrating a semantic big model, and the method comprises: Acquire an element marking instruction input by a user based on a preset template image, and determine a target image element corresponding to the preset template image based on the element marking instruction; Performing image semantic analysis on the target image element based on a preset semantic large model to generate a basic semantic result corresponding to the target image element; Filling and processing the basic semantic results based on a preset grammar template set to generate a basic language description result corresponding to the target image element; Based on the preset semantic large model and the basic language description result, a target language description result corresponding to the target image element is determined in the target display image.
2. The method according to claim 1, characterized in that The performing image semantic analysis on the target image element based on the preset semantic large model to generate a basic semantic result corresponding to the target image element includes: Performing full-dimensional feature extraction on the target image element based on a preset feature extraction algorithm in a preset semantic large model to obtain a basic visual feature set corresponding to the target image element; Image semantic analysis is performed on the target image element based on the semantic relationship between each basic visual feature in the basic visual feature set to determine a basic semantic result corresponding to the target image element.
3. The method according to claim 1, characterized in that Before filling and processing the basic semantic result based on the preset grammar template set to generate the basic language description result corresponding to the target image element, the method further includes: Obtaining a preset basic image element information set; wherein the preset basic image element information set includes a basic image element name and basic attribute information corresponding to the basic image element; Combining and processing the basic image element name and the basic attribute information corresponding to the basic image element to generate a preset attribute dictionary; Based on the preset attribute dictionary and the preset template generation rule, a preset grammar template set corresponding to the preset attribute dictionary is generated.
4. The method according to claim 1, characterized in that: The filling and processing of the basic semantic result based on the preset grammar template set to generate a basic language description result corresponding to the target image element includes: Performing an index operation on the preset grammar template set based on the target image element to obtain a target grammar template corresponding to the target image element; Fill in and process the basic semantic result based on the target grammar template to generate an original language description result corresponding to the target image element; The original language description result is verified based on the preset data processing rules to obtain a basic language description result that meets the preset data processing rules.
5. The method according to claim 4, characterized in that The preset data processing rules include: preset semantic refining rules and preset consistency checking rules.
6. The method according to claim 1, characterized in that The step of determining the target language description result corresponding to the target image element in the target display image based on the preset semantic large model and the basic language description result includes: Parsing the basic language description result based on a preset language parsing algorithm in the preset semantic big model to obtain a structured query result corresponding to the basic language description result; Based on the preset semantic analysis algorithm in the preset semantic big model, a semantic analysis is performed on the target display image to obtain a set of candidate semantic results corresponding to the target display image, and the set of candidate semantic results is matched and processed based on the structured query result to obtain a target semantic result that matches the structured query result as a target language description result corresponding to the target image element.
7. The method according to claim 1, characterized in that After determining the target language description result corresponding to the target image element in the target display image based on the preset semantic large model and the basic language description result, the method further includes: An image element adjustment instruction corresponding to the target image element is obtained, and the image element adjustment instruction is implemented in a target display image based on the target language description result.
8. A device for locating image elements based on a large language model, characterized in that: The device is applied to an intelligent agent of an integrated semantic large model, and the device comprises: An element marking module, used to obtain an element marking instruction input by a user based on a preset template image, and determine a target image element corresponding to the preset template image based on the element marking instruction; A semantic analysis module, used to perform image semantic analysis on the target image element based on a preset semantic large model, and generate a basic semantic result corresponding to the target image element; A data filling module, used for filling and processing the basic semantic results based on a preset grammar template set to generate a basic language description result corresponding to the target image element; A result generation module is used to determine the target language description result corresponding to the target image element in the target display image based on the preset semantic large model and the basic language description result.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method for locating image elements based on a large language model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for locating image elements based on a large language model according to any one of claims 1 to 7 when executed.