Zero code form automatic generation method based on AI image recognition

By integrating AI image recognition, graph neural network and reinforcement learning technology on the zero-code platform, automatically identifying and understanding form components and their logical relationships in user sketches or screenshots, the problem of automated form generation in the existing technology is solved, and efficient and accurate interactive form generation is achieved.

CN119992582AActive Publication Date: 2025-05-13NANJING DIGITAL YOUDAO TECH CO LTD

Patent Information

Application Number
CN202510453597.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The prior art is difficult to automatically identify the form components and their logical relationships in sketches or screenshots provided by users, resulting in the inability to realize the automatic generation of user design sketches or screenshots to interactive forms.

Method used

Using AI-based image recognition technology, form components are identified through object detection models, Hough transform detection connection lines are used, text detection is used to locate text content, build dynamic relationship diagrams, and analyze logical relationships through graph neural networks, combine reinforcement learning to optimize user interaction paths, and generate interactive forms.

Benefits of technology

It realizes automated generation from user sketches or screenshots to interactive forms, lowers development thresholds, improves design efficiency and user experience, and overcomes the limitations of manual configuration and manual rules in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992582A_ABST
    Figure CN119992582A_ABST
Patent Text Reader

Abstract

The invention discloses an AI image recognition-based zero code form automatic generation method, which comprises the following steps of: receiving a sketch or a screenshot provided by a user, and carrying out image Gaussian filtering noise reduction, simplified Canny algorithm edge detection and Otsu method binarization processing on the sketch or the screenshot; using an object detection model to identify form components in the preprocessed image, and outputting bounding box coordinates of each component; detecting a connecting line through Hough transform, judging whether the connecting line is an arrow or not, and extracting a potential relationship between the form components; constructing a dynamic relation graph according to the extracted potential relation, and analyzing a logic relation in the dynamic relation graph by using a graph neural network; a user interaction path is optimized through reinforcement learning, and an interactive form is generated according to the logic relation obtained through analysis and the optimized interaction path; according to the method, the interactive form can be automatically generated from the sketch or the screenshot provided by the user, the form development threshold is lowered, and the form development efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent-driven application development technology, and in particular to a method for automatically generating zero-code forms based on AI image recognition. Background Art

[0002] With the rapid development of information technology, the demand for form design and development by enterprises and individual users is increasing. Traditional form development methods usually require professional developers to implement manual coding. This method not only has a long development cycle and high cost, but also makes it difficult to quickly respond to changes in user needs. In recent years, zero-code or low-code development platforms have gradually emerged. Such platforms allow users to quickly build forms by dragging components through a graphical interface, greatly improving development efficiency. However, even for zero-code platforms, users still need to manually select and layout form components, especially when faced with complex forms or a large number of form components. The user's design efficiency and accuracy are still greatly limited. In addition, existing zero-code platforms usually lack the ability to automatically recognize and understand user sketches or existing form screenshots, and cannot automatically generate interactive forms directly from design sketches or screenshots provided by users. Users still need to perform a lot of repetitive work, which seriously affects the efficiency of form design and user experience.

[0003] On the other hand, deep learning technology in artificial intelligence technology has made significant progress in areas such as image recognition, object detection, and relationship reasoning. Technologies such as the YOLO series of algorithms, graph neural networks (GNNs), and reinforcement learning (RL) have gradually matured and been widely used. However, current technical means cannot realize a method to automatically identify form components in sketches or screenshots provided by users and accurately infer the logical relationship between components, resulting in the inability to automatically generate interactive forms from user-designed sketches or screenshots. In addition, existing means usually rely on manually defined rules in inferring the relationship between form components, lack a deep understanding of the complex logical relationship between form components, and it is difficult to achieve efficient and accurate automatic form generation. Therefore, how to accurately and efficiently realize the automatic generation of user sketches or screenshots to interactive forms has become a technical problem that needs to be solved urgently in this field. Summary of the invention

[0004] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0005] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides a method for automatically generating a zero-code form based on AI image recognition to solve the problems raised in the background technology.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for automatically generating a zero-code form based on AI image recognition, comprising: Receiving a sketch or screenshot provided by a user, performing image preprocessing on the sketch or screenshot, and using an object detection model to identify form components in the preprocessed image, and extracting potential relationships among the form components; According to the potential relationships in the form components, the logical relationships in the form components are obtained and a dynamic relationship graph is constructed, and the logical relationships in the dynamic relationship graph are analyzed using a graph neural network; The user interaction path is optimized through reinforcement learning, and an interactive form is generated based on the logical relationship obtained through analysis and the optimized interaction path.

[0007] As a preferred solution of the method for automatically generating a zero-code form based on AI image recognition according to the present invention, the image preprocessing of the sketch or screenshot includes: The sketch or screenshot is regarded as an image, image noise is suppressed by Gaussian filtering, image edges are detected by a simplified Canny algorithm, and the image is binarized by the Otsu method.

[0008] As a preferred solution of the zero-code form automatic generation method based on AI image recognition described in the present invention, the form components in the preprocessed image are identified using an object detection model, including: Taking YOLOv5 as the object detection model, the graphic image annotation tool labelme is used to generate a set of labeled images of form components. The preprocessed image is taken as input, and the form components in the preprocessed image are identified according to the set of labeled images of the form components, and the bounding box coordinates of each form component are output.

[0009] As a preferred solution of the zero-code form automatic generation method based on AI image recognition described in the present invention, the potential relationship in the form components is extracted, including: Based on the Hough transform, the connecting lines in the form are detected. By setting the minimum line length and maximum gap and analyzing the pixel distribution near the end points of the connecting line, it is determined whether the connecting line is an arrow. If it is an arrow, the coordinates of the starting point and end point of the arrow and the direction of the arrow are output.

[0010] As a preferred solution of the zero-code form automatic generation method based on AI image recognition described in the present invention, wherein: according to the potential relationship in the form component, the logical relationship in the form component is obtained and a dynamic relationship diagram is constructed, including: Locate the text position and the text content at the position by using a text detection method, and obtain the logical relationship in the form component according to the coordinates of the starting point and end point of the arrow and the direction of the arrow; Each form component is defined as a node in a dynamic relationship graph, and the node attributes include bounding box coordinates; If a connection line is detected, a directed edge is added between the two nodes. If no connection line is detected, but the spatial distance between the two nodes is less than the preset threshold and there is a dependency relationship between the categories, an undirected edge is added; Connect all the directed or undirected edges formed.

[0011] As a preferred solution of the method for automatically generating a zero-code form based on AI image recognition described in the present invention, wherein: using a graph neural network to infer the logical relationship between the nodes includes: According to the dynamic relationship graph, a node feature matrix is ​​created; The node features are updated through the attention mechanism of the graph neural network, and each edge where the node is located is classified to output a dynamic relationship graph with labels.

[0012] As a preferred solution of the method for automatically generating a zero-code form based on AI image recognition described in the present invention, the method optimizes the user interaction path through reinforcement learning, including: The user interaction path is modeled through the Markov decision process. The state in the Markov decision process is represented as the form filling state, the action is represented as the user operation process, and the reward is defined as the efficiency of the user filling in the form.

[0013] As a preferred solution of the method for automatically generating a zero-code form based on AI image recognition described in the present invention, it also includes: A proximal strategy optimization algorithm is used in the Markov decision process to update the user interaction path in real time according to the state and action.

[0014] As a preferred solution of the method for automatically generating a zero-code form based on AI image recognition described in the present invention, generating an interactive form according to the analyzed logical relationship and the optimized interaction path includes: The edge relationships in the labeled dynamic relationship graph are converted into JavaScript rules and embedded into static elements rendered based on HTML / CSS to generate interactive forms.

[0015] Compared with the prior art, the invention has the following beneficial effects: 1. The present invention uses AI image recognition technology to automatically identify form components from sketches or screenshots provided by users and extract their potential and logical relationships, eliminating the tedious steps of manual configuration. In addition, the method does not require programming knowledge, greatly lowering the development threshold, allowing non-professional users to quickly complete form design, reducing repetitive work and improving design efficiency; 2. Using graph neural network (GNN) technology, we analyze the potential and logical relationships in form components based on dynamic relationship graphs, which can deeply understand the dependencies and interaction logic between components, overcome the limitations of traditional manual rules, and lay the foundation for the formation of fully functional interactive forms; 3. Through reinforcement learning (RL) technology, based on the Markov decision process and proximal strategy optimization algorithm, the user interaction path is intelligently optimized, and the filling process is adjusted according to the design of the form, making user operations smoother and more efficient, effectively improving the convenience of form use and user satisfaction, and solving the problem of insufficient user experience in existing technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them: Figure 1 The present invention is an overall flow chart of a method for automatically generating a zero-code form based on AI image recognition according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0018] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0019] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0020] The present invention is described in detail with reference to schematic diagrams. When describing the embodiments of the present invention, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.

[0021] At the same time, in the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper, lower, inner and outer" are based on the directions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. In addition, the terms "first, second or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0022] In the present invention, unless otherwise clearly specified and limited, the terms "install, connect, connect" should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection or an integral connection; it can also be a mechanical connection, an electrical connection or a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0023] Example 1 Reference Figure 1 , which is the first embodiment of the present invention, and provides a method for automatically generating a zero-code form based on AI image recognition, comprising: S1. Receive a sketch or screenshot provided by a user, perform image preprocessing on the sketch or screenshot, and use an object detection model to identify form components in the preprocessed image, and extract potential relationships among the form components; Specifically, the user uploads a hand-drawn form sketch (drawn on paper and photographed) or a screenshot (which can be a screenshot of a website page or a custom screenshot, such as a WeChat screenshot, etc.), saves the sketch or screenshot in a common image format such as JPG or PNG, and places it in a corresponding picture folder; It should be explained that since the original image (image file) placed in the picture folder may have noise, uneven lighting or blurred edges, if the original image is directly processed by the object detection model, it will affect the accuracy of the recognition form component. Therefore, the original image needs to be pre-processed before recognition; Furthermore, the noise of the original image is suppressed by Gaussian filtering, where Gaussian filtering is an image processing filtering technology based on the normal distribution function. Its main function is to smooth the image, thereby reducing noise. It can retain more image details than simple mean filtering, and achieve the purpose of accurately identifying form components. The specific steps are as follows: First, set or select the Gaussian kernel size and standard deviation to generate a Gaussian kernel, i.e., a weight matrix; the Gaussian kernel size is obtained by the size of the Gaussian filter; then, take the original image as input, perform a two-dimensional convolution with the original image through the Gaussian filter, and for each pixel in the original image, take the pixel value and the pixel values ​​around the pixel, perform weighted summation according to the weight matrix, and form a new pixel value; repeat the convolution operation until the entire image is traversed through the pixel points to obtain the denoised original image; It should be noted that the noise in the original image is removed by Gaussian filtering, which can make the original image clearer; Furthermore, the image edge is detected by a simplified Canny algorithm, wherein the Canny edge detection algorithm is an edge detection method used in image processing. Its main feature is that it can accurately detect the edge in the image to highlight the boundary of the form component in the image. The specific steps are as follows: Based on the original image processed by the Gaussian filter (after denoising), the direction of a certain pixel in the image is taken out and divided into the horizontal (x) component and the vertical (y) component. By using the Sobel operator to associate with the image, the gradient value of the pixel on the horizontal component and the vertical component is calculated to obtain the gradient value of the pixel, and the gradient value is used as the edge area of ​​the image extraction; It should be explained that, considering that the form borders in the image are usually straight lines or rectangles, they have high contrast, simple structure, and obvious boundaries, unlike the complex edges in natural images, based on the actual characteristics of the form borders, the form borders can be extracted using the gradient image with simple threshold processing, without the need for non-maximum suppression and double threshold processing, which simplifies the Canny algorithm for the detection of image edges and improves the efficiency of image processing; Furthermore, the image after edge detection is binarized by the Otsu method to distinguish the form components from other objects in the image under different lighting (uniform) conditions. The Otsu method is an automatic threshold selection algorithm for image segmentation. Its purpose is to select an optimal threshold to maximize the grayscale difference between the foreground and background in the image, thereby clearly distinguishing the two types of pixels. The specific steps are as follows: According to the image after edge detection, the number of pixels of each gray level (0-255) in the image is counted, and the probability of each gray level is calculated. All gray levels are used as thresholds. By traversing all gray levels, the image is divided into foreground class and background class, and the inter-class variance of the foreground class and the background class is calculated. The maximum threshold of the inter-class variance is found, and the pixel values ​​in the image greater than the maximum threshold are represented as foreground, and the pixel values ​​less than or equal to the maximum threshold are represented as background; It should be explained that in an image, form components and backgrounds usually have different grayscale features (for example, text boxes are usually dark and backgrounds are usually light). The Otsu method can automatically find the optimal threshold based on the grayscale distribution of the image, eliminating the need for manual setting. Furthermore, YOLOv5 is used as the object detection model, and a set of labeled images of form components is generated through the graphic image annotation tool labelme. The preprocessed image is used as input, and the form components (text boxes, drop-down menus, etc.) in the preprocessed image are identified according to the set of labeled images of the form components, and the bounding box coordinates of each form component are output; It should be noted that there is an essential difference between a bounding box and a text box. A bounding box only represents a location and does not contain input content, while a text box can input content. In addition, each form component has a bounding box, and the component body wrapped by the bounding box is a specific form component. It should be noted that labelme is an open source image annotation tool that supports bounding box annotation. Its operation steps are: Open labelme, load the form image, select the "Create Rectangle" tool, draw bounding boxes, enter (label) names for each drawn bounding box, and save the labeling results to get the corresponding JSON file; It should be noted that YOLOv5 can only recognize rectangles, that is, the input image must be a matrix. If you want to use images of other shapes, you can do so by padding or cropping. In addition, since YOLOv5 uses TXT files to store annotation information, and labelme generates JSON files, you need to convert the JSON files to the format required by YOLOv5. Specifically, the JSON file is converted into a TXT file so that each image in the annotated image set corresponds to a TXT file, and each line in the TXT file can be represented as a form component. YOLOv5 is used to traverse each line in the TXT file and output the bounding box coordinates of each form component in the TXT file. It should be noted that to detect the connection line in the form and determine whether the connection line is an arrow, Hough Transform can be used. Hough Transform is a feature extraction technology, mainly used to detect geometric shapes in images, and arrows usually play a guiding and prompting role in user sketches of form design, for example, "basic information → contact information → confirm submission", "fill in the ID number first → then upload the ID photo", etc.; Furthermore, the connection lines (straight lines) in the form are detected based on the Hough transform, and the potential relationship between the form components is obtained by setting the minimum line length and the maximum gap and analyzing the pixel distribution near the end points of the line segments of the connection lines; Specifically, a minimum line length is set, and only line segments whose length exceeds the minimum line length will be detected; a maximum gap is set, and if the broken gap in the line segment is smaller than the maximum gap, it will be connected into a complete line segment (forming a connecting line), and all the detected complete line segments are gathered into a connecting line list, in which each connecting line is represented by the starting point coordinates and the end point coordinates; Furthermore, it is determined whether the connecting line is an arrow. If it is an arrow, the coordinates of the starting point and the end point of the arrow and the direction of the arrow are output; It should be noted that arrows typically have an extra pixel at one end that forms a tip (such as a triangle or V shape); Furthermore, a small area composed of pixels is defined near the starting point coordinates and the end point coordinates of each connecting line, such as a rectangle of 10×10 pixels; the number of dark pixels in the area is counted, and if the pixel density near any of the endpoints is greater than or equal to the pixel density of the main body of the connecting line, it is considered that there is an arrow; if there is no obvious arrow feature at both ends, that is, the pixel density near the two endpoints is lower than the pixel density of the main body of the connecting line, it is considered that there is no arrow; It should be noted that, for existing arrows, an arrow-shaped template may be predefined, and by comparing the template with the existing arrows, misidentification of the arrows may be prevented; Specifically, the direction of the arrow is determined according to the position of the arrow, and the coordinates of the starting point and the end point of the arrow and the direction of the arrow are output. Usually, one end of the arrow is the end point and the other end is the starting point. For example, suppose the starting point coordinates of the connecting line are (x1, y1) and the end point coordinates are (x2, y2). If there is an arrow near the end point coordinates (x2, y2), the direction of the arrow is from (x1, y1) to (x2, y2). If there is an arrow near the end point coordinates (x1, y1), the direction of the arrow is from (x2, y2) to (x1, y1). It should be noted that the latent relations are represented as the connection structure between form components on the user sketch, but lack semantic depth, so it is necessary to understand the interaction between form components instead of just generating connections between form components; S2. According to the potential relationships in the form components, the logical relationships in the form components are obtained and a dynamic relationship graph is constructed, and the logical relationships in the dynamic relationship graph are analyzed using a graph neural network; Furthermore, the text position and the text content at the position are located by a text detection method, and the logical relationship in the form component is obtained according to the coordinates of the starting point and the end point of the arrow and the direction of the arrow; It needs to be explained that the logical relationships in the form component can be: data linkage, conditional jump, and process control; It should be noted that the text detection method can be implemented by using a text detection model EAST or an optical character recognition technology OCR, outputting each bounding box in the form, and obtaining the text position and text content according to the bounding box; Specifically, for the starting point of the arrow, the distance between the starting point of the arrow and all bounding boxes is calculated, and the nearest bounding box is found, which is considered to be the "source" of the arrow; for the end point of the arrow, the distance between the end point of the arrow and all bounding boxes is calculated, and the nearest bounding box is found, which is considered to be the "target" of the arrow; if there are multiple bounding boxes or arrows, there may be overlap; in this case, further judgment can be made based on the direction of the arrow and the nearest text position of the bounding box; after determining the "source" and "target", each arrow will be recorded as a "source text → target text" pair; For example, assume that the user sketch of the form design has the following elements: a drop-down box A, labeled "Country"; a drop-down box B, labeled "City"; a check box C, labeled "Agree"; a text box D, labeled "Details"; a button E, labeled "Submit"; a text box F, labeled "Server"; Arrow 1: points from drop-down box A to drop-down box B; Arrow 2: points from check box C to text box D; Arrow 3: points from button E to text box F; According to arrow 1, the selection of drop-down box A will affect the change of drop-down box B, for example, when "Country" is "China", the "City" drop-down box will display "Beijing, Shanghai", etc.; its logical relationship is expressed as data linkage; according to arrow 2, the state of check box C will select the display content of text box D, for example, if check box C is in the selected state, text box D will change from hidden state to display state for user to fill in; its logical relationship is conditional jump; according to arrow 3, the triggering of button E will control the transmission state of data; its logical relationship is process control; It should be noted that the transition from potential relationship to logical relationship gives form components semantic depth, which enables form components to have interactive logic. Further, a dynamic relationship graph is constructed, and each form component is defined as a node in the dynamic relationship graph, and the node attributes include bounding box coordinates; Furthermore, if a connecting line is detected, a directed edge is added between the two nodes. If no connecting line is detected, but the spatial distance between the two nodes is less than a preset threshold, and there is a dependency relationship between the categories, an undirected edge is added; It should be explained that since the detected connection lines are directly mapped to directed edges in the relationship graph, for connection lines with arrows, the direction of the arrow determines the direction of the directed edge; for example, if there is an arrow pointing from "country" to "city" in the sketch, then a directed edge from the "country" node to the "city" node will be added to the relationship graph, indicating that "country" directly affects "city", which is a direct relationship; if no connection line is detected, an undirected edge will be added, for example, if the "name" text box and the "address" text box are adjacent in the form design and there is no arrow connecting them, but semantically "name" and "address" are both personal information, then an undirected edge will be added between them to indicate that there is an indirect relationship between them; Specifically, the spatial distance between two nodes (two-dimensional) The calculation formula is:

[0024] in, and are represented as the horizontal coordinates of node i and node j respectively, and Represented as the ordinate (vertical) coordinates of node i and node j respectively; Specifically, the preset threshold refers to the middle value between the average distance and the maximum distance, and the distance is usually in pixels; Specifically, category dependency refers to the existence of dependency relationships between the same category or different categories. For example, there is a same category dependency relationship between drop-down boxes (the "Country" drop-down box and the "City" drop-down box), and there is a different category dependency relationship between check boxes and text boxes (the "Agree" check box and the "Details" text box). Specifically, all formed directed edges or undirected edges are connected to obtain a dynamic relationship graph; It should be noted that by constructing a dynamic relationship diagram, the user sketch can be transformed from a static interaction logic relationship to a dynamic interaction logic relationship, so that non-professional users can also quickly complete the form design; Furthermore, the logical relationship in the dynamic relationship graph is analyzed by a graph neural network, a node feature matrix is ​​created according to the dynamic relationship graph, the node features are updated by the attention mechanism of the graph neural network, and each edge where the node is located is classified, and a labeled dynamic relationship graph is output; It should be explained that the graph neural network (GNN) is a deep learning model that specializes in processing graph structure data. It can mainly make up for the shortcomings of potential relationships and infer deeper logical relationships. In the solution of the present invention, it serves as an optimization operation for the dynamic relationship graph; Specifically, the bounding box coordinates in the node attributes are normalized to the range [0,1], and the width w and height h of the bounding box coordinates are defined to obtain a 4-dimensional vector [x, y, w, h]:

[0025]

[0026]

[0027]

[0028] Among them, x min Expressed as the leftmost x-coordinate of the bounding box, indicating the left edge of the bounding box; x max Expressed as the x-coordinate of the rightmost edge of the bounding box, indicating the right edge of the bounding box; y min Represented as the y coordinate of the top of the bounding box, indicating the upper boundary of the bounding box; y max It is represented as the y-coordinate of the bottom of the bounding box, indicating the lower boundary of the bounding box; Specifically, the form component is represented as a 3-dimensional vector using one-hot encoding, and then the text content in the form component is converted into a text vector form, for example, "country" is converted into a 768-dimensional vector; Furthermore, the 4-dimensional vector, the 3-dimensional vector and the text vector are concatenated to obtain a node feature vector; Specifically, the concatenation method is a summation operation between vectors; Furthermore, all node feature vectors obtained by the above method are stacked into a node feature matrix F, which is expressed as: , where N represents the total number of nodes, D represents the feature dimension, and R is a real number; Furthermore, for each node, the attention weights of the node and its neighboring nodes are calculated. , assuming j is a neighbor node, we get:

[0029] in,

[0030] in, and Represent the current node feature vectors of node i and node j respectively, Represented as the concatenation operation of node feature vectors, Represented as the parameter vector of the attention mechanism, is the "key" in the attention mechanism, which is represented by the index of node i and provides query operations for node i; Then, update the characteristics of node i:

[0031] in, represents the updated node i; Furthermore, by updating the features of node i in the same way, we can update the neighbor node j and obtain , and then stitched again to get the stitching features , input the splicing features into a fully connected neural network in a graph neural network, the fully connected neural network assigns a corresponding logical relationship label to each edge, and outputs a labeled dynamic relationship graph; It should be noted that by outputting a dynamic relationship diagram with labels, the system can understand the overall logical structure of the form design in the user's sketch, thereby optimizing the user interaction process; S3. Optimize the user interaction path through reinforcement learning, and generate an interactive form based on the analyzed logical relationship and the optimized interaction path; It should be noted that, in order to optimize the user's filling experience in the interactive form, the Markov decision process (MDP) is used in the solution of the present invention to model the user interaction path; Specifically, the state in the Markov decision process is represented as the filling state of the form, and then the filling status of each component in the form (text box, drop-down box, check box, etc.) is abstracted into a vector; among them, unfilled: 0; filled: 1; for example, the form contains 3 components (text box: "name", drop-down box: "country", drop-down box: "city"), the state vector is: [1,0,0], indicating that the name is filled in, and the country and city are not filled in; Specifically, the actions in the Markov decision process are represented as user operation processes, for example, filling in a text box (entering "name"), selecting "country" as "China" through a drop-down box, etc.; It should be noted that the rewards are designed to encourage users to complete the form quickly and efficiently; Specifically, the reward in the Markov decision process is defined as the efficiency of the user filling out the form. The reward is divided into time efficiency and operation fluency, where time efficiency is the time to fill in the reward (i.e., the shorter the time, the higher the reward); for operation fluency, if the user's action is logical (such as filling in "name" first and then "country"), a positive reward is given; if the user's action is illogical (such as filling in contact information without filling in any personal information), a negative reward is given; It should be explained that, in order to optimize the user interaction path under the MDP framework, the proximal policy optimization (PPO) algorithm is adopted in the solution of the present invention. The PPO algorithm is an efficient reinforcement learning method that can learn the optimal strategy while maintaining the stability of the strategy; Specifically, the PPO algorithm learns through a policy network and a value network. The policy network outputs the probability distribution of actions based on the state in the Markov decision process, for example, the probability of "filling in the name" is 0.6, and the probability of "selecting a country" is 0.4. The value network estimates the value of each state, which is used to evaluate the long-term return of taking actions in the current state. Specifically, by collecting the interaction data between users and forms, calculating the generalized advantage function, and updating the parameters of the above strategy network and value network, the optimal user interaction path is gradually learned; It should be explained that the advantage function is the performance of a specific action in a certain state compared to the average situation, which is usually defined as A(s, a) = Q(s, a) - V(s), where s represents the current state (such as the filling state of a form), a represents the action taken (something filled in the form), and V(s) represents the expected cumulative return value of following the current strategy in state s; if A(s, a)>0, it means that action a is better than the average in state s, and the strategy needs to increase the probability of selecting this action; if A(s, a)<0, it means that action a is better than the average in state s, and the strategy needs to reduce the probability of selecting this action; if A(s, a)=0, it means that action a is the same as the average in state s, and the strategy does not need to change; Specifically, the generalized advantage function It is expressed as:

[0032] in, It represents the weighted contribution value of all future time steps (infinite steps) starting from the current moment. In fact, due to the discount factor and smoothing parameters The existence of will gradually reduce the influence of the time step, so the summation can be approximated as a finite step; Expressed as the timing difference error, where , represents the timing difference error (such as the efficiency of filling), Represents the value estimate of the state, obtained from the output of the value network; Expressed as the discounted value of the next state; It should be noted that by continuously adjusting the generalized advantage function, the policy network continuously outputs the probability distribution of high-advantage actions, thereby achieving the effect of gradually optimizing the user interaction path; Specifically, according to the optimal user interaction path gradually learned, the edge relationships in the labeled dynamic relationship graph are converted into JavaScript rules and embedded into static elements rendered based on HTML / CSS to generate an interactive form. It should be noted that the JavaScript rules are not written manually by the user, but are automatically generated by the system based on the output of the labeled dynamic relationship diagram. <script>标签将该JavaScript规则嵌入到HTML / CSS渲染的静态元素中,实现表单页面前端的交互功能。

[0033] 本领域内的技术人员应明白,本发明实施例可提供为方法、系统、或计算机程序产品。因此,本申请可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本申请可采用在一个或多个其中包含有计算机可用程序代码的计算机可用存储介质(包括但不限于磁盘存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。本申请实施例中的方案可以采用各种计算机语言实现,例如,面向对象的程序设计语言Java和直译式脚本语言JavaScript等。

[0034] 本申请是参照根据本申请实施例的方法、设备(系统)、和计算机程序产品的流程图和/或方框图来描述的。应理解可由计算机程序指令实现流程图和/或方框图中的每一流程和/或方框、以及流程图和/或方框图中的流程和/或方框的结合。可提供这些计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理设备的处理器以产生一个机器,使得通过计算机或其他可编程数据处理设备的处理器执行的指令产生用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的装置。

[0035] 这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理设备以特定方式工作的计算机可读存储器中,使得存储在该计算机可读存储器中的指令产生包括指令装置的制造品,该指令装置实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能。

[0036] 这些计算机程序指令也可装载到计算机或其他可编程数据处理设备上,使得在计算机或其他可编程设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程设备上执行的指令提供用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的步骤。

[0037] 尽管已描述了本申请的优选实施例,但本领域内的技术人员一旦得知了基本创造性概念,则可对这些实施例作出另外的变更和修改。所以,所附权利要求意欲解释为包括优选实施例以及落入本申请范围的所有变更和修改。

[0038] 显然,本领域的技术人员可以对本申请进行各种改动和变型而不脱离本申请的精神和范围。这样,倘若本申请的这些修改和变型属于本申请权利要求及其等同技术的范围之内,则本申请也意图包含这些改动和变型在内。< / script>

Claims

1. A method for automatically generating zero-code forms based on AI image recognition, characterized in that: include: Receiving a sketch or screenshot provided by a user, performing image preprocessing on the sketch or screenshot, and using an object detection model to identify form components in the preprocessed image, and extracting potential relationships among the form components; According to the potential relationships in the form components, the logical relationships in the form components are obtained and a dynamic relationship graph is constructed, and the logical relationships in the dynamic relationship graph are analyzed using a graph neural network; The user interaction path is optimized through reinforcement learning, and an interactive form is generated based on the logical relationship obtained through analysis and the optimized interaction path.

2. The method for automatically generating a zero-code form based on AI image recognition according to claim 1, characterized in that: Performing image preprocessing on the sketch or screenshot, including: The sketch or screenshot is regarded as an image, image noise is suppressed by Gaussian filtering, image edges are detected by a simplified Canny algorithm, and the image is binarized by the Otsu method.

3. The method for automatically generating a zero-code form based on AI image recognition according to claim 1, characterized in that: Use object detection models to identify form components in preprocessed images, including: Taking YOLOv5 as the object detection model, the graphic image annotation tool labelme is used to generate a set of labeled images of form components. The preprocessed image is taken as input, and the form components in the preprocessed image are identified according to the set of labeled images of the form components, and the bounding box coordinates of each form component are output.

4. The method for automatically generating a zero-code form based on AI image recognition as claimed in claim 3, characterized in that: Extract potential relationships in form components, including: Based on the Hough transform, the connecting lines in the form are detected. By setting the minimum line length and maximum gap and analyzing the pixel distribution near the end points of the connecting line, it is determined whether the connecting line is an arrow. If it is an arrow, the coordinates of the starting point and end point of the arrow and the direction of the arrow are output.

5. The method for automatically generating a zero-code form based on AI image recognition as claimed in claim 3 or 4, characterized in that: According to the potential relationships in the form components, the logical relationships in the form components are obtained and a dynamic relationship diagram is constructed, including: Locate the text position and the text content at the position by using a text detection method, and obtain the logical relationship in the form component according to the coordinates of the starting point and end point of the arrow and the direction of the arrow; Each form component is defined as a node in a dynamic relationship graph, and the attributes of the node include bounding box coordinates; If a connection line is detected, a directed edge is added between the two nodes. If no connection line is detected, but the spatial distance between the two nodes is less than the preset threshold and there is a dependency relationship between the categories, an undirected edge is added; All the formed directed edges or undirected edges are connected to obtain a dynamic relationship graph.

6. The method for automatically generating a zero-code form based on AI image recognition as claimed in claim 5, characterized in that: Analyzing the logical relationship in the dynamic relationship graph using a graph neural network includes: According to the dynamic relationship graph, a node feature matrix is ​​created; The node features are updated through the attention mechanism of the graph neural network, and each edge where the node is located is classified to output a dynamic relationship graph with labels.

7. The method for automatically generating a zero-code form based on AI image recognition according to claim 6, characterized in that: Optimize user interaction paths through reinforcement learning, including: The user interaction path is modeled through the Markov decision process. The state in the Markov decision process is represented as the form filling state, the action is represented as the user operation process, and the reward is defined as the efficiency of the user filling in the form.

8. The method for automatically generating a zero-code form based on AI image recognition as claimed in claim 7, characterized in that: Also includes: A proximal strategy optimization algorithm is used in the Markov decision process to update the user interaction path in real time according to the state and action.

9. The method for automatically generating a zero-code form based on AI image recognition according to claim 6, characterized in that: Generate an interactive form based on the analyzed logical relationship and optimized interaction path, including: The edge relationships in the labeled dynamic relationship graph are converted into JavaScript rules and embedded into static elements rendered based on HTML / CSS to generate interactive forms.

Citation Information

Patent Citations

  • Cross-layer routing optimization method and device, equipment and storage medium

    CN116566884A

  • PID drawing identification and reconstruction system based on end-to-end deep learning

    CN117373051A

  • Multi-behavior sequence recommendation method and system based on graph neural network

    CN119493908A

  • Power grid line loss optimization method and system based on graph attention perception and reinforcement learning decision

    CN119513703A

  • Handwritten Diagram Recognition Using Deep Learning Models

    US20210073530A1

Cited By

  • Zero-code multi-terminal application automatic construction method based on AI semantic understanding

    CN120144107A

  • Flow chart image analysis and structured reconstruction method and device, and storage medium

    CN120808375A