Interface design specification evaluation-oriented image-text large model fine tuning method and related device
By constructing a comprehensive picture-text data set and training a picture-text model, the problems of low efficiency and strong subjectivity of software interface design specification evaluation are solved, and intelligent interface design specification evaluation is realized, which improves the accuracy and user experience of evaluation.
Patent Information
- Application Number
- CN202411975468.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The prior art is inefficient and subjective in the evaluation of software human-computer interactive interface design specifications, making it difficult to effectively evaluate whether the software interface complies with the general design specifications in the industry.
By constructing a comprehensive picture-text data set for software interface design specification evaluation, including the first, second and third type picture-text data sets, it is used to improve the ability of the general picture-text model to understand, describe and evaluate the software interface content, and combine the fine-tuning strategy to train the picture-text model to establish a picture-text model for interface design specification evaluation.
It realizes intelligent evaluation of static design specifications of software interfaces, improves the accuracy and efficiency of evaluation, provides intelligent assisted decision-making, promotes the optimization of software interface design specifications, and improves user interaction experience.
Smart Images

Figure CN119917095A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to a large-scale graphic model fine-tuning method and related devices for interface design specification evaluation. Background Art
[0002] With the continuous advancement of technology and the deepening of research, the design and evaluation of software human-computer interaction interfaces has become an important research direction in the field of human-computer interaction. The software human-computer interaction interface intuitively and effectively displays the software's operating interface through the reasonable layout and design of graphical interface elements, while simplifying the user operation process and improving the software's usability and user satisfaction. However, the differences in functions, professional fields, target user groups, and designers' aesthetic preferences of different software interfaces have led to the diversity of software interfaces. Whether its design follows the general design specifications in the industry can be determined through specific analysis and evaluation. However, for a long time, the evaluation of software human-computer interaction interface design specifications has mostly relied on manual methods. In terms of compliance inspection of various design specifications, software designers need to always maintain a rigorous and serious attitude during the design process. After the design is completed, each page is comprehensively checked with reference to the given software interface design requirements, which often has problems of low efficiency and strong subjectivity.
[0003] For the static interface of software, in the past, machine vision methods were used to fully mine and explore the elements in the interface, and then analyze the recognition results. With the rapid development of large language models, similar large image and text models have also entered the public eye. As a multimodal large model, it can combine image and text information to comprehensively process problems. The large image and text model further integrates the ability of natural language processing on the basis of machine vision. This fusion enables the machine to not only "see" the content in the image, but also "understand" the relationship between the image content and the natural language description.
[0004] The rapid development of graphic and text big models has brought revolutionary progress to many fields, but there is still little research on the specific application of software interface design specification evaluation. At the same time, most open source graphic and text big models are general graphic and text big models, and their performance in software human-computer interaction interface design specification evaluation is average. Summary of the invention
[0005] The purpose of this application is to provide a large graphic model fine-tuning method and related devices for interface design specification evaluation, which can improve the accuracy of software interface design specification evaluation.
[0006] To achieve the above objectives, this application provides the following solutions:
[0007] In a first aspect, the present application provides a method for fine-tuning a large graphic model for interface design specification evaluation, including:
[0008] Construct a comprehensive image-text pair dataset for software interface design specification evaluation, wherein the comprehensive image-text pair dataset includes a first type of image-text pair dataset, a second type of image-text pair dataset, and a third type of image-text pair dataset. The first type of image-text pair dataset is used to improve the ability of the general image-text large model to understand the software interface content. The second type of image-text pair dataset is used to improve the ability of the general image-text large model to describe the interface content. The third type of image-text pair dataset is used to improve the ability of the general image-text large model to evaluate the software interface design specification. The general image-text large model is an existing pre-trained model. The first type of image-text pair dataset, the second type of image-text pair dataset, and the third type of image-text pair dataset all include software interface pictures, prompt words related to the software interface, and standard answers. The standard answer refers to an evaluation text generated according to a list of general design specifications for software interfaces. The list of general design specifications for software interfaces includes comprehensive interface design requirements and specific interface design specifications.
[0009] The software interface image and the prompt word are used as feature data, the standard answer is used as label data, and the general image-text model is trained in combination with a fine-tuning strategy to obtain a trained image-text model for interface design specification evaluation.
[0010] Optionally, the specific construction process of the first type of image-text pair data set includes:
[0011] The specific construction process of the first type of image-text pair dataset includes:
[0012] Obtain various types of question-answer data pairs about software interface images from existing public question-answer data sets, wherein each of the question-answer data pairs includes a question about the software interface image and a corresponding answer;
[0013] According to the research task requirements and the software interface general design specification list, the question-answer data pairs about the software interface pictures are screened to obtain screened question-answer data pairs about the software interface pictures;
[0014] Fine-tune the required data format according to the visual instruction, and convert the format of the question-and-answer data pair about the software interface image to obtain the question-and-answer data about the software interface image after format conversion;
[0015] The first type of picture-text pair data set is obtained based on the question-and-answer data about the software interface pictures after conversion into the format.
[0016] Optionally, the specific construction process of the second type of image-text pair data set includes:
[0017] Acquire multiple effective expression forms of software interface images from existing public software interface data sets, wherein the effective expression forms refer to the expression forms of software interfaces that can improve the accuracy, comprehensiveness and richness of software interface descriptions;
[0018] Encoding the file data corresponding to the effective expression form to obtain encoded file data, wherein the file data includes a software interface image and corresponding semantic information;
[0019] Generate a software evaluation text according to the encoded file data and the prompt words designed by the prompt engineering;
[0020] Performing format alignment processing on the software evaluation text and the software interface image in the file data to obtain an aligned software evaluation text and an aligned software interface image;
[0021] The second type of image-text pair data set is obtained according to the aligned software evaluation text, the aligned software interface picture and the prompt words.
[0022] Optionally, the third type of image-text pair data set includes a standard image-text pair data set and an enhanced image-text pair data set, and the specific construction process of the enhanced image-text pair data set includes:
[0023] According to the software interface general design specification list, construct a standard software interface design theme template;
[0024] Randomly modifying the interface elements and / or design layout in the standard software interface design theme template to generate software interface sample image data that violates the software interface general design specification list;
[0025] Generate text evaluation information based on the randomly modified information;
[0026] Performing format alignment processing on the software interface sample image data and the text evaluation information to obtain aligned software interface sample image data and aligned text evaluation information;
[0027] The enhanced image-text pair data set is obtained according to the aligned software interface sample image data, prompt words and the aligned text evaluation information, wherein the prompt words are questions about the aligned software interface sample image data obtained by prompt engineering design.
[0028] Optionally, the comprehensive interface design requirements include interface consistency, design simplicity, intuitiveness principle, accessibility requirements and compliance with user habits; the specific interface design specifications include bright element colors, clear text meaning, standardization of interface elements, identification of interactive elements, unified text fonts, reasonable element layout, complete text display and coordinated font colors.
[0029] Optionally, the fine-tuning strategy includes: using the DeepSpeed deep learning optimization library, using the LoRA strategy, using a general graphic model of a preset scale, using a training checkpoint saving function, using the Lazy Preprocess data preprocessing strategy, and using Wand for fine-tuning records.
[0030] Optionally, after executing the step of "using the software interface image and the prompt word as feature data, using the standard answer as label data, and combining the fine-tuning strategy to train the general image-text model to obtain a trained image-text model for interface design specification evaluation", the image-text model fine-tuning method for interface design specification evaluation further includes:
[0031] Get relevant interface pictures of the software to be tested;
[0032] Designing test language instructions according to the software interface general design specification list;
[0033] Taking the relevant interface pictures and the test language instructions as input, outputting answers using the trained large picture and text model for interface design specification evaluation;
[0034] Calculate the score based on the number of problems and the severity of violations in each of the relevant interface images determined by the answers, where the number of problems refers to the number of interface elements that violate the software interface general design specification list, and the severity of violations refers to the degree of violation of the software interface general design specification list.
[0035] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the large graphic and text model fine-tuning method for interface design specification evaluation as described in any one of the first aspects above.
[0036] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for fine-tuning a large graphic and text model for interface design specification evaluation as described in any one of the first aspects above.
[0037] In a fourth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the large graphic and text model fine-tuning method for interface design specification evaluation described in any one of the first aspects above.
[0038] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0039] The present application provides a method and related devices for fine-tuning a large graphic model for interface design specification evaluation. By studying existing software interface design specifications, the present application formulates a software interface design specification list for the graphic large model, and based on this, constructs a comprehensive graphic data set for the interface design specification evaluation task. By fine-tuning the general graphic large model, a graphic large model for software interface design specification evaluation is established, thereby realizing intelligent evaluation of static design specifications of software interfaces. The intelligent evaluation results output by the model can provide intelligent auxiliary decision-making for software interface developers, promote the optimization of software interface design specifications, and enhance user interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0041] Figure 1 A schematic diagram of a flow chart of a method for fine-tuning a large graphic model for interface design specification evaluation provided in Example 1 of the present application;
[0042] Figure 2 This is a schematic diagram of the construction process of the first type of image-text pair data set in Example 1 of the present application;
[0043] Figure 3 This is a schematic diagram of the construction process of the second type of image-text pair data set in Example 1 of the present application;
[0044] Figure 4 This is a schematic diagram of the construction process of the third type of image-text pair data set in Example 1 of the present application. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0046] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0047] Example 1
[0048] like Figure 1As shown, the embodiment of the present application provides a method for fine-tuning a large graphic model for interface design specification evaluation, which is executed by a computer device, specifically, it can be executed by a computer device such as a terminal or a server alone, or it can be executed by a terminal and a server together. The method is applied to a server as an example for description, and includes the following steps 101 to 108. Among them:
[0049] Step 101, constructing a comprehensive image-text pair dataset for evaluating software interface design specifications, wherein the comprehensive image-text pair dataset includes a first type of image-text pair dataset, a second type of image-text pair dataset, and a third type of image-text pair dataset. The first type of image-text pair dataset is used to improve the ability of the general image-text large model to understand software interface content, the second type of image-text pair dataset is used to improve the ability of the general image-text large model to describe interface content, and the third type of image-text pair dataset is used to improve the ability of the general image-text large model to evaluate software interface design specifications. The general image-text large model is an existing pre-trained model. The first type of image-text pair dataset, the second type of image-text pair dataset, and the third type of image-text pair dataset all include software interface pictures, prompt words related to the software interface, and standard answers. The standard answers refer to evaluation texts generated according to a list of general design specifications for software interfaces. The list of general design specifications for software interfaces includes comprehensive interface design requirements and specific interface design specifications.
[0050] Step 102, taking the software interface image and the prompt word as feature data, taking the standard answer as label data, and combining the fine-tuning strategy to train the general image-text model, to obtain a trained image-text model for interface design specification evaluation.
[0051] In order to make the specific implementation process of the above steps 101 to 102 more clear to those skilled in the art, the following Figure 2-Figure 4 Give a detailed explanation.
[0052] Step 1: Create a software interface design specification list for large graphic models.
[0053] Before executing step 101, this embodiment, in response to the current situation of lack of systematic and comprehensive general design specifications for software static interfaces, formulates design specification selection principles and conversion methods through literature research and case analysis, and obtains a software static interface design specification list (i.e., software interface design specification list) for large graphic models.
[0054] This embodiment conducts extensive research on software static interface design specifications from both theoretical and practical aspects, and formulates a comprehensive list of software interface design specifications for graphic and text models by selecting principles and transformation methods, including comprehensive interface design requirements and specific interface design specifications, the form and content of which are shown in Table 1 and Table 2 respectively. The specific research steps are as follows:
[0055] Step 1.1: Literature research and extraction. Through extensive reading of relevant books and materials, extract the systematic general software design specifications, investigate the design manuals written by software developers in specific professional fields, and explore the design concepts mentioned in detail.
[0056] Step 1.2: Case analysis and summary. In-depth analysis of widely used professional software on the market, summarize its design features, extensively research the recognized best practices and guidelines in many industries, combine user evaluation and feedback, and extract key design specifications from the perspective of user experience.
[0057] Step 1.3: Standardization and transformation. To ensure the practicality, operability and consistency of the specifications and transform them into specifications that are easy to understand and model learning, a series of clear selection principles and transformation methods are followed when organizing the general design specifications of the static interface of the software.
[0058] Based on the characteristics of large-scale image and text model recognition and analysis, the interface design principles are screened according to the characteristic attributes of the software static interface, and the selection principles when obtaining design requirements mainly include the following four points:
[0059] (1) Model differentiability.
[0060] Choose specifications that are easy for the model to recognize and analyze. Objects in the specification should be large enough, clearly distinguishable, and conform to the basic style of the element type. This helps the model recognize interface elements and layout features more accurately.
[0061] (2) Clarity of examples.
[0062] Prioritize specifications that clearly define and design positive and negative examples. Such standards will facilitate the construction of data sets with software interface data specification graphics and text, improve the fine-tuning effect and verification efficiency of the model. Through example clarity, designers can quickly understand and apply specifications, thereby avoiding common design errors.
[0063] (3) Define operability.
[0064] Choose rules that are easy to define and quantify. The specification should clearly define the criteria and examples for adjectives such as "fancy" or "simple", such as the "uniform text font" rule, which requires that simplicity be improved by avoiding the mixing of multiple font styles. It should be specific to using the same font for text in the same block, without abrupt bold and italic interfaces, so as to improve the enforceability of the specification.
[0065] (4) Static identifiability.
[0066] The selected specifications must be able to be judged in a static interface. Since the research object and subsequent research methods are all static software interfaces, the plans that require dynamic feedback cannot be verified and therefore need to be eliminated.
[0067] Based on the existing interface design requirements, the industry often has some corresponding design specifications to meet the requirements. The regulations that are not universal or other regulations that are not suitable for analysis of large graphic models need to be modified and converted. The conversion methods of the regulations mainly include the following three techniques:
[0068] (1) Double-sided guiding principle.
[0069] The specification is expressed as a combination of positive and negative guiding principles, indicating both how the design should be done and how it should not be done. This expression method makes it easier to judge and divide the actual situation and improves the efficiency of inspection.
[0070] (2) Conditional result statement.
[0071] Use conditional result statements to describe specifications, that is, if the interface design meets certain conditions, it can be considered to follow the corresponding design specifications or violate a certain design specification. This method provides clear judgment criteria and can avoid vague statements affecting the implementation of the inspection.
[0072] (3) Instantiation specifications.
[0073] Provide practical application examples for each specification, which can be cases of design success or failure. Use some simple examples to help inspectors or models understand the meaning of the expression.
[0074] Table 1 Interface comprehensive design requirements
[0075]
[0076]
[0077] Table 2 Specific interface design specifications
[0078]
[0079] Step 2: Training of large graphic and text models for interface design specification evaluation.
[0080] Step 2 includes three steps: construction of fine-tuning dataset, implementation of fine-tuning strategy and benchmark verification. In terms of data construction, this embodiment proposes three methods for constructing image-text datasets for software interface design specification evaluation: a public QA dataset screening and format alignment method (i.e., the construction process of the first type of image-text dataset), a multi-type interface data text evaluation matching method (i.e., the construction process of the second type of image-text dataset), and a software interface data enhancement and structured evaluation text method for automatically constructing datasets (i.e., the construction process of the third type of image-text dataset). In terms of fine-tuning strategy, this embodiment studies the classic architecture of large image-text models, uses a variety of fine-tuning techniques and adjusts the hyperparameter settings to balance fine-tuning efficiency and fine-tuning quality. In terms of model verification, this embodiment proposes a benchmarking method for software interface design specification evaluation for large image-text models, constructs three types of benchmark questions and answers, and compares the scores of model results through GPT comparison.
[0081] The software interface design specification list established in step 1 above provides theoretical guidance for fine-tuning the general graphic model. Specifically, the software interface design specification list in step 1 is applied to step 2 by guiding the direction of data set construction and providing model verification standards, ensuring that the graphic model can effectively evaluate whether the software interface violates the software interface design specification list.
[0082] For the structure of the large model, enter the picture X v Transformed into vector Z through image encoder g v ; Language instruction X q Converted to embedding through a text encoder, denoted as H q ; W is a pre-trained image-text mapper (by inputting a large number of matching image-text pairs to build the relationship between text and corresponding images, such as a picture of a kitten can correspond to the text "kitten"), the purpose of which is to convert the image vector Z v Align with the corresponding text vector in the vector space and transform it into an embedding, denoted as H v ; Finally, H q ,H v All input to the large language model f φ Get answer X a In this task, the answers obtained here are for different types of questions in a single picture and prompt word. In order to obtain the final evaluation, it is necessary to perform scoring statistics. The specific process is shown in step 2.3. The fine-tuning of the large picture-text model in this embodiment mainly adjusts the picture-text mapper g and the large language model f. φ The fine-tuning steps are mainly divided into three steps: construction of fine-tuning dataset, implementation of fine-tuning strategy and benchmark verification.
[0083] Step 2.1: Software interface design specification image-text pair data set construction (i.e., construction of comprehensive image-text pair data set).
[0084] In order to improve the ability of the large image and text model in software interface evaluation tasks, the large image and text model is fine-tuned using the software interface design specification image and text dataset. A single fine-tuning training dataset should include three parts:
[0085] Software interface images: Provide visual information that the model needs to understand and analyze.
[0086] Prompt words (i.e. language instructions): text inputs that guide the model to generate specific outputs, questions or instructions related to the software interface.
[0087] Standard answer to the corresponding question: The model should learn to generate the target output, which is the accurate response to the prompt word.
[0088] This embodiment further proposes three methods for constructing interface design specification image-text pair data sets. Each method is based on the software interface design specification list formulated in step 1, but constructs its own data set for different capabilities of the model. Among them, the third type of image-text pair data set focuses on the main task type that needs to be completed in the application process of the large image-text model - identification and evaluation of interface defects. The first two types of data sets are used as auxiliary, and are constructed for interface content understanding and interface description capabilities respectively, because improving the capabilities of these two types of tasks is significantly helpful for improving the identification and evaluation capabilities of interface defects. The construction flow charts of the three types of software interface design specification image-text pair data are shown respectively. Figure 2 , Figure 3 and Figure 4 .
[0089] Method 1: Construction of the first type of image-text pair dataset - screening and format alignment of public QA datasets.
[0090] a) Obtaining question-answer data pairs about software interfaces from existing public question-answer datasets. This embodiment extracts various types of question-answer data pairs about software interfaces (such as extraction of software interface text content, understanding of digital meaning, etc.) from the public software interface dataset ScreenQA (a dataset containing three types of large software interface question-answer pairs and divided into datasets).
[0091] b) According to this research task (i.e., the research task requirements) and the list of general software interface design specifications formulated in step 1, the question-answer data pairs obtained in step a are screened by considering relevance, balance, and quality to obtain screened question-answer data pairs. For example, since this research task requires the evaluation of the interface content without indicating the spatial location of the evaluation content, the location coordinates and other information originally included in the answer to the question are removed for relevance considerations; for balance considerations, for question-answer pairs with multiple different expressions for the same question in the ScreenQA dataset, only one is randomly selected as the answer to avoid overfitting, etc.
[0092] c) According to the data style after data screening and the data format required for visual instruction fine-tuning, write a Python program to convert the screened question-answer data pairs into a unified format, and organize the software interface pictures and texts, questions about the software interface, and answers obtained through step b into the image input, prompt words, and answers in the visual instruction fine-tuning data. Among them, the prompt words contain the questions of the question-answer pair and some basic prompt engineering skills, such as "If you are an expert in software interface evaluation, please observe the software interface and answer the question <the question part of the question-answer pair> after calm thinking", and the answer part is the answer in the selected question-answer pair; for each set of questions and answers, a globally unique id for visual instruction fine-tuning is given, and it corresponds to the id of the picture in the original dataset ScreenQA and the software interface picture path. This part containing id, prompt words, and answers is the text part of the picture-text pair dataset used for fine-tuning, and the picture part refers to the picture pointed to by the software interface picture path.
[0093] d) Based on the format-converted question-answer data, a picture-text pair dataset is obtained to improve the content comprehension capability of the picture-text large model interface, namely, the first type of picture-text pair dataset.
[0094] The first type of image-text dataset provides a large number of question-answering examples, forcing the model to deeply understand the details and context in the image during the learning process. This involves not only the recognition of individual interface elements, but also the understanding of their functions and relationships in the overall layout. The model improves its ability to understand and parse complex interface content by optimizing its internal representation so that visual features and text features are better aligned in the multimodal space.
[0095] For example: id:"5-1-short-answers-test"; image:"RICO / 5.jpg" (this is a path to an image, which is a screenshot of a UI interface); question (from the user): "How many exercises are there in total on the interface"; answer: "12 exercises".
[0096] Among them, id is a unique identifier, the image is performed using a file path, and the conversation consists of two parts: questions and answers. It is extracted based on the public dataset mentioned in the first dataset construction method.
[0097] The first type of image-text pair dataset is mainly constructed based on question-answering datasets, especially question-answering pairs for software interface content understanding. This type of dataset contains images and corresponding question-answering pairs. The purpose is to help the model understand the interface elements and their functions in the image by providing a large number of question-answering examples.
[0098] Images: Screenshots or graphical representations of software interfaces (e.g. buttons, text boxes, icons, etc.).
[0099] Prompt words: manually designed basic understanding questions about the interface content, such as the number and type of numbers, texts or controls displayed on the interface.
[0100] Evaluation text: A simple description or query of the interface content.
[0101] The first type of image-text pair dataset helps the model understand the basic meaning and function of interface elements. Through direct question-answering, the model learns how to extract specific information from images (such as the number of buttons, the content of text boxes, etc.). This training enables the model to better identify visual elements in the interface and understand the basic functions and positional relationships of these elements.
[0102] Method 2: Construction of the second type of image-text pair dataset - multi-type interface data text evaluation matching.
[0103] aa) Obtain multiple effective expressions of software interfaces from existing public software interface datasets, wherein the effective expressions refer to expressions of software interfaces that can improve the accuracy, comprehensiveness, and richness of software interface descriptions. This embodiment extracts multiple expressions of software interfaces from the public software interface dataset Enrico (a dataset containing a large number of software interfaces in multiple representations. For the same software interface, its expressions include software interface screenshots, semantic wireframes, semantic annotations, DOM-like tree structures, application metadata, etc.).
[0104] bb) Screening multiple representation methods of software interfaces that are helpful for improving the accuracy, comprehensiveness and richness of describing software interfaces and for evaluating software design specifications, such as UI screenshots provided in JPG format with high resolution of 540x960 pixels to provide direct data for visual analysis; semantic wireframes in PNG format with the same resolution to provide auxiliary reference for layout analysis of software interfaces; semantic annotations provided in JSON format to provide richer GUI information.
[0105] cc) For the various software interface expressions selected in step bb, the file data (i.e., JPG screenshots, PNG semantic wireframes, and JSON semantic annotation files) are encoded and passed to ChatGPTAPI together with the prompt words designed by the prompt project (a project that uses a variety of prompt word design techniques to improve the performance of the output content of the large model so as to meet human expectations) to generate software evaluation text. This method uses the formatted output, role prompts, providing thinking time, task splitting and other techniques of the prompt project to guide the large language model to generate the expected output based on the above materials provided, including software interface description, software interface general design specification list and specification compliance.
[0106] dd) Performing format alignment processing on the software interface screenshots selected in step bb and the software evaluation text generated in step cc, and obtaining a picture-text pair dataset for improving the interface description capability of the picture-text large model according to the formatted software interface screenshots, software evaluation text and prompt words, i.e., the second type of picture-text pair dataset.
[0107] The second type of image-text pair dataset requires the model to generate detailed natural language descriptions, which requires the model to integrate visual information and language generation mechanisms at a higher level in its internal representation. Through this training, the model not only learns to recognize interface elements, but also learns how to organize these elements into coherent and logical language descriptions, thereby improving its interface description capabilities.
[0108] For example: id:"04221207_0523093923242023"; image:"test / augmentation / 04221207_0523093923242023.png"; Question (from user): "Please evaluate the interface design based on the provided software interface image to determine whether it follows the general design standards for software. List each UI control with design issues. Describe the design issues of each control, including specific issues related to geometric positioning, visual elements, and user interaction. For each issue, provide a description of the general software design standards that are violated"; Answer: Element: <element name> Problem: <problem attribute type> Type: <specific problem type>; Description: <specific problem description and impact>; Violation standard: <specific violated general design standard>.
[0109] The question part is through the prompt words of the engineering design, and the answer part here is the result returned by calling the gpt api mentioned in the second data set construction process.
[0110] The second type of image-text dataset requires the model to generate a detailed natural language description based on a given software interface image. This dataset focuses on improving the model's ability to convert interface elements and layouts into coherent and reasonable language descriptions.
[0111] Image: A detailed screenshot of the software interface, showing the various controls and layout of the interface.
[0112] Hint: Complex descriptive questions require the model to give a detailed interface description or analysis. For example, the model needs to describe whether the interface design meets common design standards, and whether there are problems with element position, fonts, colors, etc.
[0113] Evaluation text: The model is required to output a detailed and structured interface analysis based on visual information. For example, for a description of a button position, it can point out the button size problem, asymmetric position, etc., and even give the design specifications that are violated.
[0114] The second type of image-text dataset allows the model to learn how to convert the visual information of interface elements into text and evaluate it from a design perspective. Through this type of training, the model not only understands the interface content, but also generates a coherent and in-depth description of the relationship between interface elements and the normativeness of the design.
[0115] Method 3: Software interface data enhancement and structured evaluation text.
[0116] aaa) According to the software interface design specification list formulated in step 1, some standard design theme templates are manually constructed. First, the Qt Designer component of PyQt 5 is used to form as rich software elements as possible through simple drag and drop operations and property settings, so as to draw a simple form interface, notification window, etc. Then, since what is needed is a screenshot of the static interface of the software, there is no need to design dynamic properties such as the interactive actions of the controls, so the ui file format is directly selected for saving, which is a special form of the xml format used to describe the layout and properties of the user interface. Finally, a simple rendering code can be written in Python to obtain the corresponding static interface of the software.
[0117] bbb) According to the regulations on the software interface design specification list, the interface elements and / or design layout in the standard software interface design theme template are randomly modified, and the corresponding data enhancement methods for generating negative samples are designed one by one, so as to obtain a large amount of software interface image data that violates the corresponding specifications. First, in Python, the xml parsing library xml.etree.ElementTree is used to read the properties of each control in the ui file and create an entity object to save information, which facilitates the subsequent modification of element attributes. Here, the main attributes such as geometric information, text content, font and color are selected for method research and design of design specification negative samples. Secondly, based on the general design specifications obtained from the survey, the above attributes are randomly modified, such as random offset of geometric positions, random replacement of text content with similar meanings but unclear expressions, etc., and negative samples that violate the corresponding general design specifications after certain attributes of certain elements of the interface are changed are automatically generated.
[0118] ccc) records the information of random changes in step bbb, generates formatted evaluation text according to the element attribute change information, and obtains text evaluation data corresponding to the interface image that violates the specification. For example, in the judgment of geometric position change, it is considered that when the amount of change in the horizontal or vertical direction exceeds a certain threshold, the effect description of this change is that the horizontal or vertical direction is greatly offset from other controls. This large offset may affect the user's positioning and identification of interface elements and violate the pre-selected general design specification of user interface layout continuity. Similar judgments and evaluation text generation are performed for all change information, and finally organized into a formatted evaluation text Json file of the control name and problem list (element attributes, attribute descriptions, and violations of general design specifications).
[0119] ddd) Perform format alignment processing on the software interface image generated in step bbb and the evaluation text generated in step ccc, and combine with voice commands to obtain a picture-text pair dataset for improving the evaluation ability of the picture-text large model for a given interface design specification list, that is, the third type of picture-text pair dataset.
[0120] The third type of image-text pair dataset further improves the model's ability to understand and apply design specifications. The model needs to learn the details of the design specifications and be able to identify the parts that do not meet the specifications in the actual interface. This requires the model to embed knowledge of design standards in its internal representation and identify specific design defects through fine-grained feature analysis capabilities. This involves not only the detection of visual elements, but also the analysis of their geometric position, color contrast, layout rationality, and many other aspects.
[0121] For example: id:"04221207_0510150331753286";
[0122] image:"ui_data / augmentation / 04221207_0510150331753286.png"; Question (from user): "Please evaluate the interface design based on the provided software interface image to determine whether it follows the general design standards for software, and provide issues according to the following JSON structure: List each UI control with design issues. Describe the design issues of each control, including specific issues related to geometric positioning, visual elements, and user interaction. For each issue, provide a description of the general software design standards that are violated. The answer should be formatted according to the following JSON object; if there are multiple issues, please provide a separate entry for each issue: {\"control\":\"<control name>\",\"issues\":{<issue attribute type>:\"Type:<specific issue type>; Description:<specific issue description and impact>; Violated Standard:<Specific violation of the general design standard>\"; Answer: {\"element\":\"<Confirm password>QLabel\",\"issues\":{\"font\":\"Type: font family, font is 'CourierNew'; Description: Non-uniform fonts may be inconsistent with other fonts in the interface, affecting user experience; Violated Standard: interface consistency. \"}.
[0123] The third type of image-text dataset aims to help the model deeply understand the software interface design specifications and identify the parts of the interface that violate the design specifications. This dataset contains a detailed list of specifications, and the model needs to evaluate the interface based on these specifications and point out design defects and non-compliance with the specifications.
[0124] Images: They are still screenshots of software interfaces, but these interfaces have been “enhanced” in a certain way, that is, the style deliberately violates the design specifications (such as irregular element positions or inconsistent fonts).
[0125] Prompt words: Require the model to evaluate whether the design in the image complies with predefined design specifications, point out the illegal designs, and generate formatted evaluation text.
[0126] Comment text: Describe the specific design flaw and give the design standard that was violated. For example, "The button sizes are inconsistent, which makes the interface cluttered and violates the interface consistency standard."
[0127] The third type of image-text pair dataset focuses on the model’s deep understanding of design specifications. Through this type of data, the model learns how to analyze and evaluate the interface according to specific design standards, can identify defects in the design, and give reasonable and standardized evaluations. This is crucial for the model to be able to automatically evaluate interface design.
[0128] The first type of image-text pair dataset: improves the model’s basic understanding of software interface content (such as the number of elements, functions, etc.).
[0129] The second type of image-text dataset: improves the model's ability to generate detailed descriptions of interfaces, especially descriptions of the layout and design between elements.
[0130] The third type of image-text pair dataset: improves the model’s understanding of interface design specifications, helping it automatically identify and evaluate defects in interface design.
[0131] These three types of data sets complement each other, improving the comprehensive capabilities of the model in the field of interface design from basic understanding and description capabilities to in-depth design evaluation.
[0132] Step 2.2: Fine-tune strategy research.
[0133] The comprehensive image-text pair data set for software interface design specification evaluation constructed in step 2.1 is the necessary data for fine-tuning in this step. The instruction fine-tuning process is as follows:
[0134] Define the task goal: Let the model learn to generate accurate evaluations based on the input pictures and prompt words.
[0135] Set training parameters: including learning rate, batch size, optimizer (such as AdamW), etc.
[0136] Choose a loss function: Choose an appropriate loss function, such as cross entropy loss, to evaluate the difference between the model output and the ground truth answer.
[0137] Specific training process:
[0138] Forward propagation: Input the picture and prompt word into the model to get the predicted answer.
[0139] Calculate loss: Compare the predicted answer with the standard answer and calculate the loss value.
[0140] Back propagation: Calculate the gradient based on the loss value and update the model parameters.
[0141] Iterative training: Repeat the above steps until the preset number of training rounds is reached or the loss converges.
[0142] Therefore, this section mainly studies the fine-tuning strategy, and conducts experiments by adopting different strategies and adjusting the hyperparameters of the fine-tuning process to balance computing resources and fine-tuning efficiency. Finally, the following strategies were selected for fine-tuning:
[0143] 1. Use the deepspeed deep learning optimization library to accelerate and expand the fine-tuning process of large models;
[0144] 2. Use the LoRA strategy to adjust the model weights by introducing a low-rank matrix to achieve rapid learning;
[0145] 3. Use a large graph and text model of appropriate size, adjust the training hyperparameters, and modify the settings repeatedly based on the model performance and fine-tuning efficiency obtained after fine-tuning to balance computing resources and training efficiency;
[0146] 4. Use the training checkpoint saving function to save the model checkpoint every given step in order to record the training phase results and backup;
[0147] 5. Use lazypreprocess data preprocessing strategy to reduce memory burden by gradually loading and preprocessing data;
[0148] 6. Use wandb to record fine-tuning, conduct experimental tracking, parameter optimization, and model visualization during the fine-tuning process.
[0149] This embodiment achieves fine-tuning of the visual instructions of the general image and text model in the evaluation of software interface design specifications, and improves the task capability of the general image and text model in the evaluation of software interface design specifications. The constructed image and text model is similar to a visual dialogue assistant in form, supporting the input of a picture and a prompt word, and the output is the text answer of the model, which can conduct multiple rounds of dialogue.
[0150] Step 2.3: Benchmark verification of fine-tuned graph-text model.
[0151] The large image and text model established by fine-tuning the visual instructions in step 2.2 is the object of verification in this section. In order to verify the effectiveness of the fine-tuning method and that the model realizes the evaluation of software interface design specifications, a benchmark test for the software interface design specification evaluation task for the large image and text model is designed. The implementation process is as follows:
[0152] 1. Design various types of benchmark test questions and answers, including QA question-and-answer tests, interface description tests, and interface design specification compliance tests.
[0153] Input the benchmark test questions into the large graph-text model before and after fine-tuning, and collect the model output for each question.
[0154] 2. The model answers before fine-tuning and the benchmark answers are grouped together, and the model answers after fine-tuning and the benchmark answers are grouped together in another group. The designed prompt words are passed into ChatGPTAPI to obtain a comprehensive evaluation and score of the answer text.
[0155] 3. The scores before and after fine-tuning are counted and compared by benchmark test question and answer type, which shows the effectiveness of the fine-tuning method.
[0156] This embodiment designs a benchmark test for the intelligent evaluation of software interface design specifications for this research task, which verifies the effectiveness of the large graphic model established in step 2.2 and the ability of the large graphic model in the task of evaluating software interface design specifications.
[0157] Step 3: Evaluate the software interface design specifications based on the trained graphic and text model for interface design specification evaluation.
[0158] After fine-tuning and training in step 2, the large graphic and text model can achieve component-level content understanding, defect identification, and specification evaluation on software interface design specification issues. After inputting pictures and prompt words of the software interface, it can get answers about the design problems and violations of the design specifications of the software interface.
[0159] This part designs question-answering tests for each specification based on the established software interface design specifications, observes the identification of problems by the analysis model, and then makes a comprehensive quantitative evaluation of the software interface design specifications based on the identified problems and the number of elements, including:
[0160] Get relevant interface pictures of the software to be tested.
[0161] Design test language instructions according to the software interface general design specification list.
[0162] The relevant interface images and the test language instructions are used as input, and the trained large picture and text model for interface design specification evaluation is used to output the answer.
[0163] A score is calculated based on the number of problems and the severity of violations in each of the relevant interface images determined by the answers, wherein the number of problems refers to the number of interface elements that violate the software interface general design specification list, and the severity of the violation refers to the degree of violation of the software interface general design specification list.
[0164] For example, for several related pages of a certain software i , the number of elements on each page is e i , elements include controls and labels. Each page is used as the input of the large graphic model, and the number and type of violations of design specifications identified in the output are extracted. The page p iThe number of violations of the kth design specification is n i,k , the severity of the kth design specification is s k , there is a score calculation formula:
[0165]
[0166] This section proposes a quantitative evaluation method for the design specifications of the entire software interface based on the identification results and evaluation of a single software interface defect using a large graphic model. This method aims to comprehensively evaluate the degree of compliance with the software interface design specifications and provide quantitative indicators for the core of the entire task - software interface design specification evaluation.
[0167] For the static graphical user interface of the software, this embodiment can provide a large graphic model for interface design specification evaluation by analyzing the layout and content of interface elements, studying existing software interface design specifications and industry practices, and using technologies such as fine-tuning of graphic large model instructions. It also designs a benchmark test for the static interface design specifications of the graphic large model software to verify the performance of the model on this task, thereby achieving intelligent and accurate evaluation of the software interface design specifications.
[0168] The software interface design specification research method for the large graphic model involved in this embodiment can formulate a systematic software static interface design specification list for the large graphic model, and provide a theoretical basis for the intelligent evaluation of the large graphic model.
[0169] The method for fine-tuning a large graphic and text model for interface design specification evaluation involved in this embodiment can construct a high-quality, multi-type graphic and text pair data set based on a formulated design specification list, achieve efficient fine-tuning, establish an intelligent evaluation model, and provide a systematic quantitative indicator and evaluation method for the fine-tuned model.
[0170] The intelligent evaluation method for software interface design specifications based on a large graphic and text model involved in this embodiment can realize intelligent evaluation of software interface design specifications. Based on the formulated design specification list, it can accurately identify elements of the static interface that do not meet the design specifications and violate regulations and conduct a comprehensive evaluation, thereby providing auxiliary decision-making suggestions for software interface designers.
[0171] The present application also provides an application scenario, which applies the above-mentioned large-scale graphic model fine-tuning method for interface design specification evaluation. Specifically: the large-scale graphic model fine-tuning method for interface design specification evaluation provided in this embodiment can be applied in the scenario of user experience evaluation. This scenario includes a design specification analysis link, a model fine-tuning link, and an effect verification link. The large-scale graphic model fine-tuning method for interface design specification evaluation provided in this embodiment is a key technology in the model fine-tuning link. By verifying user feedback and actual application effects, it is ensured that the model provides accurate evaluation in different design scenarios, thereby optimizing user experience and design effects.
[0172] Example 2
[0173] The present embodiment provides a computer device, which can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store any data in a method for fine-tuning a large model of graphics and text for interface design specification evaluation provided in Example 1. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for fine-tuning a large model of graphics and text for interface design specification evaluation provided in Example 1.
[0174] Example 3
[0175] This embodiment provides a computer device, including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements a large graphic model fine-tuning method for interface design specification evaluation provided in the above embodiment 1.
[0176] Example 4
[0177] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method for fine-tuning a large graphic model for interface design specification evaluation provided in the above-mentioned embodiment 1 is implemented.
[0178] Example 5
[0179] This embodiment provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for fine-tuning a large graphic model for interface design specification evaluation provided in the above-mentioned embodiment 1.
[0180] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0181] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0182] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0183] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0184] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for fine-tuning a large graphic model for interface design specification evaluation, characterized in that: The method for fine-tuning a large graphic model for interface design specification evaluation includes: Construct a comprehensive image-text pair dataset for software interface design specification evaluation, wherein the comprehensive image-text pair dataset includes a first type of image-text pair dataset, a second type of image-text pair dataset, and a third type of image-text pair dataset. The first type of image-text pair dataset is used to improve the ability of the general image-text large model to understand the software interface content. The second type of image-text pair dataset is used to improve the ability of the general image-text large model to describe the interface content. The third type of image-text pair dataset is used to improve the ability of the general image-text large model to evaluate the software interface design specification. The general image-text large model is an existing pre-trained model. The first type of image-text pair dataset, the second type of image-text pair dataset, and the third type of image-text pair dataset all include software interface pictures, prompt words related to the software interface, and standard answers. The standard answer refers to an evaluation text generated according to a list of general design specifications for software interfaces. The list of general design specifications for software interfaces includes comprehensive interface design requirements and specific interface design specifications. The software interface image and the prompt word are used as feature data, the standard answer is used as label data, and the general image-text model is trained in combination with a fine-tuning strategy to obtain a trained image-text model for interface design specification evaluation.
2. The method for fine-tuning a large graphic model for interface design specification evaluation according to claim 1 is characterized in that: The specific construction process of the first type of image-text pair dataset includes: Obtain various types of question-answer data pairs about software interface images from existing public question-answer data sets, wherein each of the question-answer data pairs includes a question about the software interface image and a corresponding answer; According to the research task requirements and the software interface general design specification list, the question-answer data pairs about the software interface pictures are screened to obtain screened question-answer data pairs about the software interface pictures; Fine-tune the required data format according to the visual instruction, and convert the format of the question-and-answer data pair about the software interface image to obtain the question-and-answer data about the software interface image after format conversion; The first type of picture-text pair data set is obtained based on the question-and-answer data about the software interface pictures after conversion into the format.
3. The method for fine-tuning a large graphic model for interface design specification evaluation according to claim 1 is characterized in that: The specific construction process of the second type of image-text pair dataset includes: Acquire multiple effective expression forms of software interface images from existing public software interface data sets, wherein the effective expression forms refer to the expression forms of software interfaces that can improve the accuracy, comprehensiveness and richness of software interface descriptions; Encoding the file data corresponding to the effective expression form to obtain encoded file data, wherein the file data includes a software interface image and corresponding semantic information; Generating a software evaluation text according to the encoded file data and the prompt words designed by the prompt engineering; Performing format alignment processing on the software evaluation text and the software interface image in the file data to obtain an aligned software evaluation text and an aligned software interface image; The second type of image-text pair data set is obtained according to the aligned software evaluation text, the aligned software interface picture and the prompt words.
4. The method for fine-tuning a large graphic model for interface design specification evaluation according to claim 1 is characterized in that: The third type of image-text pair dataset includes a standard image-text pair dataset and an enhanced image-text pair dataset. The specific construction process of the enhanced image-text pair dataset includes: According to the software interface general design specification list, construct a standard software interface design theme template; Randomly modifying the interface elements and / or design layout in the standard software interface design theme template to generate software interface sample image data that violates the software interface general design specification list; Generate text evaluation information based on the randomly modified information; Performing format alignment processing on the software interface sample image data and the text evaluation information to obtain aligned software interface sample image data and aligned text evaluation information; The enhanced image-text pair data set is obtained according to the aligned software interface sample image data, prompt words and the aligned text evaluation information, wherein the prompt words are questions about the aligned software interface sample image data obtained by prompt engineering design.
5. The method for fine-tuning a large graphic model for interface design specification evaluation according to claim 1 is characterized in that: The comprehensive interface design requirements include interface consistency, design simplicity, intuitive principles, accessibility requirements and compliance with user habits; the specific interface design specifications include bright element colors, clear text meaning, standardization of interface elements, identification of interactive elements, unified text fonts, reasonable element layout, complete text display and coordinated font colors.
6. The method for fine-tuning a large graphic model for interface design specification evaluation according to claim 1 is characterized in that: The fine-tuning strategies include: using the DeepSpeed deep learning optimization library, using the LoRA strategy, using a general large image and text model of a preset scale, using the training checkpoint saving function, using the Lazy Preprocess data preprocessing strategy, and using Wand for fine-tuning records.
7. The method for fine-tuning a large graphic model for interface design specification evaluation according to claim 1 is characterized in that: After executing the step of "using the software interface image and the prompt word as feature data, using the standard answer as label data, and combining the fine-tuning strategy to train the general image-text model to obtain a trained image-text model for interface design specification evaluation", the image-text model fine-tuning method for interface design specification evaluation also includes: Get relevant interface pictures of the software to be tested; Designing test language instructions according to the software interface general design specification list; Taking the relevant interface pictures and the test language instructions as input, outputting answers using the trained large picture and text model for interface design specification evaluation; A score is calculated based on the number of problems and the severity of violations in each of the relevant interface images determined by the answers, wherein the number of problems refers to the number of interface elements that violate the software interface general design specification list, and the severity of the violation refers to the degree of violation of the software interface general design specification list.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the large graphic model fine-tuning method for interface design specification evaluation as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for fine-tuning a large graphic model for evaluating interface design specifications as described in any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for fine-tuning a large graphic model for evaluating interface design specifications as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Optimization method and equipment for interface design task allocation
CN116700680A
Image processing method and related equipment
CN118212231A
Image evaluation model construction method and device, evaluation method and device and storage medium
CN118230091A
General image aesthetic evaluation method based on state space model
CN118822992A
Cited By
Intelligent process automation method and device based on large model and machine vision
CN120315704A