A / b evaluation method, device, system, apparatus and storage medium
Patent Information
- Application Number
- CN202611007641.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0009]针对现有技术中的问题,本发明的目的在于提供A/B评估方法、装置、系统、设备及存储介质,以解决现有技术A/B评估方案可用性及一致性差的技术问题
[0026]通过将当前待评估设计对应的多模态信息与历史实验知识库中的历史A/B案例进行关联检索,并将检索得到的历史A/B案例及对应的相似度值作为推理依据输入大语言模型,使大语言模型在生成预测评估结果时能够结合历史实验经验进行分析,而不仅依赖模型自身的参数知识进行推理,从而增强预测评估结果与历史实验数据之间的关联性,提高预测评估结果的一致性和可信度。
Smart Images

Figure CN122816591A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to an A / B evaluation method, apparatus, system, device, and storage medium. Background Technology
[0002] In the field of product design and interaction optimization, it is usually necessary to predict the differences in the effects of different design solutions during the design phase in order to assist in design decisions.
[0003] In existing technologies, a common practice is to conduct A / B testing (user behavior segmentation experiments) after the product or feature is launched to compare and analyze different versions of the design scheme, in order to obtain the performance differences of each version in business metrics such as click-through rate, conversion rate, or user dwell time. This type of method relies on real user traffic and online experimental environment, so relatively reliable analysis results can generally only be obtained after the design scheme has been implemented and launched.
[0004] Meanwhile, some technical solutions attempt to incorporate historical experimental data into the design phase for reference and analysis of new design schemes. These solutions typically archive and store historical A / B test results, and retrieve relevant cases from historical data based on image or text information in the design draft when needed, to assist designers in making experience-based judgments.
[0005] In addition, there are technical approaches that utilize machine learning models or large language models to assist in the analysis of design schemes. In this type of approach, the design draft information and related descriptive information are typically input into the model, which then generates descriptive analysis results on the advantages and disadvantages of different design schemes based on existing training knowledge or external input information.
[0006] However, the aforementioned existing technologies still have certain limitations. For example, when conducting analysis during the design phase, the correlation between historical A / B experimental data and the current design scheme usually relies on manual screening or simple rule matching, which affects the stability and consistency of the model output results.
[0007] Therefore, improving the usability and consistency of the A / B evaluation process has become a pressing technical problem to be solved in this field.
[0008] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0009] In view of the problems in the prior art, the purpose of this invention is to provide an A / B evaluation method, apparatus, system, device and storage medium to solve the technical problems of poor usability and consistency of the existing A / B evaluation scheme.
[0010] The first aspect of this disclosure provides an A / B evaluation method applied to a server, including: It receives the baseline design draft to be evaluated, candidate design draft, and context metadata sent by the client, and uses the baseline design draft, candidate design draft, and context metadata to construct a multimodal query representation; Based on multimodal query representation, multimodal similarity retrieval is performed in the historical experimental knowledge base to obtain historical A / B cases and their corresponding similarity values; The multimodal query representation, historical A / B cases, and similarity values are input into the large language model, and the predicted evaluation results of the baseline design and candidate design are output.
[0011] In some implementations, multimodal query representations, historical A / B test cases, and similarity values are input into a large language model, including: Based on similarity values, multiple historical A / B cases are sorted and weighted to generate a structured evidence sequence, which includes at least a summary of historical cases sorted by similarity weight and corresponding experimental conclusions. The multimodal query representation, baseline design draft, candidate design draft, and structured evidence sequence are assembled into a cue context, which is then input into a large language model.
[0012] In some implementations, a multimodal query representation is constructed based on baseline design drafts, candidate design drafts, and contextual metadata, including: The first visual feature vector of the baseline design draft, the second visual feature vector of the candidate design draft, and the text feature vector of the context metadata are extracted using a pre-trained multimodal encoder. Using the cross-attention mechanism, the text feature vector is used as the query vector, and the first visual feature vector and the second visual feature vector are used as the key vector and value vector respectively to fuse them, so as to obtain the baseline multimodal fusion vector and the candidate multimodal fusion vector. The baseline multimodal fusion vector and the candidate multimodal fusion vector are concatenated to generate a multimodal query representation.
[0013] In some implementations, multimodal similarity retrieval is performed in a historical experimental knowledge base based on multimodal query representation to obtain historical A / B cases and corresponding similarity values, including: Extract image and text feature components from the multimodal query representation; Vector retrieval and inverted text retrieval are performed based on image feature components and text feature components respectively to obtain two candidate sets; The reciprocal ranking fusion algorithm is used to rearrange the two candidate sets to extract historical A / B cases, and the similarity value is normalized to the specified target interval based on the corresponding similarity monotonic transformation function.
[0014] In some implementations, the A / B evaluation method further includes the following steps before inputting the multimodal query representation, historical A / B cases, and similarity values into the large language model for inference: Computer vision rules are used to scan the baseline and candidate design drafts to identify the differences in page layout and the offset of the center of gravity of visually significant areas between the baseline and candidate design drafts, and then convert them into a structured difference feature matrix. The structured difference feature matrix is incorporated into the input context of the large language model.
[0015] In some implementations, after outputting the predictive evaluation results of the baseline design and candidate design, the A / B evaluation method further includes: The predicted evaluation results are subjected to structured inverse parsing to extract fields related to the expected A / B experiment results. These fields include at least: the expected dominant version direction, the expected relative effect range or level, and the relevant historical A / B case reference identifier. Deserialization and parsing are performed based on the function call interface of the large language model; If deserialization parsing fails, a fallback parsing mechanism based on regular expressions and a lightweight sequence labeling model will be enabled.
[0016] In some implementations, the A / B evaluation method also includes a post-implementation closed-loop write-back step: When the real online control experiment conclusion data corresponding to the baseline design draft and candidate design draft is detected to be generated, the real online control experiment conclusion data is obtained. The conclusion data of the real online controlled experiment are associated and encapsulated with the corresponding baseline design draft, candidate design draft and contextual metadata, and then written back to the historical experiment knowledge base.
[0017] In some implementations, the A / B evaluation method also includes offline evaluation and optimization steps: An evaluation sample set was constructed based on the conclusions of multiple sets of real online controlled experiments accumulated over history. The multimodal similarity retrieval, reasoning, and structured parsing steps are replayed for the samples in the evaluation sample set to obtain the expected A / B evaluation result for each sample; Compare the expected A / B evaluation results with the corresponding real online control experiment conclusions, and calculate an evaluation index matrix that includes the expected direction consistency rate and the retrieval relevance score. Based on whether the evaluation index matrix reaches the preset threshold, adjust the fusion weight when performing multimodal similarity retrieval, or adjust the prompt word template input into the large language model.
[0018] A second aspect of this disclosure provides an A / B evaluation method applied to a client, comprising: In response to user-triggered evaluation commands, retrieve the baseline design draft, candidate design draft, and associated contextual metadata in the current design canvas; The baseline design draft, candidate design draft, and contextual metadata are sent to the server so that the server can generate corresponding prediction and evaluation results based on historical knowledge base retrieval and large model inference. Receive and display the prediction and evaluation results returned by the server.
[0019] A third aspect of this disclosure provides an A / B evaluation system, comprising: The client responds to the evaluation command triggered by the user, obtains the baseline design draft, candidate design draft, and associated context metadata in the current design canvas; sends the baseline design draft, candidate design draft, and context metadata to the server, and receives and displays the predicted evaluation results from the server. The server constructs a multimodal query representation using baseline design drafts, candidate design drafts, and contextual metadata; it performs multimodal similarity retrieval in the historical experimental knowledge base based on the multimodal query representation to obtain historical A / B cases and corresponding similarity values; it inputs the multimodal query representation, historical A / B cases, and similarity values into a large language model to output the prediction and evaluation results of the baseline design draft and candidate design draft.
[0020] A fourth aspect of this disclosure provides an A / B evaluation apparatus applied to a server, comprising: The receiving module receives the baseline design draft to be evaluated, candidate design draft, and context metadata sent by the client, and uses the baseline design draft, candidate design draft, and context metadata to construct a multimodal query representation; The similarity retrieval module performs multimodal similarity retrieval in the historical experimental knowledge base based on multimodal query representation to obtain historical A / B cases and their corresponding similarity values; The evaluation module takes the multimodal query representation, historical A / B cases, and similarity values as input to the large language model and outputs the prediction evaluation results of the baseline design and candidate design.
[0021] The fifth aspect of this disclosure provides an A / B evaluation apparatus applied to a client, comprising: The response module responds to user-triggered evaluation commands by obtaining the baseline design draft, candidate design draft, and associated contextual metadata in the current design canvas. The sending module sends the baseline design draft, candidate design draft and contextual metadata to the server, so that the server can generate corresponding prediction and evaluation results based on historical knowledge base retrieval and large model inference. The display module receives and displays the predicted evaluation results returned by the cloud evaluation server.
[0022] A sixth aspect of this disclosure provides an electronic device, characterized in that it includes: a processor; and a memory storing executable instructions of the processor; wherein the processor is configured to perform an A / B evaluation method of any of the above embodiments by executing the executable instructions.
[0023] The seventh aspect of this disclosure provides a computer-readable storage medium for storing a program that, when executed, implements the A / B evaluation method of any of the above embodiments.
[0024] The eighth aspect of this disclosure provides a computer program product having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the A / B evaluation method of any of the above embodiments.
[0025] The A / B evaluation method, apparatus, system, device, and storage medium proposed in this disclosure have the following advantages: In this embodiment, the server receives the baseline design draft, candidate design draft, and contextual metadata sent by the client, and constructs a multimodal query representation using the baseline design draft, candidate design draft, and contextual metadata; performs multimodal similarity retrieval in the historical experimental knowledge base based on the multimodal query representation to obtain historical A / B cases and corresponding similarity values; further, the multimodal query representation, historical A / B cases, and similarity values are input into a large language model to output the prediction and evaluation results of the baseline design draft and candidate design draft.
[0026] By associating the multimodal information corresponding to the current design to be evaluated with historical A / B cases in the historical experimental knowledge base, and inputting the retrieved historical A / B cases and their corresponding similarity values as the basis for reasoning into the large language model, the large language model can combine historical experimental experience for analysis when generating prediction and evaluation results, rather than relying solely on the model's own parameter knowledge for reasoning. This enhances the correlation between the prediction and evaluation results and historical experimental data, and improves the consistency and credibility of the prediction and evaluation results.
[0027] Meanwhile, by unifying the baseline design draft, candidate design draft, and contextual metadata into a multimodal query representation, information from different modalities can jointly participate in the retrieval of the historical experiment knowledge base and subsequent reasoning process. This helps improve the matching accuracy between the design to be evaluated and historical experimental cases, thereby improving the usability of the A / B evaluation process, providing a more reliable evaluation basis for the design scheme before its official launch, and reducing the uncertainty brought about by human experience judgment.
[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0029] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings.
[0030] Figure 1 This is a flowchart illustrating an A / B evaluation method for servers provided in one embodiment of this disclosure; Figure 2 This is a flowchart illustrating an A / B evaluation method applied to a client, provided in one embodiment of this disclosure. Figure 3 This is a schematic diagram of the architecture of an A / B evaluation system provided in one embodiment of this disclosure; Figure 4 This is a schematic diagram illustrating the functional layering and module interaction principle of an A / B evaluation system provided in one embodiment of this disclosure. Figure 5 This is a diagram showing the core data flow and refined steps of a full-link A / B evaluation method provided in one embodiment of this disclosure; Figure 6 This is a detailed flowchart of the offline evaluation and optimization steps in an A / B evaluation method provided in one embodiment of the present disclosure; Figure 7 This is a time-series interaction diagram of an A / B evaluation method provided in one embodiment of this disclosure in an edge-cloud interaction scenario; Figure 8 This is a structural block diagram of an A / B evaluation device for a server provided in one embodiment of the present disclosure; Figure 9 This is a structural block diagram of an A / B evaluation device applied to a client, provided in one embodiment of this disclosure; Figure 10 This is a schematic diagram of the hardware structure of an electronic device suitable for implementing the A / B evaluation method according to an embodiment of the present disclosure. Detailed Implementation
[0031] To make the technical solution, the technical problem solved, and the technical effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the described embodiments are merely exemplary embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can make various modifications and variations to the embodiments without departing from the spirit and scope of the present invention, and all such modifications and variations should be considered to fall within the scope of protection of the present invention.
[0032] Techniques, methods, and apparatus known to those skilled in the art are not discussed in detail where appropriate, but should be considered part of this specification. In all examples shown herein, any specific numerical values or parameters should be interpreted as exemplary only and not as limiting the invention.
[0033] Figure 1 A flowchart illustrating an A / B evaluation method provided by an embodiment of this disclosure is shown. The method can be executed by a server, which can be deployed on a cloud computing platform or a local server cluster, and is used to perform predictive evaluation processing on baseline design drafts and candidate design drafts during the design phase.
[0034] The A / B evaluation method provided in this disclosure can be applied to design evaluation systems, product design analysis platforms, or design tool backend services to analyze and compare the expected effects of different design schemes during the design phase.
[0035] like Figure 1 As shown, the A / B evaluation method includes, but is not limited to, the following steps: Step 110: Receive the baseline design draft, candidate design draft, and context metadata sent by the client, and construct a multimodal query representation using the baseline design draft, candidate design draft, and context metadata; Step 120: Based on the multimodal query representation, perform multimodal similarity retrieval in the historical experimental knowledge base to obtain historical A / B cases and their corresponding similarity values; Step 130: Input the multimodal query representation, historical A / B cases, and similarity values into the large language model, and output the prediction and evaluation results of the baseline design and candidate design.
[0036] In this embodiment, the baseline design draft and candidate design draft can be image data, structured layout data, or a combination thereof. Contextual metadata is used to characterize the business attribute information or application scenario information corresponding to the design draft. A multimodal query representation is used to uniformly express the baseline design draft, candidate design draft, and contextual metadata, forming a unified input feature for subsequent retrieval.
[0037] The historical experiment knowledge base stores historical A / B test cases, which include at least historical design pairs, historical experiment results, and corresponding metadata. Through multimodal similarity retrieval, historical A / B test cases matching the current multimodal query representation can be selected from the historical experiment knowledge base, and corresponding similarity values are output to characterize the degree of similarity between the current design pair and historical cases.
[0038] Among them, the predictive evaluation results are used to characterize the predictive analysis results of the possible differences in performance between the baseline design and the candidate design without conducting real online experiments.
[0039] In this embodiment, a multimodal query representation is used to uniformly represent the combined features of image and text information, enabling data from different modalities to be processed in a unified representation space. Similarity values are used to characterize the degree of matching between the current design pair and historical A / B cases, thus providing a reference for subsequent model analysis. A large language model is used for semantic reasoning and result generation based on the input multimodal information and historical case information.
[0040] Through the above steps, after inputting the design draft data and context information, a unified multimodal query representation is first constructed; then, based on this query representation, similar A / B cases are retrieved from the historical experimental knowledge base and similarity values are obtained; finally, the above information is input into the large language model to generate prediction and evaluation results.
[0041] Therefore, compared to existing technologies that typically rely on real A / B test results after the design is completed for effect analysis, or rely solely on manual experience based on static historical data, lacking the ability to perform unified retrieval and model-assisted analysis of multimodal design drafts during the design phase, this disclosure constructs a multimodal query representation. This enables baseline and candidate design drafts, along with their contextual metadata, to participate in subsequent processing in a unified form. Based on this unified representation, a multimodal similarity retrieval is performed in the historical experiment knowledge base to obtain historical A / B cases and similarity values that correspond to the current design context. Furthermore, by inputting historical A / B cases and similarity values along with the multimodal query representation into a large language model, the large language model can perform association analysis and result generation under a unified input context, thereby reducing analytical bias caused by inconsistencies in information expression between different input sources.
[0042] Therefore, through the above-mentioned unified representation construction, multimodal similarity retrieval, and joint input processing mechanism based on historical A / B cases and similarity values, the A / B evaluation process can obtain structurally consistent input information and reusable historical reference information during the design phase, thereby improving the usability and consistency of the A / B evaluation process.
[0043] In one implementation, step 110 is used to perform unified representation processing on the baseline design draft, candidate design draft and contextual metadata to be evaluated, so as to generate a multimodal query representation for subsequent retrieval, thereby providing a standardized input basis for subsequent similarity analysis based on historical experimental knowledge base.
[0044] In its implementation, the A / B evaluation method is executed collaboratively by the client and server. The client responds to the user's evaluation trigger action on the design interface, retrieving the baseline and candidate design drafts from the design canvas, and simultaneously acquiring contextual metadata related to the design pair. The baseline and candidate design drafts can be screenshots, structured descriptions of the design drafts, or a combination of both. The contextual metadata characterizes the business semantics, interaction attributes, or usage scenario information corresponding to the design draft. The client then unifies and encapsulates this data, generating a standardized evaluation request data packet, which is sent to the server via a network interface.
[0045] Upon receiving the evaluation request data packet, the server parses it to obtain the baseline design draft, candidate design draft, and contextual metadata. In one embodiment, the server uses a pre-trained multimodal encoder to perform feature extraction processing on the parsed input data. The baseline design draft is encoded to obtain a first visual feature vector, the candidate design draft is encoded to obtain a second visual feature vector, and the contextual metadata is encoded to obtain a text feature vector. The pre-trained multimodal encoder may include a combination of a visual encoding network and a text encoding network, used to map data from different modalities to a unified feature space to support subsequent cross-modal fusion computation.
[0046] After obtaining the visual and text feature vectors, the server fuses the features from different modalities based on a cross-attention mechanism. Specifically, the text feature vector is used as the query vector, and the first and second visual feature vectors are used as the key and value vectors, respectively, to perform attention calculations, thereby obtaining the baseline multimodal fusion vector and the candidate multimodal fusion vector. By introducing text features as query information, visual features can be weighted and aggregated under semantic context constraints, thereby enhancing the alignment between visual and semantic information.
[0047] In one implementation, the server concatenates the baseline multimodal fusion vector and the candidate multimodal fusion vector to generate a multimodal query representation. This multimodal query representation is used to simultaneously characterize the joint distribution relationship between the baseline design draft and the candidate design draft in terms of visual structural features and semantic context features, and serves as a unified input representation for subsequent multimodal similarity retrieval using a historical A / B test knowledge base. After completing this processing, the server can use the multimodal query representation in subsequent retrieval processes and can return the processing status or intermediate results to the client to support the client-side interface display and workflow integration.
[0048] Through the processing method in step 110 above, a data interaction mechanism based on standardized evaluation requests between the client and the server is realized, enabling design data from different sources to enter the processing flow in a unified interface form. At the same time, image data and text data are mapped to a unified feature space, reducing the information fragmentation problem between different modalities. Furthermore, multimodal feature fusion is achieved under semantic constraints through a cross-attention mechanism, thereby improving the consistency and expression quality of multimodal representation.
[0049] In one embodiment, step 120 is used to perform a multimodal similarity retrieval in the historical experimental knowledge base based on the multimodal query representation, so as to obtain historical A / B cases that match the current design pair to be evaluated and the corresponding similarity values, thereby providing historical reference information for subsequent prediction and evaluation.
[0050] In the specific implementation, after receiving the multimodal query representation generated in step 110, the server parses the multimodal query representation to extract the image feature components and text feature components. The image feature components are used to characterize the visual structural information of the baseline design and the candidate design, while the text feature components are used to characterize the corresponding contextual semantic information.
[0051] In one implementation, the server performs two retrieval processes based on image feature components and text feature components, respectively. One process performs vector retrieval in the vector index space of the historical experiment knowledge base based on image feature components to obtain a set of historical candidate A / B cases that are similar to the visual structure of the current design draft; the other process performs inverted text retrieval in the text index structure of the historical experiment knowledge base based on text feature components to obtain a set of historical A / B cases that match the semantic information of the current context.
[0052] After obtaining two candidate sets, the server uses a reciprocal ranking fusion algorithm to merge and sort the two candidate sets. This process comprehensively re-ranks historical A / B cases obtained under different search paths, and extracts the top-ranked historical A / B cases as the final search result set. The reciprocal ranking fusion algorithm is used to comprehensively consider the ranking consistency under different search paths, thereby improving the overall recall quality and stability of historical A / B cases.
[0053] In one embodiment, the server further normalizes the retrieved original similarity values based on a preset similarity monotonic transformation function, mapping the similarity values to a specified target interval, so as to output normalized similarity values corresponding to each historical A / B case, thereby making the similarity results from different sources comparable.
[0054] Through the processing method described in step 120 above, historical A / B cases in the historical experimental knowledge base can be retrieved and matched simultaneously from both visual and semantic feature dimensions, thereby improving the completeness of historical case recall. At the same time, the reciprocal ranking fusion algorithm is used to uniformly sort the multi-way retrieval results, reducing the bias caused by a single retrieval method. Furthermore, through the normalization of similarity values, the degree of similarity between different historical cases is expressed under a unified dimension, thereby providing a more stable and consistent input basis for the subsequent evaluation and reasoning of the large language model.
[0055] In one implementation, step 130 is used to input the multimodal query representation, historical A / B cases, and similarity values into the large language model, and the large language model outputs the prediction and evaluation results of the baseline design and the candidate design.
[0056] In the specific implementation, after completing the multimodal similarity retrieval in step 120, the server first sorts the multiple historical A / B cases based on the similarity value, and then assigns corresponding weight values to each historical A / B case according to the similarity value, thereby generating a structured evidence sequence.
[0057] The structured evidence sequence includes at least historical A / B case summaries sorted by similarity weights and historical experimental conclusions corresponding to each historical A / B case, in order to characterize the reference contribution of different historical cases to the current evaluation task.
[0058] In one implementation, the server uniformly encapsulates the multimodal query representation, baseline design draft, candidate design draft, and structured evidence sequence, assembling them into a cue context. This cue context is used to uniformly express the input information and historical reference information of the current design pair to be evaluated, ensuring that the large language model can simultaneously acquire current design information and historical evidence information during a single inference process.
[0059] Subsequently, the server inputs the contextual information into the large language model. The large language model then performs joint semantic reasoning based on the current design semantic information represented by the multimodal query representation and the historical reference information represented by the structured evidence sequence, thereby outputting the predictive evaluation results for the baseline design and candidate design. These predictive evaluation results are used to characterize the predictive analysis output of the relative performance differences between the two design schemes in the absence of a real online A / B experiment.
[0060] In one implementation, the large language model can be a generative model based on the Transformer architecture. The server can execute the inference process by API calls or local deployment and return the output to the client for display.
[0061] Through the processing method described in step 130 above, historical A / B cases are no longer directly input into the model in their original form. Instead, they are processed through similarity-driven ranking and weighting to form a structured evidence sequence, thereby enhancing the consistency of expression and the distinguishability of contribution of historical information in the model reasoning process. At the same time, by constructing a prompting context together with the current multimodal query representation and the structured evidence sequence, the model can perform joint reasoning under a unified input context, thereby improving the stability and reference value of the prediction and evaluation results.
[0062] In one embodiment, before inputting the multimodal query representation, historical A / B cases, and similarity values into the large language model for inference, the baseline design draft and candidate design draft are further subjected to computer vision rule scanning processing to generate structured difference features to supplement the model input information, thereby enhancing the expressive completeness of subsequent inference input.
[0063] In the specific implementation, the server performs image analysis processing on both the baseline design draft and the candidate design draft. This image analysis processing can be performed based on a page element detection model or a rule-driven computer vision analysis module to identify the structured layout information on the page. During this process, the server extracts the page layout difference features between the baseline and candidate design drafts, such as changes in module position, changes in component hierarchy, or changes in area proportion, and further calculates the spatial distribution of visually salient areas based on the layout information.
[0064] In one embodiment, the server calculates the centroid of the visually salient regions in the baseline design and the candidate design, and obtains the centroid position of the corresponding visually salient regions. Then, it calculates the centroid offset between the baseline design and the candidate design to characterize the degree of difference in visual attention distribution between the two designs.
[0065] After completing the above calculations, the server organizes the page layout differences and the center-of-gravity offsets of visually salient areas into a structured feature matrix. This structured feature matrix is used to express the differences between the baseline design and the candidate design in terms of spatial layout and visual focus using a unified data structure.
[0066] In one implementation, the server incorporates the structured difference feature matrix as supplementary information into the input context of the large language model, and together with the multimodal query representation, historical A / B cases, and similarity values, it constitutes complete model input information, enabling the large language model to utilize semantic information, historical evidence information, and explicit structured difference information simultaneously when performing reasoning.
[0067] By employing the above processing methods, the visual structural differences between the baseline design draft and the candidate design draft can be explicitly quantified before entering the large language model for inference. This transforms the layout differences and attention distribution differences that were originally implicit in the image into structured feature inputs, enhancing the expressive richness and interpretability of the model's input information, and thus improving the stability and consistency of the prediction and evaluation results.
[0068] In one embodiment, after outputting the prediction evaluation results of the baseline design draft and the candidate design draft, the method further includes performing structured inverse parsing processing on the prediction evaluation results to extract structured fields used to characterize the prediction conclusions of the A / B experiment, thereby enabling the unstructured model output to be converted into a standard data format that can be computed and stored.
[0069] In the specific implementation, the server first receives the prediction evaluation results output by the large language model. The prediction evaluation results can be text data containing natural language descriptions. The server performs structural parsing processing on the text data to identify key information fields related to the A / B evaluation, including the expected dominant version direction, the expected relative effect range or level, and the associated historical A / B case reference identifiers.
[0070] In one implementation, the server prioritizes structured deserialization parsing of the prediction evaluation results based on the function call interface provided by the large language model. The model output is directly mapped to structured fields through a predefined function call protocol to reduce the reliance on unstructured text parsing.
[0071] When the deserialization parsing based on the function call interface is successful, the server directly obtains the structured output result and stores and returns it as a standardized prediction evaluation result.
[0072] In one implementation, when an anomaly is detected in the function call interface parsing, the output format does not conform to the predefined structure, or fields are missing, the server initiates a fallback parsing mechanism. This fallback mechanism includes regular expression-based rule matching and a lightweight sequence labeling model for information extraction. Regular expressions are used for fast matching and extraction of fixed-format fields, while the sequence labeling model is used to identify and extract entity and relational information from unstructured text.
[0073] Through the above dual-path parsing method, the server can prioritize high-precision function call parsing when the model output structure is stable, and perform compensatory parsing by combining rules and models when the structure is abnormal or incomplete, thereby improving the stability and usability of the structured output of the prediction and evaluation results.
[0074] Through the above processing, the prediction and evaluation results in the original natural language form are converted into structured field outputs, which can be directly called or stored by subsequent system modules, while improving the standardization of A / B evaluation results and the system's integrability.
[0075] In one embodiment, the A / B evaluation method further includes a post-hoc closed-loop write-back step and an offline evaluation and optimization step, which are used to continuously update and optimize the historical experimental knowledge base and retrieval and reasoning strategies based on the results of real online control experiments, thereby forming a data-driven closed-loop iterative mechanism.
[0076] In practice, when the evaluation system detects the generation of real online comparative experiment conclusion data corresponding to the baseline design draft and candidate design draft, the server retrieves the real online comparative experiment conclusion data from the experiment analysis system or the event tracking statistics system. The real online comparative experiment conclusion data may include conversion metrics, click metrics, or other business evaluation metrics of the baseline and candidate solutions under actual traffic conditions, and their comparison results.
[0077] In one implementation, the server associates and encapsulates the real online control experiment conclusion data with the corresponding baseline design draft, candidate design draft, and contextual metadata to form structured experimental record data. This structured experimental record data is used to simultaneously characterize the mapping relationship between design input information and real experimental output results, and is used for subsequent retrieval, model training, or evaluation.
[0078] After encapsulation, the server writes the structured experimental record data back to the historical experimental knowledge base to incrementally update the historical experimental knowledge base, so as to continuously accumulate real online control experimental samples, thereby providing a richer and higher quality source of historical cases for subsequent multimodal similarity retrieval.
[0079] In one implementation, after constructing an evaluation sample set based on the conclusions of multiple sets of real online comparative experiments accumulated in history, the server replays the multimodal similarity retrieval steps, reasoning steps, and structured parsing steps for each sample in the evaluation sample set in sequence to generate the expected A / B evaluation results of the corresponding samples, thereby simulating the complete online reasoning process.
[0080] Subsequently, the server compares the expected A / B evaluation results with the corresponding real online control experiment conclusion data, calculates the evaluation index matrix, which includes at least the expected direction consistency rate and the retrieval relevance score, and is used to characterize the consistency between the predicted results and the actual results as well as the quality of historical case recall.
[0081] In one implementation, the server adaptively adjusts system parameters based on whether the evaluation index matrix reaches a preset threshold. When the evaluation index does not reach the preset threshold, the fusion weight parameters in the multimodal similarity retrieval process can be adjusted to optimize the contribution ratio of different modal features in the retrieval ranking; or the prompt word templates input into the large language model can be adjusted to optimize the model's expression in terms of structured evidence utilization and predictive output.
[0082] Through the above processing methods, the system can continuously expand the historical experimental knowledge base based on real online comparative experimental results, while continuously optimizing the retrieval strategy and model input strategy, thereby improving the long-term consistency and predictive stability of A / B evaluation results and enhancing the system's adaptability under different design scenarios.
[0083] Figure 2 A flowchart illustrating an A / B evaluation method provided by an embodiment of this disclosure is shown. The method can be executed by a client, which can be a design tool plugin, a design editor client, or a front-end interactive module in a design collaboration platform, used to initiate evaluation requests and present predicted evaluation results during the design phase.
[0084] The A / B evaluation method provided in this disclosure can be applied to scenarios such as product design review, interface interaction optimization, comparison of operational page schemes, and analysis of design draft version iterations. It is used to predictively analyze and compare the expected effects of different design schemes before conducting a real online A / B experiment.
[0085] like Figure 2 As shown, the A / B evaluation method includes, but is not limited to, the following steps: Step 210: In response to the user-triggered evaluation command, obtain the baseline design draft, candidate design draft, and associated contextual metadata in the current design canvas; Step 220: Send the baseline design draft, candidate design draft, and contextual metadata to the server so that the server can generate corresponding prediction and evaluation results based on historical knowledge base retrieval and large model inference; Step 230: Receive and display the prediction evaluation results returned by the server.
[0086] In this way, by uniformly acquiring and encapsulating the baseline and candidate design drafts in the design canvas on the client side and sending them to the server for centralized analysis and processing, the design stage can obtain predictive evaluation results based on historical experimental knowledge base and large language model reasoning, thereby realizing the preliminary evaluation and comparative display of different design schemes.
[0087] Compared to existing technologies that typically require A / B testing to obtain feedback on the effectiveness of a design solution after it has been launched and real traffic has been distributed, the implementation method disclosed in this paper can output predictive evaluation results in advance during the design phase, reduce the solution verification cycle, and improve the efficiency of design decisions and the timeliness and reference value of solution comparison.
[0088] In one embodiment, the A / B evaluation method is applied to the client, which can be a design tool plugin, a web-based design editor, or a local design software extension module, to initiate an A / B evaluation request during the design phase and receive the predicted evaluation results returned by the server.
[0089] In the specific implementation, the client responds to user evaluation triggers in the design interface. For example, after editing or comparing the baseline and candidate design drafts, the user can initiate the A / B evaluation process via button clicks, shortcuts, or automatic trigger rules. Upon responding to the evaluation command, the client retrieves the baseline and candidate design drafts from the current design canvas. These drafts can be combinations of image screenshots, vector design data, or structured layout description data. Simultaneously, the client acquires contextual metadata related to the current design task. This contextual metadata characterizes the business scenario, page functional attributes, or interactive semantic information corresponding to the design draft.
[0090] In one implementation, the client encapsulates the acquired baseline design draft, candidate design draft, and contextual metadata to generate standardized evaluation request data. This data is then sent to the server via a network communication interface. The server then performs a multimodal similarity search based on a historical experimental knowledge base and uses a large language model to generate corresponding predictive evaluation results. The client and server can interact via HTTP, RPC, or message queues to achieve asynchronous or synchronous transmission of the evaluation request.
[0091] After the server generates the prediction and evaluation results, the client receives the results and visualizes them in the design interface. The prediction and evaluation results can be displayed in a structured card format, including a comparison of the baseline and candidate solutions, the expected dominant direction, and a description of the relative effects, to support user decision-making during the design phase.
[0092] The above processing method enables clients to directly initiate A / B evaluation requests during the design phase and obtain predictive evaluation results based on historical knowledge bases and large model reasoning through standardized data interaction with the server. This achieves pre-design evaluation capabilities and improves the efficiency of design decisions and the intuitiveness of scheme comparison.
[0093] Figure 3 This illustration shows a schematic diagram of an A / B evaluation system provided by an embodiment of the present disclosure. The A / B evaluation system includes a client 31 and a server 32, which interact with each other via a network communication interface to achieve predictive A / B evaluation functionality during the design phase.
[0094] In one embodiment, client 31 responds to the user's evaluation trigger operation in the design interface by retrieving the baseline design draft, candidate design draft, and associated contextual metadata from the current design canvas. The baseline and candidate design drafts can be image data, vector design data, or structured layout data, while the contextual metadata represents the business semantic information, interaction attribute information, or application scenario information corresponding to the design draft. After retrieving the data, client 31 performs standardized encapsulation processing to generate evaluation request data, which is then sent to server 32 via a network interface.
[0095] In one embodiment, after receiving the baseline design draft, candidate design draft, and contextual metadata sent by the client 31, the server 32 parses and processes the data, and constructs a multimodal query representation based on the baseline design draft, candidate design draft, and contextual metadata. The multimodal query representation is used to uniformly express image information and textual semantic information, forming a standardized input for subsequent retrieval and reasoning.
[0096] In one embodiment, server 32 performs a multimodal similarity retrieval in a historical experiment knowledge base based on a multimodal query representation, thereby obtaining historical A / B cases that match the current design pair and their corresponding similarity values. The historical experiment knowledge base stores historical A / B experiment data, and historical A / B cases include at least historical design pair data, historical experiment conclusions, and corresponding metadata information. The similarity value characterizes the degree of similarity between the current design pair and historical A / B cases.
[0097] In one implementation, server 32 inputs multimodal query representations, historical A / B cases, and similarity values into a large language model for joint inference to generate predictive evaluation results for baseline and candidate design drafts. These predictive evaluation results characterize the predictive analysis of the relative performance differences between different design schemes without conducting a real online A / B experiment.
[0098] After generating the prediction and evaluation results, the server 32 returns the prediction and evaluation results to the client 31. After receiving the results, the client 31 visualizes them for the user to use as a reference for design decisions.
[0099] Through the above system architecture and interaction method, the client 31 can complete the design data collection and request initiation, the server 32 completes the construction of multimodal representation, historical similarity retrieval and large language model reasoning generation, and sends the prediction evaluation results back to the client 31 for display, thereby realizing the design phase A / B evaluation function based on historical experimental knowledge base and large model reasoning ability.
[0100] Figure 4 This is a schematic diagram of the architecture of an A / B evaluation system provided for an embodiment of this disclosure. Figure 4 As shown, the A / B evaluation system includes: an input and retrieval layer 41, a reasoning and parsing layer 42, and an output layer 43.
[0101] In some implementations, the input and retrieval layer 41 is responsible for accessing the data to be evaluated, constructing multimodal query representations, and performing multimodal similarity retrieval. It contains four core components and forms three cross-layer data streams and one intra-layer data stream.
[0102] The baseline design draft and the candidate design draft serve as the data input sources to be evaluated, and their data flows across layers to the large language model inference module of the inference and parsing layer 42.
[0103] Contextual metadata serves as another data input source to be evaluated, and its data also flows across layers to the large language model reasoning module of the reasoning and parsing layer 42.
[0104] The historical experiment knowledge base serves as the data storage unit for historical A / B cases, and the data stored therein flows unidirectionally to the multimodal similarity retrieval module within this layer.
[0105] The similarity calculation engine performs multimodal similarity retrieval based on the historical experimental knowledge base, obtains historical A / B cases and their corresponding similarity values, and then flows the historical A / B cases and similarity values across layers to the evidence integration module of the reasoning and analysis layer 42.
[0106] In some implementations, the reasoning and parsing layer 42 is responsible for receiving the data provided by the input and retrieval layer 41 and generating the prediction and evaluation results. It contains four processing modules and has an internally converged serial structure.
[0107] The large language model inference engine receives cross-layer data from baseline design drafts, candidate design drafts, and contextual metadata, completes prediction evaluation inference, and then sends the inference results to the structured parser.
[0108] The evidence integration module receives historical A / B cases and similarity values from the similarity calculation engine, organizes and processes the historical A / B cases and similarity values, and then sends the processed structured evidence to the structured parsing module.
[0109] The structured parser, as the data aggregation node of this layer, receives input from the large language model inference engine and the evidence integration module, performs structured parsing on the prediction evaluation results, and then sends the data unidirectionally to the result evaluation module.
[0110] The results evaluation module receives the structured parsing results, performs integrity and consistency processing on the prediction evaluation results, and then sends the prediction evaluation results across layers to the machine-readable output module of the output layer 43.
[0111] In some implementations, the output layer 43 is responsible for the formatted output and multi-channel display of the prediction and evaluation results.
[0112] The machine-readable output module receives the prediction evaluation results from the result evaluation module, converts the prediction evaluation results into a preset machine-readable format, and then outputs the data in two parallel paths. One path is output to the client to display the prediction evaluation results, and the other path is output to the external system through an interface.
[0113] Figure 5 A flowchart illustrating an A / B evaluation method provided by an embodiment of this disclosure is shown. This method uses a client as the trigger and a server as the computation end, employing a hierarchical data flow architecture to perform the A / B evaluation. Figure 5 As shown, the A / B evaluation method includes, but is not limited to, the following steps.
[0114] In the input layer, perform the following steps: Step 510: The client triggers the A / B evaluation process, starting the processing flow in response to the evaluation command triggered by the user in the design tool or collaboration interface.
[0115] Step 511: Obtain the baseline design draft and candidate design draft to be evaluated. The baseline design draft and candidate design draft form a four-way parallel data stream: The first flow path is step 520, used for visual encoding of the manuscript; The second flow is directed to step 512, which is used for cascading extraction of context metadata; The third flow is to step 522, which is used for knowledge base interaction; The fourth flow is to step 521, which is used to participate in text and metadata extraction.
[0116] Step 512: Obtain context metadata, including page identifier, terminal type, business scenario, market information, and user segmentation. The context metadata forms four parallel data streams: The first path flows to step 520, serving as an auxiliary context for visual encoding; The second flow path is to step 521, which is used to extract features from plain text and metadata. The third flow direction is step 522, which serves as a coarse screening condition for the historical experimental knowledge base retrieval; Fourth cascade flow direction step 513 (obtain design change intent).
[0117] Step 513: Obtain design change intent, which characterizes the user's business-driven intent regarding design version differences. The design change intent data flows in three branches: The first path flows to the upper left, leading to step 520 (visual encoding of two manuscript pages - image modality). The second path flows directly downwards to step 521 (text and metadata extraction); The third path flows to the upper right towards step 522 (multimodal historical experiment knowledge base).
[0118] In the feature processing and knowledge retrieval layer, the following steps are performed: Step 520: Perform visual encoding of the two versions of the design. Using a pre-trained multimodal encoder, the design image from step 511, the auxiliary context metadata from step 512, and the design change intent from step 513 are received in parallel to perform visual modal space representation, generate visual feature vectors, and flow to step 530.
[0119] Step 521: Extract text and metadata features. As the text feature aggregation point of this layer, it receives the original text information from step 511, the contextual metadata from step 512, and the design change intent from step 513 in parallel, extracts a unified text semantic feature vector, and flows to step 530.
[0120] Step 522: Maintain a multimodal historical experiment knowledge base. Serving as a dynamic / static historical data foundation, it receives design draft data from step 511, metadata from step 512, and change intentions from step 513 in parallel, and stores the data in the knowledge base.
[0121] Step 523: Perform multimodal indexing and record retrieval of the knowledge base. Retrieve historical A / B test case data according to the retrieval strategy and proceed to step 530.
[0122] At the evaluation and analysis level, the following steps are performed: Step 530: Construct a multimodal query representation. As the first data aggregation center of the core layer of evaluation and analysis, this step receives the visual feature vector from step 520, the textual semantic feature vector from step 521, and the historical case data from step 523. It constructs a unified multimodal query representation through a multimodal feature alignment mechanism and then flows to step 531.
[0123] Step 531: Retrieve Top-K Similar Historical Cases. Based on the multimodal query representation, perform vector retrieval in the historical experiment knowledge base index, output the top K most similar historical cases, and proceed to step 532.
[0124] Step 532: Calculate the similarity score and hit summary. Quantify the similarity weight of each historical case and extract the core conclusion summary; the results flow to step 535.
[0125] Step 533 (parallel step): Perform a quality scan. Perform a computer vision rule scan on the baseline and candidate design drafts to identify page layout differences and rule features. The results flow to step 535.
[0126] Step 534: Perform difference feature evaluation. The similar historical data hit in step 532 and the rule difference features obtained from step 533 are intersected, recombined, and conflict evaluated. The comprehensive feature set is output and flows out of this layer.
[0127] In the inference layer of the large model, the following steps are performed: Step 540, Assemble the context. Receive the comprehensive feature set from step 534, assemble the multimodal query representation, prompt word template, retrieval evidence, and scanning features into a structured context that conforms to the input of the large model, and proceed to step 541.
[0128] Step 541, Large Model A / B Evaluation Reasoning. The assembled structured context is input into the large language model for reasoning, generating causal chain evaluation conclusions.
[0129] Step 542: Stream the A / B evaluation suggestions, output the inference results in the form of token-streamed text, and then proceed to step 550.
[0130] In the parsing and presentation layer, perform the following steps: Step 550, Structured Data Parsing. Using function call interfaces or fallback regular expression rules, reverse-parse the streaming evaluation recommendations to extract structured fields including the expected dominant version direction, expected risk / reward range, and historical case reference identifiers, and then proceed to step 551.
[0131] Step 551: The experimental results UI is displayed, and the fields after structured parsing are visualized and rendered on the client front-end interface.
[0132] In the closed-loop iteration layer, perform the following steps: Step 560: Write back the actual experimental conclusions. When the actual online A / B evaluation corresponding to the design draft is completed and the actual control experiment conclusion data is generated, obtain the data and write it back to the historical experiment knowledge base in step 522.
[0133] Step 561, Offline Evaluation Driven. A standard evaluation sample set is constructed based on the write-back data, the mainstream watershed is replayed, and the evaluation index matrix is calculated. The results flow to step 562.
[0134] Step 562: The strategy is continuously iterated. Based on whether the evaluation index matrix has reached the preset threshold, the multimodal similarity retrieval fusion weight or the large language model prompt word template is dynamically optimized.
[0135] Figure 6 This illustration shows a schematic diagram of an offline A / B evaluation and optimization process provided by an embodiment of the present disclosure. It is used to evaluate the A / B evaluation method offline based on the conclusion data of real online comparative experiments, and to continuously optimize the multimodal similarity retrieval strategy and the large language model inference strategy based on the evaluation results.
[0136] In some implementations, the offline evaluation process includes the following steps: Step 610: Evaluation Results Display. Specifically, the various evaluation indicators obtained from the offline evaluation are statistically summarized and visualized through reports or dashboards. The evaluation results serve as the basis for subsequent system optimization and then proceed to step 620.
[0137] Step 620: Strategy optimization configuration. Specifically, based on the evaluation results, the fusion weights, prompt word templates of the large language model, and related processing strategies in the multimodal similarity retrieval process are adjusted, and the optimized configuration flows to step 630.
[0138] Step 630: Construct the evaluation sample set. Specifically, an evaluation sample set is constructed based on historically accumulated real online controlled experiment conclusion data. The evaluation sample set includes historical design data, contextual metadata, and corresponding real online controlled experiment conclusion data, and is then input into step 640.
[0139] Step 640, Replay the A / B evaluation process. This includes: replaying the A / B evaluation process based on the evaluation sample set. The replay of the A / B evaluation process includes the following sub-steps.
[0140] Step 641, Multimodal Similarity Retrieval. Specifically, a multimodal query representation is constructed based on the design data in the evaluation samples, and a multimodal similarity retrieval is performed in the historical experimental knowledge base to obtain historical A / B cases and their corresponding similarity values, which then proceeds to step 642.
[0141] Step 642, Large Language Model Inference. Specifically, the multimodal query representation, historical A / B cases, and similarity values are input into the large language model to generate the corresponding predicted A / B evaluation results, which then flow to step 643.
[0142] Step 643, Structured Parsing. Specifically, the predicted A / B evaluation results are structured to obtain structured fields such as the expected dominant version direction, the expected relative effect range or level, and the associated historical A / B case reference identifiers, and then proceed to step 644.
[0143] Step 644, Result Comparison. Specifically, the predicted A / B evaluation results obtained from structured analysis are compared with the corresponding real online control experiment conclusion data in the evaluation sample. The comparison results are then directed to steps 650, 660, and 670, respectively.
[0144] In some implementations, the statistical analysis of evaluation metrics includes the following steps: Step 650, Expected Accuracy Assessment. Specifically, based on the comparison results of step 644, statistical analysis is performed on the consistency between the predicted A / B evaluation results and the actual online control experiment conclusion data, and evaluation indicators such as directional consistency rate, estimation error, and interval coverage rate are calculated respectively.
[0145] Step 651 is used to calculate the direction consistency rate; Step 652 is used to statistically estimate the error; Step 653 is used to calculate the interval coverage.
[0146] After summarizing the above statistical results, return to step 610 for visualization.
[0147] Step 660, retrieval quality assessment. Specifically, based on the comparison results of step 644, the retrieval performance of the historical experimental knowledge base is statistically analyzed, and evaluation indicators such as Top-K hit rate, similarity, and consistency with experimental conclusions are calculated.
[0148] Step 661 is used to calculate the Top-K hit rate; Step 662 is used to calculate the consistency between the similarity and the experimental conclusions.
[0149] After summarizing the above statistical results, return to step 610 for visualization.
[0150] Step 670: Output quality assessment. Specifically, based on the comparison results of step 644, statistical analysis is performed on the structured output quality of the predicted A / B assessment results, and evaluation indicators such as the completeness rate of required fields and the traceability of evidence are calculated respectively.
[0151] Step 671 is used to calculate the completeness rate of required fields; Step 672 is used for statistical evidence traceability.
[0152] The above statistical results are summarized and returned to step 610 for visualization, and used as the basis for strategy optimization configuration in step 620.
[0153] This embodiment provides an end-to-end interactive flow sequence diagram for an A / B evaluation method. Combined with... Figure 7 , Figure 7 The data interaction process between the user, client 31, and server 32 is shown.
[0154] The client 31 can be deployed in design tool plugins, design collaboration platforms, or other client devices capable of acquiring design draft data. The server 32 can be deployed in cloud servers, private servers, or other computing devices capable of executing A / B evaluation processes, and it integrates a historical experimental knowledge base, a large language model, and corresponding data processing modules.
[0155] In this embodiment, the user initiates an A / B evaluation request through the client 31. The client 31 is responsible for obtaining the data to be evaluated and sending it to the server 32. After the server 32 completes the multimodal similarity retrieval, large language model inference and prediction evaluation results, it returns the results to the client 31 for display, forming a complete A / B evaluation interaction process.
[0156] In step 710, the user submits the baseline design and candidate design drafts to be evaluated through client 31. Specifically, the user can select the two design drafts to be evaluated in the current design canvas of the design tool and trigger the A / B evaluation instruction.
[0157] In response to the A / B evaluation command, client 31 retrieves the baseline design and candidate design drafts from the current design canvas. The baseline design and candidate design drafts can be page screenshot image data, structured description data of the design drafts, or a combination of both.
[0158] In step 720, the user inputs contextual metadata corresponding to the baseline design draft and candidate design draft through client 31. The contextual metadata is used to describe the business semantic information, interaction attributes, or usage scenario information corresponding to the current design task, and may include page identifier, terminal type, business scenario, market information, user segmentation, design change intent, or other contextual information related to the current design task.
[0159] In step 730, the client 31 sends the baseline design draft, candidate design draft, and contextual metadata to the server 32.
[0160] Specifically, the client 31 can encapsulate the acquired image data and context metadata into an evaluation request according to a preset data format, and send it to the server 32 through a network interface, so that the server 32 can use the baseline design draft, candidate design draft and context metadata to construct a multimodal query representation and execute the subsequent A / B evaluation process.
[0161] In step 740, server 32 receives the evaluation request sent by client 31 and constructs a multimodal query representation using baseline design drafts, candidate design drafts, and contextual metadata.
[0162] Subsequently, in step 750, server 32 performs multimodal similarity retrieval in the historical experimental knowledge base based on the multimodal query representation to obtain historical A / B cases and their corresponding similarity values.
[0163] In step 760, server 32 sorts and weights the retrieved historical A / B cases based on similarity values to generate a structured evidence sequence. It then assembles the multimodal query representation, baseline design draft, candidate design draft, and structured evidence sequence into a prompt context, inputs it into a large language model for reasoning, and outputs the prediction evaluation results corresponding to the baseline design draft and the candidate design draft.
[0164] In some implementations, server 32 may also perform computer vision rule scanning on baseline design drafts and candidate design drafts before large language model inference to obtain a structured difference feature matrix, and incorporate the structured difference feature matrix into the prompt context to further improve the accuracy of prediction and evaluation results.
[0165] In step 770, server 32 performs structured inverse parsing on the prediction evaluation results output by the large language model, extracts structured fields such as the expected dominant version direction, the expected relative effect range or level, and the associated historical A / B case reference identifiers, and sends the parsed prediction evaluation results to client 31.
[0166] In step 780, client 31 receives the prediction and evaluation results returned by server 32 and displays them visually in the design interaction interface. Specifically, client 31 can display the prediction and evaluation results, related historical A / B cases, corresponding similarity values, and relevant analysis information, so that users can refer to the prediction and evaluation results to analyze the current design scheme during the design phase.
[0167] In step 790, the user views, compares, and performs subsequent design processing on the baseline design draft and candidate design draft based on the prediction and evaluation results displayed on client 31, in order to assist in completing the design scheme review or version selection.
[0168] In some implementations, after the design scheme is launched and the real online control experiment conclusion data is obtained, the client 31 can also obtain the real online control experiment conclusion data and send it to the server 32.
[0169] Server 32 associates and encapsulates the real online control experiment conclusion data with the corresponding baseline design draft, candidate design draft and contextual metadata, and writes it back to the historical experiment knowledge base to enrich the data source for subsequent multimodal similarity retrieval.
[0170] In a further implementation, server 32 can also construct an evaluation sample set based on historically accumulated real online comparative experimental conclusion data. By replaying the multimodal similarity retrieval, large language model reasoning, and structured parsing process, it can calculate evaluation index matrices such as expected direction consistency rate and retrieval relevance score. Based on the evaluation index matrix, it can adjust the fusion weights in the multimodal similarity retrieval process or the prompt word templates of the large language model, thereby achieving continuous optimization of A / B evaluation capabilities.
[0171] pass Figure 7 The interactive process shown involves client 31, which is responsible for collecting the data to be evaluated, sending requests, and displaying the predicted evaluation results. Server 32 is responsible for constructing multimodal query representations, multimodal similarity retrieval, large language model reasoning, and generating predicted evaluation results. The two work together to complete the A / B evaluation in the design phase and combine the conclusion data of real online comparative experiments to continuously update the historical experimental knowledge base, thereby continuously improving the accuracy, usability, and consistency of subsequent A / B evaluation processes.
[0172] In this disclosure, an A / B evaluation apparatus is also provided. For example... Figure 8 As shown, the A / B evaluation device 800 can be deployed on a server-side cloud computing platform, a privately deployed data center, or an edge computing node to perform A / B evaluation processing flow based on multimodal retrieval-enhanced large language model reasoning.
[0173] like Figure 8 As shown, the A / B evaluation device 800 includes a receiving module 810, a similarity retrieval module 820, and an evaluation module 830. The modules work together to realize a complete processing chain from receiving client requests and retrieving historical knowledge to generating predictive evaluation results.
[0174] In one embodiment, the receiving module 810 receives the baseline design draft, candidate design draft, and contextual metadata sent by the client, and parses and standardizes the data for subsequent multimodal representation construction. During this process, the receiving module 810 may also perform data integrity verification and format adaptation processing to ensure that the input data conforms to predefined multimodal input specifications.
[0175] In one embodiment, the similarity retrieval module 820 is used to construct a multimodal query representation based on the baseline design draft, candidate design draft and contextual metadata, and to perform multimodal similarity retrieval in the historical experimental knowledge base based on the multimodal query representation to obtain historical A / B cases and corresponding similarity values.
[0176] In its specific implementation, the similarity retrieval module 820 first uses a pre-trained multimodal encoder to extract the first visual feature vector of the baseline design draft, the second visual feature vector of the candidate design draft, and the text feature vector of the context metadata. Then, it uses a cross-attention mechanism to fuse the text feature vector as the query vector and the first and second visual feature vectors as the key and value vectors to obtain the baseline multimodal fusion vector and the candidate multimodal fusion vector. Finally, it concatenates the baseline multimodal fusion vector and the candidate multimodal fusion vector to generate a multimodal query representation.
[0177] After obtaining the multimodal query representation, the similarity retrieval module 820 decomposes the multimodal query representation, extracts the image feature components and text feature components, performs vector retrieval based on the image feature components, and performs inverted text retrieval based on the text feature components to obtain two candidate historical A / B case sets. Subsequently, the two candidate sets are rearranged using a reciprocal ranking fusion algorithm, and the final historical A / B cases are extracted. At the same time, the similarity value normalized to the target interval is output based on the similarity monotonic transformation function, thereby forming a structured historical retrieval result.
[0178] In one embodiment, the evaluation module 830 is used to input multimodal query representations, historical A / B cases and similarity values into a large language model, and output the prediction evaluation results of the baseline design and the candidate design.
[0179] In its implementation, the evaluation module 830 first sorts and weights multiple historical A / B cases based on similarity values to generate a structured evidence sequence. The structured evidence sequence includes at least a summary of historical cases sorted by similarity weight and the corresponding experimental conclusion information. Then, the multimodal query representation, baseline design draft, candidate design draft, and structured evidence sequence are assembled into a prompt context, and the prompt context is input into the large language model to perform joint inference and generate a predictive evaluation result.
[0180] In one embodiment, before inputting the prompt context into the large language model, the evaluation module 830 may also call the visual analysis submodule to perform computer vision rule scanning on the baseline design and candidate design to identify page layout differences and the center-of-gravity offset of visually significant areas, convert them into a structured difference feature matrix, and incorporate the structured difference feature matrix into the prompt context to enhance the structural expressiveness of the input information.
[0181] In one embodiment, after outputting the prediction evaluation result, the evaluation module 830 also performs structured reverse parsing on the prediction evaluation result, and performs deserialization parsing through the large language model function call interface to extract fields related to the expected A / B experiment results; when the parsing is abnormal, a fallback parsing mechanism based on regular expressions and a lightweight sequence labeling model is activated to ensure the stability of the structured output.
[0182] In one embodiment, the A / B evaluation device 800 can further collaborate with the historical experiment knowledge base update module and the offline evaluation module. Specifically, when real online control experiment conclusion data is detected, this data is associated and encapsulated with the corresponding design draft and contextual metadata, and written back to the historical experiment knowledge base. Simultaneously, an evaluation sample set is constructed based on the accumulated real online control experiment conclusion data, and the multimodal retrieval, reasoning, and parsing processes are replayed and evaluated. Furthermore, the retrieval fusion weights or prompt word templates are dynamically optimized based on the evaluation index matrix, thereby achieving continuous iterative optimization of the system.
[0183] The aforementioned device structure enables the server to modularly implement multimodal query representation construction, historical similarity retrieval, large language model reasoning, and structured result parsing. Furthermore, the closed-loop write-back and offline evaluation mechanisms facilitate continuous optimization of system capabilities, thereby enhancing the stability and scalability of A / B evaluation results.
[0184] It should be noted that the receiving module 810, the similarity retrieval module 820, and the evaluation module 830 mentioned above correspond to the implementation process of steps 110 to 130 of the A / B evaluation method in the embodiments of this disclosure. Their functions, processing flow, and application scenarios can be found in the specific descriptions in the corresponding method embodiments, and will not be repeated here.
[0185] In this disclosure, an A / B evaluation apparatus is also provided. For example... Figure 9 As shown, the A / B evaluation device 900 can be deployed in design clients, design editor plugins, or front-end interactive applications to initiate A / B evaluation requests, transmit data, and visualize the predicted evaluation results during the design phase.
[0186] like Figure 9 As shown, the A / B evaluation device 900 includes a response module 910, a sending module 920, and a display module 930. The modules work together to realize a complete processing chain from client-side evaluation triggering, design data acquisition, server request sending, and prediction evaluation result display.
[0187] In one embodiment, the response module 910 is used to respond to the evaluation command triggered by the user and obtain the baseline design draft, candidate design draft, and associated contextual metadata information from the current design canvas. The baseline design draft and candidate design draft can be image data, design structure data, or a combination thereof, and the contextual metadata is used to characterize the business semantic information, interaction attribute information, or application scenario information corresponding to the current design.
[0188] In one embodiment, the sending module 920 is used to standardize and encapsulate the baseline design draft, candidate design draft and contextual metadata, and send the encapsulated evaluation request data to the server, so that the server can generate corresponding predictive evaluation results based on historical experimental knowledge base retrieval and large language model inference.
[0189] In one embodiment, the display module 930 is used to receive the prediction evaluation results returned by the server, and to parse and visualize the prediction evaluation results so that users can intuitively obtain the comparison information of expected effects between the baseline design draft and the candidate design draft during the design stage.
[0190] It should be noted that the response module 910, the sending module 920 and the display module 930 mentioned above correspond to the implementation process of steps 210 to 230 of the A / B evaluation method in the embodiments of this disclosure. Their functions, processing flow and application scenarios can be found in the specific descriptions in the corresponding method embodiments, and will not be repeated here.
[0191] Those skilled in the art will understand that the various embodiments of this disclosure can be implemented as systems, methods, apparatuses, electronic devices, or computer program products. Therefore, the various modules, units, apparatuses, or platforms in this disclosure can be implemented in a completely hardware manner, a completely software manner (including firmware, executable program instructions stored in memory, microcode, etc.), or a hardware-software co-implementation manner, and used to perform corresponding data processing operations.
[0192] Based on the same inventive concept, this disclosure also provides an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the method described above by executing the executable instructions. Since the principle by which this electronic device embodiment solves the problem is similar to that of the above method embodiments, the implementation of this electronic device embodiment can refer to the implementation of the above method embodiments, and repeated details will not be described again.
[0193] Based on the same inventive concept, this disclosure also provides an electronic device, which is described below with reference to... Figure 10This invention describes an electronic device 1000 according to an embodiment of the present disclosure. The electronic device 1000 includes a processor (processing unit) 1010 and a memory (storage unit) 1020 for storing executable instructions of the processor. In specific industrial applications, the electronic device 1000 can be embodied as a live streaming operation control host, a live streaming media processing server, or a central computing unit integrated into a live streaming automation control system.
[0194] like Figure 10 As shown, the electronic device 1000 is manifested in the form of a general-purpose computing device. The components of the electronic device 1000 may include, but are not limited to: at least one processing unit 1010, at least one storage unit 1020, a bus 1030 connecting different system components (including storage unit 1020 and processing unit 1010), a display unit 1040, etc.
[0195] The storage unit stores program code, which can be executed by the processing unit 1010 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 1010 can perform the above-described steps. Figure 1 The steps of the method shown.
[0196] Storage unit 1020 may include readable media in the form of volatile storage units, such as random access memory (RAM) 1021 and / or cache memory 1022, and may further include read-only memory (ROM) 1023.
[0197] Storage unit 1020 may also include a program / utility 1024 having a set (at least one) program module 1025, such program module 1025 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0198] Bus 1030 can represent one or more of several bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures. Electronic device 1000 can also communicate with one or more external devices 1070 (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable user interaction with electronic device 1000, and / or any device that enables electronic device 1000 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed through input / output (I / O) interface 1050. Furthermore, electronic device 1000 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1060. As shown, network adapter 1060 communicates with other modules of electronic device 1000 via bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0199] In some embodiments, this disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described A / B evaluation method.
[0200] In some embodiments, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described A / B evaluation method.
[0201] Computer-readable storage media can be readable signal media or readable storage media. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), USB flash drive, portable hard disk, optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0202] Furthermore, computer-readable storage media may also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof.
[0203] A readable signal medium can be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Optionally, program code contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0204] In practical implementation, program code for performing the operations of this disclosure can be written in any combination of one or more programming languages. These programming languages include object-oriented programming languages—such as Java and C++—as well as conventional procedural programming languages—such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device.
[0205] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0206] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0207] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
Claims
1. An A / B evaluation method, characterized in that, Applied to servers, including: Receive the baseline design draft, candidate design draft, and context metadata sent by the client, and construct a multimodal query representation using the baseline design draft, candidate design draft, and context metadata; Based on the multimodal query representation, a multimodal similarity retrieval is performed in the historical experimental knowledge base to obtain historical A / B cases and their corresponding similarity values; The multimodal query representation, the historical A / B cases, and the similarity value are input into a large language model, and the predicted evaluation results of the baseline design and the candidate design are output.
2. The A / B evaluation method according to claim 1, characterized in that, The step of inputting the multimodal query representation, the historical A / B cases, and the similarity value into the large language model includes: Based on the similarity value, multiple historical A / B cases are sorted and weighted to generate a structured evidence sequence, wherein the structured evidence sequence includes at least a summary of historical cases sorted by similarity weight and corresponding experimental conclusions. The multimodal query representation, the baseline design draft, the candidate design draft, and the structured evidence sequence are assembled into a cue context, and the cue context is input into the large language model.
3. The A / B evaluation method according to claim 1, characterized in that, The construction of a multimodal query representation based on the baseline design draft, the candidate design draft, and the context metadata includes: The first visual feature vector of the baseline design draft, the second visual feature vector of the candidate design draft, and the text feature vector of the context metadata are extracted using a pre-trained multimodal encoder. Using the cross-attention mechanism, the text feature vector is used as the query vector, and the first visual feature vector and the second visual feature vector are used as the key vector and value vector respectively to fuse them, so as to obtain the baseline multimodal fusion vector and the candidate multimodal fusion vector. The baseline multimodal fusion vector and the candidate multimodal fusion vector are concatenated to generate the multimodal query representation.
4. The A / B evaluation method according to claim 1, characterized in that, The multimodal query representation performs multimodal similarity retrieval in the historical experimental knowledge base to obtain historical A / B cases and their corresponding similarity values, including: Extract the image feature components and text feature components from the multimodal query representation; Based on the image feature components and the text feature components, vector retrieval and inverted text retrieval are performed respectively to obtain two candidate sets; The two candidate sets are rearranged using a reciprocal ranking fusion algorithm to extract the historical A / B cases, and the similarity value is normalized to the specified target interval based on the corresponding similarity monotonic transformation function.
5. The A / B evaluation method according to claim 1, characterized in that, Before inputting the multimodal query representation, the historical A / B cases, and the similarity values into the large language model, the A / B evaluation method further includes: The baseline design and the candidate design are subjected to computer vision rule scanning to identify the page layout differences and the center-of-gravity offset of the visually significant regions between the baseline design and the candidate design, and then converted into a structured difference feature matrix. The structured difference feature matrix is incorporated into the input context of the large language model.
6. The A / B evaluation method according to claim 1, characterized in that, After outputting the predicted evaluation results of the baseline design and the candidate design, the A / B evaluation method further includes: The predicted evaluation results are subjected to structured inverse parsing to extract fields related to the expected A / B experiment results. These fields include at least: the expected dominant version direction, the expected relative effect range or level, and the associated historical A / B case reference identifier. Deserialization and parsing are performed based on the function call interface of the large language model; If deserialization parsing fails, a fallback parsing mechanism based on regular expressions and a lightweight sequence labeling model will be enabled.
7. The A / B evaluation method according to claim 1, characterized in that, The A / B evaluation method also includes a post-event closed-loop write-back step: When the generation of real online control experiment conclusion data corresponding to the baseline design draft and the candidate design draft is detected, the real online control experiment conclusion data is acquired. The conclusion data of the real online control experiment are associated and encapsulated with the corresponding baseline design draft, candidate design draft and context metadata, and then written back to the historical experiment knowledge base.
8. The A / B evaluation method according to claim 7, characterized in that, The A / B evaluation method also includes offline evaluation and optimization steps: An evaluation sample set was constructed based on the conclusion data of multiple sets of real online controlled experiments accumulated over history. The multimodal similarity retrieval, the input large language model, and the structured parsing steps are replayed for the samples in the evaluation sample set to obtain the expected A / B evaluation result for each sample; Compare the expected A / B evaluation results with the corresponding real online control experiment conclusion data, and calculate an evaluation index matrix that includes the expected direction consistency rate and the retrieval relevance score; Based on whether the evaluation index matrix reaches a preset threshold, the fusion weights during the multimodal similarity retrieval are adjusted, or the prompt word templates input into the large language model are adjusted.
9. An A / B evaluation method, characterized in that, Applied to the client side, including: In response to user-triggered evaluation commands, retrieve the baseline design draft, candidate design draft, and associated contextual metadata in the current design canvas; The baseline design draft, the candidate design draft, and the context metadata are sent to the server so that the server can generate corresponding prediction and evaluation results based on historical knowledge base retrieval and large model inference. Receive and display the prediction evaluation results returned by the server.
10. An A / B evaluation system, characterized in that, include: The client responds to the evaluation command triggered by the user and obtains the baseline design draft, candidate design draft and associated context metadata in the current design canvas; The baseline design draft, the candidate design draft, and the context metadata are sent to the server, and the prediction evaluation results are received from the server and displayed. The server uses the baseline design draft, candidate design draft, and contextual metadata to construct a multimodal query representation; Based on the multimodal query representation, a multimodal similarity retrieval is performed in the historical experimental knowledge base to obtain historical A / B cases and their corresponding similarity values; The multimodal query representation, the historical A / B cases, and the similarity values are input into a large language model, and the predicted evaluation results of the baseline design and the candidate design are output.
11. An A / B evaluation device, characterized in that, Applied to servers, including: The receiving module receives the baseline design draft, candidate design draft, and context metadata sent by the client, and uses the baseline design draft, candidate design draft, and context metadata to construct a multimodal query representation; The similarity retrieval module performs multimodal similarity retrieval in the historical experimental knowledge base based on the multimodal query representation to obtain historical A / B cases and their corresponding similarity values; The evaluation module inputs the multimodal query representation, the historical A / B cases, and the similarity value into the large language model, and outputs the prediction evaluation results of the baseline design and the candidate design.
12. An A / B evaluation device, characterized in that, Applied to the client side, including: The response module responds to user-triggered evaluation commands by obtaining the baseline design draft, candidate design draft, and associated contextual metadata in the current design canvas. The sending module sends the baseline design draft, the candidate design draft, and the context metadata to the server, so that the server can generate corresponding prediction and evaluation results based on historical knowledge base retrieval and large model inference. The display module receives and displays the predicted evaluation results returned by the cloud evaluation server.
13. An electronic device, characterized in that, include: processor; as well as A memory in which executable instructions of the processor are stored; The processor is configured to execute the A / B evaluation method of any one of claims 1 to 9 by executing the executable instructions.
14. A computer-readable storage medium for storing a program, characterized in that, When the program is executed, it implements the A / B evaluation method as described in any one of claims 1 to 9.
15. A computer program product having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the A / B evaluation method according to any one of claims 1 to 9.