Product detail page generation method, model training method, and electronic device
Patent Information
- Application Number
- CN202611041261.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]针对以上问题,本发明提供了一种商品详情页的生成方法、模型训练方法及电子设备,用以至少解决现有技术中如何高效构建商品各型号与图文元素间的映射关系以提升商品详情页的生成效率的技术问题
[0008]可选地,在数据获取步骤中,多模态数据包括说明书,确定页面渲染图和元素片段集合具体包括:获取目标商品的说明书,将说明书逐页渲染为多张页面渲染图;通过版面解析将说明书拆分为多个片段,多个片段具体包括文字提取结果以及由存储路径表示的图片元素;基于拆分得到的多个片段构建元素片段集合,并为元素片段集合中的每个片段分配标识符;在详情页生成步骤中,提取原始内容回填至页面模板具体包括:针对包含文字提取结果的片段,直接将文字提取结果回填至页面模板;针对包含图片元素的片段,则根据对应的存储路径调取相应的图片文件进行回填至页面模板。
Smart Images

Figure CN122816610A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and multimodal data processing technology, specifically to a method for generating product detail pages, a model training method, and an electronic device. Background Technology
[0002] In e-commerce, industrial product procurement, and commercial transactions, the product detail page is a crucial page that carries core information such as product name, technical parameters, specifications, and application instructions. With the increasing variety of products, generating product detail pages quickly and with high quality has become a vital step in improving information distribution efficiency. Typically, the raw materials for product detail pages mainly come from the multimodal data of the target product, such as instruction manuals or technical manuals containing various elements including text, tables, and images.
[0003] Currently, when generating product detail pages based on raw multimodal data, in addition to relying on manual layout, automated information extraction and generation schemes are often used. Since the same target product often corresponds to multiple specific sub-models (i.e., a model list containing multiple models), and the page template contains multiple different module titles, automated systems need to process a large amount of cross-modal information. Existing automated generation schemes typically require targeted information retrieval, matching, and content integration for each module title or different model in the page template.
[0004] However, as application scenarios place increasing demands on generation efficiency and layout accuracy, especially when dealing with complex multimodal data containing a large number of model lists and massive element fragments, how to more efficiently construct a layout structure that accurately reflects the mapping relationship between each model, module title and underlying element fragment, and how to better constrain the output format of the data during the generation process to ensure the legality of subsequent parsing, are technical problems that those skilled in the art continue to focus on and urgently need to solve. Summary of the Invention
[0005] To address the above problems, this invention provides a method for generating product detail pages, a model training method, and an electronic device, which at least solves the technical problem in the prior art of how to efficiently construct the mapping relationship between various product models and graphic elements to improve the generation efficiency of product detail pages.
[0006] In one or more embodiments of this application, a method for generating a product details page is provided, comprising: a data acquisition step, acquiring a list of model numbers of a target product, determining a page rendering image and a set of element fragments with identifiers corresponding to the target product based on multimodal data of the target product, and determining a module title according to a page template corresponding to the target product; an end-to-end inference step, inputting the page rendering image, the set of element fragments, the module title, and the list of model numbers into an end-to-end visual-language model for inference, and outputting a layout structure of the target product, wherein the layout structure includes a mapping relationship between the model numbers in the model number list and the module title and element references, and the element references are composed of identifiers; and a details page generation step, extracting original content from the set of element fragments and backfilling it into the page template according to the element references in the layout structure, thereby generating a product details page for the target product.
[0007] By employing the above methods, an end-to-end visual-language model directly outputs identifier-based mapping relationships and performs backfilling, avoiding the need for the model to directly output complete text and image content, significantly reducing the number of output tokens and decoding time. Simultaneously, this end-to-end architecture eliminates multiple independent retrieval and model components, achieving substantial compression of overall inference latency in typical scenarios, thereby improving the running speed and layout control accuracy of the product details page generation system.
[0008] Optionally, in the data acquisition step, the multimodal data includes the instruction manual. Determining the page rendering images and element fragment set specifically includes: acquiring the instruction manual of the target product and rendering the instruction manual page by page into multiple page rendering images; splitting the instruction manual into multiple fragments through layout parsing, the multiple fragments specifically include text extraction results and image elements represented by storage paths; constructing an element fragment set based on the multiple fragments obtained from the splitting, and assigning an identifier to each fragment in the element fragment set; in the details page generation step, extracting the original content and backfilling it into the page template specifically includes: for fragments containing text extraction results, directly backfilling the text extraction results into the page template; for fragments containing image elements, retrieving the corresponding image files according to the corresponding storage paths and backfilling them into the page template.
[0009] By rendering and parsing the instruction manual in the above way, the unstructured document is transformed into a structured visual and text asset that can be accurately located by the visual-language model, providing a foundation for subsequent direct use of identifiers to backfill the original content.
[0010] Optionally, determining the page rendering image and the set of element fragments further includes: obtaining fragment metadata for each fragment in the set of element fragments, the fragment metadata including an identifier, page number, fragment type, and bounding box; determining the page rendering image corresponding to the fragment metadata based on the page number, and establishing an association between the fragment metadata and the corresponding page rendering image to assist the end-to-end vision-language model in reasoning.
[0011] By aligning fragment metadata with page rendering images page by page and inputting them together, the end-to-end visual-language model is provided with more accurate spatial location and semantic type clues, which improves the model's understanding of complex layout pages.
[0012] Optionally, in the end-to-end inference step, the layout structure is JSON data, the top-level key of the JSON data is the model in the model list, and the value corresponding to the top-level key is the mapping between the module title and the element reference.
[0013] By using the above method, the layout structure can be expressed in JSON format with the product model as the top-level key. This allows for a clear and highly condensed organization of content across different models, greatly reducing the text burden of model generation.
[0014] Optionally, the end-to-end inference steps include: enabling guided decoding during the inference process of the end-to-end visual-language model; and pruning invalid decoding paths through guided decoding to constrain the output format of the layout structure.
[0015] By pruning invalid paths in a timely manner during the inference and search phases, the model output can always be a valid data structure that can be parsed by the program, thus fundamentally avoiding the failure of post-processing after structured decoding.
[0016] This application also provides a model training method applied to an end-to-end vision-language model. The model training method includes: a training sample acquisition step, which acquires original training samples, the original training samples containing input prompts and a complete gold standard layout composed of complete gold standard elements, the input prompts including a page rendering image, a set of element fragments with identifiers, a module title, and a model list; and a training sample enhancement step, which extracts a subset of page numbers from the page number set of the page rendering image, retains only the complete gold standard elements falling within the page number subset to generate the gold standard layout corresponding to the page number subset, thereby obtaining candidate samples, and discarding those that do not meet the pre-defined requirements. The process involves several steps: First, candidate samples are defined based on validity conditions, resulting in multiple training samples. Second, a supervised fine-tuning step is performed, using the gold standard layout corresponding to a subset of page numbers as the supervision target. The initial model's parameters are fine-tuned using the training samples to obtain a supervised fine-tuned model. Third, a group relative policy optimization step is taken, where the policy model and reference model are initialized using the supervised fine-tuned model. Multiple candidate outputs are sampled using the policy model for the same input cue. The reward for each candidate output is calculated, and the policy model is updated based on the relative advantage within the group. Finally, relative divergence constraints are used to limit the deviation between the policy model and the reference model, resulting in an end-to-end visual-language model.
[0017] By employing page number sampling for training data augmentation, the training distribution can cover documents of different length ranges, improving the model's robustness to varying input lengths. The second-stage training utilizes Group Relative Policy Optimization (GRPO) combined with an independent reward evaluation mechanism, which normalizes rewards using the within-group mean and standard deviation. This avoids the memory contention and reliance on neural network reward models inherent in traditional reinforcement learning schemes, enhancing the controllability and stability of policy updates.
[0018] Optionally, the supervised fine-tuning steps include: freezing the visual encoder backbone of the initial model; and using a low-rank adaptation method, updating the model parameters only for the multimodal projection layer and the language model of the initial model through backpropagation of cross-entropy loss to obtain a supervised fine-tuned model.
[0019] By employing the above methods, using a parameter-efficient fine-tuning strategy and freezing the visual backbone, training memory can be saved while effectively injecting typographic alignment knowledge, thereby reducing hardware resource consumption and suppressing overfitting.
[0020] Optionally, in the training sample augmentation step, candidate samples that do not meet the preset validity conditions are: candidate samples with fewer than a preset threshold of non-empty modules or that do not cover any model in the model list; in the group relative strategy optimization step, the reward for each candidate output is calculated, including: calculating the reward value through a preset rule-based combined reward function, which is constructed by multiplying the weighted sum of structural legality hard gate constraints and alignment evaluation indicators of multiple dimensions; wherein, the structural legality hard gate constraint is configured as follows: when the candidate output simultaneously satisfies the following conditions: it can be parsed into a valid JSON format, the fields satisfy the predefined data pattern schema, contains a valid identifier, and the included model exists in the model list, the value is 1; otherwise, the value is 0, so as to veto invalid decoding paths through the structural legality hard gate constraint.
[0021] By eliminating degenerate samples, the model learns incorrect blank mappings. By using deterministic offline rules to calculate rewards and introducing a hard-gate mechanism with a veto, the scoring illusion of the reward model is completely avoided. This also forces the policy model to quickly abandon corrupted formats, focusing the search space on the exploration of legal structures and improving the convergence speed of optimization. Optionally, the alignment evaluation metrics across multiple dimensions may include at least one of the following: a micro F1 score for measuring the correspondence between elements and module titles; a micro F1 score or element-wise Jaccard average for measuring the correspondence between elements and model numbers; a normalized depreciation cumulative gain (NDCG) metric for auditing the order of elements under each model number and module title; and a model coverage metric for characterizing the coverage ratio between the set of model numbers with non-empty content in the predicted output and the set of corresponding gold standard layouts for the page number subsets.
[0022] By introducing dual-path F1 scores and NDCG sorting audits into the combined rewards, strict alignment was achieved for the macro-structure attribution and the micro-arrangement order of elements, improving the reading logic experience of the details page. At the same time, by imposing constraints through the model coverage index, the "specialization" behavior of the model ignoring complex or long-tail models was successfully penalized, solving the pattern collapse problem in the multi-model coexistence generation scenario, thereby improving the completeness of information generation.
[0023] This application also provides an electronic device, including a processor and a memory, wherein the memory stores program instructions, and when the program instructions are executed, they implement the above-described method for generating a product details page or the above-described model training method. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating a method for generating a product details page in some embodiments of this application.
[0025] Figure 2 This is a schematic diagram of multimodal data in the data acquisition step of some embodiments of this application.
[0026] Figure 3 This is a schematic diagram of a page template in some embodiments of this application.
[0027] Figure 4 This is a schematic diagram of a product details page generated in one of the details page generation steps of some embodiments of this application.
[0028] Figure 5 This is a flowchart of a model training method in some embodiments of this application.
[0029] Figure 6 This is a schematic diagram illustrating end-to-end vision-language model enhancement based on training data page sampling in some embodiments of this application.
[0030] Figure 7 This is a schematic diagram illustrating the process of two-stage training of the end-to-end vision-language model in some embodiments of this application. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] <How to generate a product details page> refer to Figure 1This application also provides a method for generating product detail pages, which is performed using an end-to-end visual-language model. This generation method, following the execution logic sequence, includes data acquisition steps, end-to-end inference steps, and detail page generation steps.
[0033] Data acquisition steps include: Figure 2 and Figure 3 As shown, the process retrieves a list of target product models. Based on the multimodal data of the target products, it determines the corresponding page rendering image and a set of element fragments with identifiers. Simultaneously, it determines the module title based on the page template corresponding to the target products. The page template can be a pre-configured, standardized structural framework stored in a template library, or it can be a corresponding template automatically matched and retrieved from the template library based on the category attributes of the target products. The model list is independent of the multimodal data; its source can be a pre-configured database of target product model lists or a model text field entered by the merchant.
[0034] At the implementation level, the multimodal data of the target product can be rendered into images page by page at a fixed resolution, thus obtaining the page rendering image required for the input. Simultaneously, the multimodal data is split into multiple fragments through a layout parsing process. These fragments can include text fragments and image fragments. Text fragments are the extracted text results, and image fragments can be image elements represented by storage paths. Based on the split fragments, an element fragment set is constructed, and each fragment in the element fragment set is assigned a corresponding identifier.
[0035] Multimodal data, besides product PDF manuals, can also include product brochures, technical white papers, archived web pages, and other media that provide product information. The element fragment set includes not only text and image fragments, but also standard layout elements such as formula fragments and barcode fragments. Furthermore, the module titles required for the page template, in addition to product overview, product features, dimension diagrams, specification tables, and product images, can also include business-customized titles such as model composition, usage methods, operation manuals, use cases, application industries, technical parameters, and product structure.
[0036] It should be noted that, in addition to rendering and parsing original technical documents such as PDF manuals, multimodal data can also be obtained by crawling existing online product history image and text pages, or by parsing the electronic warehouse tag data stream of products. Similarly, the page template can adopt a preset structured tree-like template. Furthermore, in other embodiments, besides structured tree-like templates, the page template can also adopt unordered webpage card sketches, flat grid layouts, or style sheets with layout tags.
[0037] The end-to-end inference process includes assembling the prepared page rendering image, element fragment set, module title, and model list, and then inputting them into the end-to-end visual-language model for inference. After processing, the end-to-end visual-language model outputs the layout structure of the target product. This layout structure includes the mapping relationship between the models in the model list and the module titles and element references. Each element reference consists of an element identifier, and each element reference contains the identifier of the corresponding fragment. It should be noted that the layout structure used to represent the mapping relationship can be in data format other than JSON, such as XML, YAML, or a custom hierarchical key-value dictionary, which can express tree-like or hierarchical mapping relationships. The end-to-end visual-language model can be a pre-trained and fine-tuned multimodal large model, or other large image-text language models with image understanding and text generation capabilities.
[0038] Furthermore, because the layout structure output by the model contains element references rather than the actual content of the elements, the model does not need to repeatedly decode large amounts of long text during inference. This mechanism reduces the model's lexical output pressure, shortens decoding time, and avoids the illusion problem that arises in generative models due to autonomous text generation.
[0039] Compared to traditional detail page generation, which often relies on independent vector libraries, retrieval systems, and sorting systems across multiple discrete stages, and frequently requires calling large models independently for each module, this implementation utilizes an end-to-end visual-language model to uniformly understand the visual layout of the page rendering and the textual semantics of the fragments. After inference, the model directly outputs the entire page layout structure, thereby constructing a structured mapping between model numbers, module titles, and element fragments. More importantly, the model does not repeatedly output the original content of elements; its output layout structure only retains element references for the corresponding fragments. This process avoids data transfer between multiple components and eliminates the need for independent vector libraries, dense retrieval systems, sparse retrieval systems, and fusion sorting systems.
[0040] The steps to generate a product detail page include: Figure 4As shown, based on the element references in the layout structure output by the model, the original content is extracted from the element fragment set and backfilled into the page template to generate the product details page of the target product. Specifically, the original content is located and backfilled from the fragment dictionary using element identifiers. Since text fragments correspond to the extracted text results, and image fragments correspond to the storage path of image elements (which can be represented as a system path), in this backfilling process, for text fragments, the extracted text results are directly backfilled into the page template; for image fragments, the corresponding image file is retrieved and loaded according to its storage path and backfilled into the page template, rather than directly embedding the original byte data of the image during data flow. It should be noted that the step of extracting and backfilling the original content can be performed not only during the front-end rendering stage, but also on the back-end server side by directly concatenating and generating static web page code through a template engine, or by parsing and binding data through local components of the mobile native application.
[0041] The end-to-end solution eliminates the need for independent vector retrieval and fusion sorting components, as well as the need to independently call large models for each module, thus significantly reducing the number of system components and system maintenance costs. By uniformly generating the layout structure through a visual-language model, the overall inference latency for each product series is compressed from minutes in traditional solutions to near the time level of a single model inference, achieving significant speedup in real-world applications and effectively improving generation and processing efficiency. Since the layout structure output by the model replaces the generation of the original image and text content with element references, this output format reduces the number of output markers compared to solutions that directly output complete image and text content, further reducing decoding time and improving generation efficiency.
[0042] Meanwhile, because the original content such as text, data, and images in the final product details page comes entirely from a pre-determined set of element fragments, it is a deterministic data backfilling rather than being generated autonomously by a large model. This fundamentally eliminates the illusion and fabrication problems that may occur when multimodal generation models generate content, ensuring the accuracy, rigor, and authenticity of the final product details page content.
[0043] The following section, with reference to the figures, uses the EX-10 series of photoelectric sensors as an example to explain in detail each step of the product details page generation method provided in this embodiment.
[0044] First, perform the data acquisition steps. See below for further details. Figure 2Obtain the instruction manual corresponding to the target product. The instruction manual can be a PDF document or other electronic document format. For the obtained instruction manual, render each page at a preset resolution to obtain multiple corresponding page rendering images. The page rendering images are used to preserve the overall visual layout information of the instruction manual, including text layout, image positions, table structure, and page hierarchy.
[0045] After obtaining the page rendering, the instruction manual undergoes layout parsing. Layout parsing can be achieved using document analysis models, OCR layout analysis models, or rule-based parsing programs. Layout parsing identifies text areas, image areas, table areas, and other content areas in the instruction manual, and breaks each area down into multiple independent fragments, thus obtaining a set of element fragments.
[0046] The fragments in the element fragment set can include one or more of the following: text fragments, image fragments, and table fragments. For example, the instruction manual for the EX-10 series photoelectric sensor can be broken down into multiple element fragments, such as product overview text, product feature text, specification parameter table, dimension diagram image, wiring instruction image, and application case image. For ease of subsequent reference, each fragment in the element fragment set is assigned a unique identifier.
[0047] Identifiers can be generated using numerical codes, string codes, hash codes, or other unique identification methods. For example: E001 corresponds to the product overview segment; E002 corresponds to the product features segment; E003 corresponds to the size image segment; and E004 corresponds to the specification table segment.
[0048] After segmenting the elements, the metadata for each segment is further obtained. The segment metadata includes: identifier; page number; segment type; and bounding box. The page number indicates the segment's location on the page in the instruction manual; the segment type indicates whether the segment is a text segment, image segment, or table segment; and the bounding box indicates the segment's location within the page. The bounding box can be represented using rectangular coordinates, including the coordinates of the top-left and bottom-right corners.
[0049] Subsequently, the page rendering image corresponding to the fragment metadata is determined based on the page number. For example, if an element fragment is located on page 5, then the element fragment is determined to correspond to the page rendering image of page 5.
[0050] In addition, such as Figure 3 As shown, the module title is determined based on the page template corresponding to the target product. The module title describes the preset display module in the product details page. The module title may include: Product Overview; Product Features; Dimension Diagram; Specification Table; Product Image; Model Composition; Usage Instructions; User Manual; Use Cases; Application Industries; Technical Parameters; Product Structure.
[0051] Subsequently, an end-to-end inference step is performed, inputting the page rendering image, element fragment set, fragment metadata, module title, and model list into the end-to-end visual-language model. The end-to-end visual-language model performs inference based on the input content and outputs the layout structure corresponding to the target product.
[0052] In this implementation, the layout structure is preferably represented using JSON data. The top-level key of the JSON data is the model number in the model list. Each model number corresponds to a key-value object. The key-value object stores the mapping relationship between module titles and element references. Element references are composed of identifiers. For example, the "Product Features" module corresponding to a certain model can reference identifier E002; the corresponding "Specifications" module can reference identifier E004; and the corresponding "Dimension Diagram" module can reference identifier E003. Therefore, the model outputs element reference relationships, rather than the original element content. This avoids repeatedly generating text, images, or tables from the instruction manual, reduces the model output length, and improves inference efficiency.
[0053] Furthermore, guided decoding is enabled during model inference. Guided decoding can be implemented based on context-free grammars, regular expression constraints, or predefined data schemas. Guided decoding performs real-time validation of candidate decoding paths during model generation. When a candidate decoding path does not meet JSON syntax requirements, it is pruned. When a candidate decoding path generates an identifier that does not exist in the element fragment set, it is pruned. When a candidate decoding path generates a model that does not belong to the model list, it is pruned. When a candidate decoding path generates a model that does not conform to the predefined field structure, it is pruned. Through these constraints, the layout structure of the model output always meets the preset format requirements.
[0054] Next is the process of generating the details page, such as... Figure 3 and Figure 4 As shown, after obtaining the layout structure, the original content associated with the corresponding identifier is retrieved from the element fragment set based on the element references in the layout structure. Then, the retrieved original content is filled back into the corresponding module position in the page template to obtain the product details page for the target product.
[0055] For example, when the layout structure specifies that the "Product Features" module references E002, the content corresponding to E002 is extracted from the element fragment set and populated into the Product Features area of the page template. When the layout structure specifies that the "Specifications Table" module references E004, the content of the table corresponding to E004 is populated into the Specifications Table area. After all elements are populated, the product details page corresponding to the target product is generated.
[0056] By uniformly inputting instruction manual renderings, element fragment sets, fragment metadata, module titles, and model lists into an end-to-end visual-language model, the model can simultaneously utilize page visual, content, and structural information for layout reasoning, thereby improving module matching accuracy. Furthermore, by establishing a relationship between fragment metadata and page renderings, the model can use information such as page numbers, fragment types, and page positions to assist in determining element relationships, improving the accuracy of layout structure generation. Additionally, the layout structure is expressed in JSON data format, and element references are used instead of raw content output, reducing the output data size, minimizing token consumption during reasoning, and improving generation efficiency.
[0057] Furthermore, by pruning invalid decoding paths through guided decoding, the model output process is subject to format constraints, which improves the standardization and parsability of the layout structure, reduces post-processing anomalies, and enhances the stability and reliability of the product details page generation process. For large product series containing multiple models, this implementation method can also improve the accuracy of the correspondence between model numbers and module content, thereby improving the completeness and consistency of the product details page.
[0058] <Model Training Methods> The following is combined Figures 5-7 This application describes some of the model training methods provided in its embodiments. These model training methods can be applied to the training process of the end-to-end visual-language model in the above embodiments to obtain an end-to-end visual-language model capable of generating layout structures based on page rendering images, element fragment sets, module titles, and model lists. (Reference) Figure 5 The model training method includes steps for obtaining training samples, training sample augmentation, supervised fine-tuning, and population relative policy optimization.
[0059] The steps to obtain training samples include: First, obtaining the original training samples. The original training samples include input prompts and a complete gold-standard layout composed of complete gold-standard elements. The input prompts include a page rendering image, a set of element fragments with identifiers, module titles, and a model list. The page rendering image can be obtained by rendering the product manual page by page; the set of element fragments can be obtained through layout parsing; the module titles represent the various display modules in the product details page; and the model list represents the multiple product models included in the target product series. The complete gold-standard layout records the layout results obtained through manual annotation, and it includes the correspondence between model numbers, module titles, and element references. Element references consist of identifiers for the corresponding element fragments.
[0060] The training sample enhancement step includes: after obtaining the original training samples, performing training sample enhancement processing on the original training samples. Specifically, the page number set corresponding to the page rendering image is obtained, and a page number subset is extracted from the page number set. The page number subset can be obtained through random sampling, and its page number can be randomly determined within a preset range. For example, for an original training sample containing a 20-page instruction manual, eight to nineteen pages can be randomly selected as the page number subset. Subsequently, only complete gold standard elements falling into the page number subset are retained, and a gold standard layout corresponding to the page number subset is generated based on the retained complete gold standard elements, thereby obtaining candidate samples. The candidate samples include the input prompts corresponding to the page number subset and the gold standard layout corresponding to the page number subset. Candidate samples that do not meet the preset validity conditions are eliminated, and candidate samples that meet the preset validity conditions are retained. By performing multiple page number subset samplings on the same original training sample, multiple training samples can be obtained, so that the training data covers input scenarios with different lengths and different page combinations.
[0061] The supervised fine-tuning process includes: after obtaining training samples, supervising the initial model using these samples. Specifically, input prompts from the training samples are fed into the initial model, which then outputs the corresponding layout structure prediction. The gold standard layout corresponding to the page number subset is used as the supervised target, and the training loss is calculated by comparing the difference between the prediction and the gold standard layout. Cross-entropy loss can be used as the training loss. Backpropagation is then performed based on the training loss to update the model parameters. After multiple rounds of iterative training, a supervised fine-tuned model is obtained. Through this process, the model learns the correspondence between page rendering images, element fragments, module titles, and model lists with the layout structure, thus acquiring basic layout generation capabilities.
[0062] The group relative policy optimization steps include: after obtaining the supervised fine-tuning model, further performing group relative policy optimization. Specifically, the supervised fine-tuning model is used to initialize the policy model and the reference model, respectively, where the policy model participates in training and updates, and the reference model remains frozen. For the same input cue, multiple candidate outputs are sampled using the policy model, each candidate output corresponding to a layout structure prediction result. Subsequently, a reward value is calculated for each candidate output to evaluate the degree of matching between the candidate output and the corresponding gold standard layout. Based on the reward values corresponding to multiple candidate outputs, the intra-group relative advantage of each candidate output is calculated. The intra-group relative advantage is used to characterize the relative quality of the current candidate output relative to other candidate outputs in the same group. Subsequently, the policy model is updated according to the intra-group relative advantage, so that the generation behavior corresponding to high-quality candidate outputs is strengthened, while the generation behavior corresponding to low-quality candidate outputs is suppressed. At the same time, the deviation between the policy model and the reference model is limited by relative divergence constraints (such as relative KL constraints), so that the policy model maintains consistency with the behavior of the reference model while continuously optimizing the output quality. After multiple rounds of policy updates, a trained end-to-end vision-language model is obtained.
[0063] After training, this end-to-end vision-language model can receive page rendering images, element fragment sets, module titles, and model lists as input, and directly output the layout structure corresponding to the target product. The layout structure records the mapping relationship between model, module title, and element references, thereby supporting content backfilling and page construction in the subsequent product details page generation process.
[0064] By constructing training samples including page rendering images, element fragment sets, module titles, and model lists, the model can simultaneously learn the relationships between page visual features, content features, and layout features. Training samples of different lengths and page combinations are generated through page number subset sampling, expanding the range of training data distribution and improving the model's adaptability to scenarios with different sizes of instruction manuals and different input lengths. Supervised fine-tuning enables the model to learn the organizational patterns in manually annotated layouts, improving the accuracy of layout structure prediction. Furthermore, group relative strategy optimization allows the model to continuously optimize the layout generation strategy based on reward feedback, improving module attribution accuracy, model matching accuracy, and layout organization rationality. Simultaneously, introducing a reference model to constrain the strategy update process enhances training stability and reduces the risk of model performance fluctuations. The resulting end-to-end visual-language model exhibits high layout generation accuracy, strong generalization ability, and good training stability, thereby improving the accuracy, completeness, and efficiency of automatically generated product detail pages.
[0065] The following section will continue to use the EX-10 series of photoelectric sensors as an example to provide a detailed explanation of the steps involved in the model training method.
[0066] like Figure 6 As shown, in the training sample augmentation step, to improve the robustness of the model under different input lengths and page combinations, each sample in the original training set (a set of complete PDF pages and its gold standard layout) is augmented in the following way: Sample 0 uses all pages and the complete gold standard; Samples 1 to k are randomly sampled multiple times from the page number set of the sample, with a size between preset upper and lower bounds (e.g., lower bound 8, upper bound is the total number of pages of the sample minus 1), and each time only the gold standard elements falling within the sampled subset are retained as training labels; if the number of non-empty modules is lower than the threshold or no longer covers any model after sampling, the degraded sample is discarded. Through the above strategy, a single original sample can be augmented into several training samples (e.g., about six), so that the training distribution covers different length ranges from short documents to complete documents.
[0067] Furthermore, such as Figure 7 As shown, this end-to-end vision-language model employs a two-stage fine-tuning strategy, consisting of a supervised fine-tuning step and a group-relative policy optimization step. Phase 1: Supervised Fine-Tuning (SFT). In this phase, the gold-standard layout JSON, after manual verification, is used as the supervision target. Low-rank adaptation (LoRA or DoRA) is employed to efficiently fine-tune the parameters of the visual-language model. To conserve GPU memory, the visual encoder backbone can be frozen, and only the multimodal projection layer and the language model are updated. Optionally, 4-bit or 8-bit quantization can be used to load the base model weights. Techniques such as cosine learning rate scheduling and gradient checkpointing can be used to reduce peak GPU memory usage. In this phase, model parameters are updated through backpropagation using cross-entropy loss, resulting in the SFT checkpoint.
[0068] Phase Two: Group Relative Policy Optimization (GRPO). In this phase, the policy model π_θ and the reference model π_ref are initialized simultaneously using the SFT checkpoints obtained in Phase One. The policy model is trainable, while the reference model is frozen. Several (e.g., N = 8) candidate outputs o_1, o_2, …, o_N are sampled for the same input. For each candidate, the following rule-based combined reward r(o_i) is calculated. The reward is normalized to the relative advantage A_i = (r(o_i)) using the within-group mean μ_r and standard deviation σ_r. μ_r) / σ_r, thus replacing the independent value function with the relative advantage within the group; the policy update objective is max_θ E_i [ A_i · log π_θ(o_i) ] β · KL(π_θ ∥ π_ref ), where β is the KL regularization coefficient. The GRPO stage provides sampling as a service through an independent inference engine (e.g., vLLM), and the training process obtains the rollout results through an interface, avoiding contention between training memory and sampling memory.
[0069] like Figure 7 As shown, the rule-based combination reward is calculated for each candidate output using the following formula: reward = r_schema × ( w_title · r_chunk_title + w_code · r_chunk_code + w_ndcg · r_ndcg + w_cov · r_coverage ) / ( w_title + w_code + w_ndcg+ w_cov ) Where: r_schema∈{0,1} is the schema validity hard gate, which is 1 only if the output can be parsed into valid JSON, the fields satisfy the predefined data schema, all element references exist in the set of valid element identifiers, and all models exist in the model list; otherwise, it is 0. r_chunk_title is the micro F1 of the (element, module title) binary pair. r_chunk_code is the micro F1 of the (element, model) binary pair. For scenarios where an element belongs to multiple models, element-wise Jaccard averaging can also be used as an alternative implementation. r_ndcg is the NDCG of the element order under each (model, module title) and the gold standard order, and the arithmetic mean is taken over all (model, module title). r_coverage is the coverage ratio of the set of models with non-empty content in the prediction to the corresponding set of the gold standard. w_title, w_code, w_ndcg, and w_cov are preset weights, which can be, for example, 3.0, 2.0, 2.0, and 1.0, respectively.
[0070] It should be noted that the aforementioned guided decoding operates in the inference phase, while the structural legality hard gate constraint operates in the training phase. The two constrain the model inference process and the model training process, respectively.
[0071] The reward function is reused in both the reinforcement learning and final evaluation phases in phase two, thus ensuring that the training objective and the online evaluation metric are consistent. Before GRPO training, SFT checkpoints can be sampled and the mean and standard deviation of the reward can be calculated. If the within-group variance is too small, the GRPO phase can be skipped to avoid insufficient training gains.
[0072] By employing a rule-based combined reward system that does not rely on an independent reward model, specific business rules are directly transformed into deterministic offline reward functions. This design eliminates the computational overhead of additionally training and maintaining neural network reward models required in traditional reinforcement learning. It fundamentally avoids the scoring illusion, logical flaws, and biases that may exist in neural network reward models themselves, enabling the policy model to receive stable, objective, predictable, and noise-free feedback signals, significantly improving the stability and controllability of training.
[0073] Meanwhile, a "one-vote veto" hard-gating mechanism is constructed by multiplying the structural legality hard-gating constraint with multi-dimensional evaluation metrics. If a candidate output has corrupted JSON format or contains false element references or model numbers outside the preset legal set, its final reward value will be directly set to zero. This strong constraint mechanism guides the model to quickly abandon invalid decoding paths, focusing the limited search space on exploring legal formats and ensuring the absolute parsability of the final generated result. Introducing two micro-level F1 scores and element-wise Jaccard averaging allows for high-precision dual alignment measurements of the correspondence between "element-module title" and "element-model number" from both macro-level attribution and micro-level overlap perspectives.
[0074] Furthermore, by utilizing the Normalized Diminished Cumulative Gain (NDCG) of the element order under each model and module title as a ranking metric, the positional penalty characteristic is fully leveraged to rigorously audit the element order in the layout sequence, significantly improving the visual and logical readability of the final generated details page. Introducing a model coverage metric with non-empty content effectively penalizes the model's "specialization" behavior—selectively ignoring certain models in pursuit of local high scores when facing long-tailed models or complex samples. This constraint plays a crucial role in addressing the pattern collapse problem in reinforcement learning, forcing the model to consider the layout completeness of all input models, thereby significantly improving the generation completeness rate of product details pages in scenarios with multiple models coexisting.
[0075] <Electronic Devices> In other embodiments, this application also provides an electronic device. The electronic device includes a processor and a memory. The memory stores program instructions, which, when executed by the processor, implement the product details page generation method or the model training method described in the foregoing embodiments.
[0076] The processor is the core of electronic devices, serving as both the computing and control center. It drives hardware resources to perform logical operations and data interactions by reading program instructions from memory.
[0077] Memory serves as the physical storage medium for program instructions and data. For example, during the execution of a method for generating a product details page, memory can be used to temporarily store multimodal data acquired in the data acquisition step and the layout structure in the end-to-end inference step. It should be noted that memory can include not only volatile memory such as random access memory, but also read-only memory, flash memory, solid-state drives, optical disc storage, or highly available network-attached storage devices.
[0078] By coordinating the processor and memory, the aforementioned product detail page generation method or model training method can be implemented, which can automatically complete the process of parsing product manual content, generating layout structure, and constructing product detail pages. It can also obtain an end-to-end visual-language model with high layout structure prediction accuracy and strong generalization ability, thereby improving the accuracy, completeness, consistency, and generation efficiency of automatic product detail page generation, reducing manual layout workload, lowering system maintenance costs, and improving the automation level and stability of the product detail page production process.
[0079] It should be noted that, in addition to being a data center server or computing cluster deployed in the cloud, this electronic device can also be a personal computer, workstation, portable laptop, or an embedded control motherboard integrated into a specific industry terminal.
[0080] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for generating a product details page, characterized in that, include: The data acquisition steps include obtaining a list of target product models, determining the page rendering image and a set of element fragments with identifiers corresponding to the target product based on the multimodal data of the target product, and determining the module title according to the page template corresponding to the target product. The end-to-end inference step involves inputting the page rendering image, the set of element fragments, the module title, and the model list into the end-to-end visual-language model for inference, and outputting the layout structure of the target product. The layout structure includes the mapping relationship between the models in the model list and the module title and element references, and the element references are composed of the identifiers. The product details page generation steps involve extracting original content from the set of element fragments based on the element references in the layout structure and backfilling it into the page template to generate the product details page for the target product.
2. The method for generating a product details page according to claim 1, characterized in that, In the data acquisition step, the multimodal data includes the instruction manual. Determining the page rendering image and the set of element fragments specifically includes: Obtain the instruction manual for the target product, and render the instruction manual page by page into multiple page rendering images; The instruction manual is divided into multiple segments by page layout analysis. These segments specifically include text extraction results and image elements represented by storage paths. The element fragment set is constructed based on the multiple fragments obtained from the splitting, and the identifier is assigned to each fragment in the element fragment set; In the details page generation step, extracting the original content and backfilling it into the page template specifically includes: For segments containing the extracted text results, the extracted text results are directly populated back into the page template; For a fragment containing the image element, the corresponding image file is retrieved according to the storage path and then backfilled into the page template.
3. The method for generating a product details page according to claim 2, characterized in that, Determining the page rendering image and the set of element fragments also includes: Obtain the fragment metadata of each fragment in the element fragment set, wherein the fragment metadata includes the identifier, page number, fragment type, and bounding box; Based on the page number, the page rendering image corresponding to the fragment metadata is determined, and the fragment metadata is associated with the corresponding page rendering image to assist the end-to-end visual-language model in reasoning.
4. The method for generating a product details page according to claim 1, characterized in that, In the end-to-end inference step, the layout structure is JSON data, the top-level key of the JSON data is the model in the model list, and the value corresponding to the top-level key is the mapping between the module title and the element reference.
5. The method for generating a product details page according to claim 1, characterized in that, The end-to-end inference step includes: Guided decoding is enabled during the inference process of the end-to-end vision-language model; The guided decoding method prunes invalid decoding paths to constrain the output format of the layout structure.
6. A model training method, characterized in that, Applied to end-to-end vision-language models, the model training method includes: The step of obtaining training samples involves obtaining original training samples, which include input prompts and a complete gold standard layout consisting of complete gold standard elements. The input prompts include a page rendering image, a set of element fragments with identifiers, a module title, and a model list. The training sample enhancement step involves extracting a subset of page numbers from the set of page numbers in the page rendering image, retaining only the complete gold standard elements that fall within the subset of page numbers to generate the gold standard layout corresponding to the subset of page numbers, thereby obtaining candidate samples, and removing candidate samples that do not meet the preset validity conditions to obtain multiple training samples. The supervised fine-tuning step uses the gold standard layout corresponding to the page number subset as the supervision target, and uses the training samples to fine-tune the parameters of the initial model to obtain the supervised fine-tuning model. The group relative policy optimization step involves initializing the policy model and the reference model using the supervised fine-tuning model, sampling multiple candidate outputs for the same input prompt using the policy model, calculating the reward for each candidate output, updating the policy model based on the relative advantage within the group, and limiting the deviation between the policy model and the reference model through relative divergence constraints to obtain the end-to-end vision-language model.
7. The model training method according to claim 6, characterized in that, The supervised fine-tuning steps include: Freeze the visual encoder backbone of the initial model; By employing a low-rank adaptation approach, the model parameters of the initial model are updated only through backpropagation of cross-entropy loss to the multimodal projection layer and the language model, thus obtaining the supervised fine-tuning model.
8. The model training method according to claim 6, characterized in that, In the training sample enhancement step, the candidate samples that do not meet the preset validity conditions are: candidate samples whose number of non-empty modules is less than a preset threshold or that do not cover any model in the model list; In the group relative policy optimization step, calculating the reward for each candidate output includes: The reward value is calculated through a preset rule-based combined reward function, which is constructed by multiplying the weighted sum of structural legality hard gate constraints and alignment evaluation indicators of multiple dimensions. The structural legality hard gate constraint is configured as follows: when the candidate output simultaneously satisfies the following conditions: it can be parsed into a valid JSON format, the fields satisfy the predefined data pattern schema, it contains a valid identifier, and the included model exists in the model list, the value is 1; otherwise, the value is 0, so as to veto invalid decoding paths through the structural legality hard gate constraint.
9. The model training method according to claim 8, characterized in that, The alignment evaluation metrics across multiple dimensions include at least one of the following: A micro F1 score used to measure the correspondence between elements and the module title; Microscopic F1 score or element-wise Jaccard average used to measure the correspondence between an element and the model number; Normalized cumulative depreciation revenue (NDCG) is used to audit the order of elements under the titles of various models and modules. The model coverage index is used to characterize the coverage ratio between the set of models with non-empty content in the prediction output and the set of models corresponding to the gold standard layout corresponding to the page number subset.
10. An electronic device, characterized in that, It includes a processor and a memory, the memory storing program instructions, which, when executed, implement the product details page generation method according to any one of claims 1-5 or the model training method according to any one of claims 6-9.