Method and apparatus for localizing cross-border commodity data based on multimodal large model
By collecting and processing product data in cross-border e-commerce using a multimodal large model, generating neutral product semantic representations and constructing localized strategy graphs, the complexity and compliance issues of data localization in cross-border e-commerce are resolved, achieving efficient and accurate localization processing of product data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QIFA SILK ROAD (BEIJING) TECHNOLOGY DEVELOPMENT CO LTD
- Filing Date
- 2025-12-12
- Publication Date
- 2026-07-17
AI Technical Summary
In cross-border e-commerce, the lack of unified multimodal semantic modeling in existing technologies leads to a complex process of localizing product data, inconsistent management across multiple countries and platforms, insufficient utilization of multimodal information, inadequate ability to identify compliance elements, and a lack of systematic verification mechanisms, making it difficult to efficiently complete data localization processing under the premise of compliance.
A multimodal large model is used to collect product text, images, and instruction manual information in a local computing environment, generate multimodal feature representations, standardize and unify them through a neutral product attribute model, and construct a localized strategy graph for generation and verification to ensure that the data complies with the specifications of the target country and platform.
It improves the accuracy and consistency of cross-border commodity data localization, enhances multimodal compliance constraints and automatic verification capabilities, reduces manual intervention, and improves processing efficiency.
Smart Images

Figure CN121808668B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and apparatus for localizing cross-border commodity data based on a multimodal large model. Background Technology
[0002] In cross-border e-commerce, platforms need to convert product data from different countries and supply chain systems into localized product information that conforms to the target country's language, units of measurement, and listing standards. In practice, product data includes not only titles, text descriptions, and parameter tables, but also multimodal materials such as product images, packaging labels, certificates of conformity, and instruction manuals. Furthermore, it is subject to the regulatory requirements of various countries regarding local data storage, data export, and product compliance labeling.
[0003] In existing technologies, cross-border commodity data localization typically employs a "rule configuration + machine translation + human proofreading" approach. This involves multilingual translation of product titles and descriptions, and simple conversion of a small number of structured fields based on mapping tables. Some solutions introduce pre-trained language models to extract attributes and generate copy, but these are mostly single-modal text processing methods. Images, scanned documents, and other information are only used for display purposes, failing to establish a unified multimodal semantic model. Furthermore, given the constraints of national regulations and platform field specifications, existing systems largely rely on fragmented rule engines and manual configuration based on experience, lacking a unified structured expression and a computable constraint model.
[0004] Against this technological backdrop, the existing technologies suffer from several key problems: First, the localization process relies primarily on translation and field mapping, lacking a unified, neutral product semantic layer for different countries and platforms, leading to complex multi-country, multi-platform management and inconsistent fields. Second, they fail to fully utilize key information in product images, packaging labels, and instruction manuals, limiting their ability to identify compliance elements such as hazardous attributes and certification marks. Third, various regulatory rules and platform specifications lack a unified structured constraint model, making it difficult to systematically verify the consistency and rules of the localization results. Fourth, under the requirements of data localization and cross-border data export compliance, existing solutions lack a technical path that balances local processing and intelligent generation capabilities, making it difficult to efficiently complete cross-border product data localization while ensuring compliance. Summary of the Invention
[0005] In view of this, embodiments of this application provide a method and apparatus for localizing cross-border commodity data based on a multimodal large model, in order to solve the problems of lack of unified neutral commodity semantic expression, insufficient utilization of multimodal commodity compliance elements, and lack of localized rule constraints and verification mechanisms in the prior art.
[0006] A first aspect of this application provides a method for localizing cross-border commodity data based on a multimodal large model, comprising: collecting and preprocessing commodity text information, commodity image information, and commodity description document information associated with the target cross-border commodity in a local computing environment corresponding to the country to which the cross-border commodity belongs, to obtain a multimodal data set; inputting the multimodal data set into a multimodal large model to obtain a multimodal feature representation for the target cross-border commodity, and extracting candidate commodity attributes; based on a preset neutral commodity attribute pattern, standardizing the attribute names, unifying the units, and normalizing the values of the candidate commodity attributes to generate a neutral commodity semantic representation independent of the platform and country; and based on the target sales country and target sales platform... Based on the rule information, a localization strategy graph is constructed to define the target localized product attribute items, the dependencies between them, and the value constraints. Under the constraints of the localization strategy graph, the neutral product semantic representation is input into the multimodal large model for localization generation. According to the attribute filling order and value constraints determined by the localization strategy graph, localized product attribute data and localized product description information that conform to the language expression of the target sales country and the field specifications of the target sales platform are obtained. Based on the multimodal feature representation and the localization strategy graph, consistency and rule verification are performed on the localized product attribute data. According to the verification results, localized product attribute data that does not meet the preset conditions are corrected to obtain localized product data.
[0007] A second aspect of this application provides a cross-border commodity data localization processing device based on a multimodal large model, comprising: a collection module, used to collect and preprocess commodity text information, commodity image information, and commodity description document information associated with the target cross-border commodity in a local computing environment corresponding to the country to which the cross-border commodity belongs, to obtain a multimodal data set; an extraction module, used to input the multimodal data set into a multimodal large model to obtain a multimodal feature representation for the target cross-border commodity, and extract candidate commodity attributes; a standardization module, used to standardize the attribute names, unify the units, and normalize the values of the candidate commodity attributes based on a preset neutral commodity attribute pattern, to generate a neutral commodity semantic representation independent of the platform and country; and a construction module, used to construct the data based on the target sales country. Based on the rules and information of the target sales platform, a localization strategy graph is constructed to define the target localized product attribute items, the dependencies between these attributes, and the value constraints. A generation module, under the constraints of the localization strategy graph, inputs the neutral product semantic representation into a multimodal large model for localization generation. Following the attribute filling order and value constraints determined by the localization strategy graph, it obtains localized product attribute data and localized product description information that conform to the language of the target sales country and the field specifications of the target sales platform. A validation module, based on the multimodal feature representation and the localization strategy graph, performs consistency and rule validation on the localized product attribute data. Based on the validation results, it corrects localized product attribute data that does not meet preset conditions, thus obtaining the localized product data.
[0008] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0009] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0010] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: By collecting and preprocessing product text information, product image information, and product description document information associated with the target cross-border product in a local computing environment corresponding to the country of origin of the cross-border product, a multimodal dataset is obtained. This multimodal dataset is then input into a large multimodal model to obtain multimodal feature representations for the target cross-border product and extract candidate product attributes. Based on a pre-defined neutral product attribute pattern, the candidate product attributes are normalized in terms of attribute names, units, and values, generating a neutral product semantic representation independent of the platform and country. Finally, based on the rules of the target sales country and the target sales platform, a limited localized product attribute item is constructed. This application presents a localization strategy graph defining the dependencies and value constraints between target localized product attribute items. Under the constraints of the localization strategy graph, neutral product semantic representations are input into a multimodal large model for localization generation. Following the attribute filling order and value constraints determined by the localization strategy graph, localized product attribute data and localized product description information conforming to the target sales country's language and the target sales platform's field specifications are obtained. Based on the multimodal feature representation and the localization strategy graph, consistency and rule checks are performed on the localized product attribute data. Localized product attribute data that does not meet preset conditions is corrected according to the check results, resulting in localized product data. This application can improve localization accuracy and consistency, and enhance multimodal compliance constraints and automatic verification capabilities. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating the cross-border commodity data localization processing method based on a multimodal large model provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of the cross-border commodity data localization processing device based on a multimodal large model provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0014] In existing technologies, cross-border e-commerce platforms typically use a combination of rule configuration, machine translation, and human proofreading to translate product titles and descriptions into multiple languages, and then use simple field mapping to localize product data. Some solutions introduce pre-trained language models to extract attributes and generate copy for product text, but overall, the approach remains primarily text-based, with multimodal materials such as product images, packaging labels, and instruction manuals mainly used for display, lacking a unified multimodal semantic modeling approach. Furthermore, to address the constraints of different national regulatory requirements and platform field specifications, the approach relies heavily on fragmented rule engines and human experience, lacking a unified structured expression and a computable constraint model.
[0015] Against this backdrop, existing technologies suffer from the following main technical problems: First, the lack of a neutral product semantic expression layer independent of platforms and countries makes it difficult to uniformly manage and reuse data for the same product across multiple countries and platforms; second, insufficient utilization of multimodal information in product images, packaging labels, and instruction manuals, resulting in inadequate identification and support for compliance elements such as hazardous attributes and certification marks; third, the lack of localized rule constraints and systematic verification mechanisms that abstract regulatory rules and platform field specifications into computable constraint models, making it difficult to promptly detect and correct inconsistencies and non-compliance issues in localization results; and fourth, the lack of a technical path that balances local processing capabilities with intelligent generation capabilities under the constraints of data localization and cross-border data export regulations.
[0016] To address the aforementioned technical issues, this application proposes a method for localizing cross-border commodity data based on a multimodal large model. This method involves collecting and preprocessing commodity text, image, and documentation information associated with the target cross-border commodity in a local computing environment to form a multimodal dataset. This dataset is then input into a multimodal large model to obtain multimodal feature representations and extract candidate commodity attributes. Based on a predefined neutral commodity attribute pattern, the candidate commodity attributes are standardized in terms of attribute names, units, and values, constructing a neutral commodity semantic representation independent of the platform and country. According to the rules of the target sales country and platform, a localization strategy graph is constructed, defining the target localized commodity attribute items, their dependencies, and value constraints. Under the constraints of this strategy graph, the multimodal large model performs localization generation, resulting in localized commodity attribute data and descriptions that conform to the language of the target sales country and the field specifications of the target sales platform. Furthermore, consistency and rule checks are performed on the localized commodity attribute data based on the multimodal feature representation and the localization strategy graph, correcting any localized commodity attribute data that does not meet the predefined conditions.
[0017] Through the above technical solutions, this application can construct a unified neutral semantic representation of commodities based on multi-source and multi-modal commodity data, and introduce a localization strategy graph to achieve constrained control and verification correction of the localization generation process, improve the accuracy and consistency of cross-border commodity data localization processing, enhance the utilization and automatic verification capabilities of multi-modal compliance elements, reduce manual intervention, and improve the automation and processing efficiency of cross-border commodity data localization processing while meeting data localization and compliance requirements.
[0018] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0019] Figure 1 This is a flowchart illustrating the cross-border commodity data localization processing method based on a multimodal large model provided in this application. Figure 1 As shown, this method for localizing cross-border commodity data based on a multimodal large model can specifically include: S101. In the local computing environment corresponding to the country to which the cross-border goods belong, collect and preprocess the product text information, product image information and product description document information associated with the target cross-border goods to obtain a multimodal data set; S102, input the multimodal dataset into the multimodal large model to obtain multimodal feature representations for the target cross-border goods and extract candidate product attributes; S103, based on the preset neutral product attribute pattern, standardizes the attribute names, unifies the units and normalizes the values of candidate product attributes, and generates a neutral product semantic representation that is independent of the platform and country; S104, Based on the rule information of the target sales country and the target sales platform, construct a localization strategy graph that limits the target localized product attribute items and the dependencies and value constraints between the target localized product attribute items; S105, Under the constraints of the localization strategy graph, the neutral product semantic representation is input into the multimodal large model for localization generation. According to the attribute filling order and value constraints determined by the localization strategy graph, localized product attribute data and localized product description information that conform to the language expression of the target sales country and the field specifications of the target sales platform are obtained. S106, based on multimodal feature representation and localization strategy graph, performs consistency and rule checks on localized product attribute data, and corrects localized product attribute data that does not meet preset conditions according to the check results, thus obtaining localized product data.
[0020] In some embodiments, product text information, product image information, and product description document information associated with the target cross-border product are collected and preprocessed to obtain a multimodal dataset, including: In the local computing environment, retrieve the original product text, original product images, and original product description electronic documents associated with the target cross-border product from business systems, production systems, and file storage systems; Perform layout structure analysis and optical character recognition on the original electronic product description document to extract paragraphs, headings, table cells and text, and generate document text fragments carrying layout position information; Perform product and label region detection and text recognition on the original product image to extract image text fragments carrying spatial location information and product appearance features; Sensitive fields are identified and desensitized based on pre-trained semantic models for original product text, document text fragments, and image text fragments. Based on the internal identifiers of the target cross-border goods, the anonymized product text information, product image information, and product description document information are uniformly numbered and associated with an index to construct a multimodal data set that corresponds one-to-one with the target cross-border goods.
[0021] Specifically, the cross-border commodity data localization processing system involved in this solution is deployed in a local data center corresponding to the country to which the target cross-border commodity belongs. The system includes a data acquisition submodule, a document parsing submodule, an image parsing submodule, a sensitive field desensitization submodule, and a multimodal indexing and association submodule. Each submodule is called through a service interface in the same intranet environment to ensure that the data is processed locally.
[0022] First, for each target cross-border product, the data acquisition submodule retrieves the original product text, original product images, and original product description electronic documents associated with that product from the business system, production system, and file storage system, based on the product's unique identifier within the enterprise. The business system can be a product management system for managing basic product information and sales attributes; the production system can be a production execution system for managing process parameters and bills of materials; and the file storage system can be a file server or document management system storing instruction manuals (PDF), test reports, and scanned copies of certificates of conformity. The data acquisition submodule pre-establishes a mapping relationship between internal identifiers and resource paths in each system through interface configuration. Upon receiving the internal identifier of the target cross-border product, it batch-retrieves multi-source data, including the original product text, main and detail images, packaging label images, and instruction manual (PDF) files, according to this mapping relationship.
[0023] Furthermore, the document parsing submodule performs layout structure analysis and optical character recognition on the original product description electronic document. In this embodiment, the document parsing submodule performs layout structure analysis on the instruction manual PDF, identifying the title area, body paragraph area, illustration area, and table area on the page, and assigning corresponding layout structure tags to each area. For the table area, it is further subdivided into table rows and table cells, and the position rectangle of each cell in the page coordinate system is determined.
[0024] Subsequently, the document parsing submodule, combined with a multilingual optical character recognition model, identifies the text content within each area, generating document text fragments containing text content, layout structure labels, and location coordinates. Examples include a "Technical Parameters" title fragment and parameter cell fragments such as "Rated Voltage 220V" and "Rated Power 1500W." In this way, each document text fragment carries not only the text content but also its layout location information within the source documentation, facilitating the subsequent establishment of precise correspondences with image information and attribute fields.
[0025] Furthermore, the image parsing submodule is used to perform product and label region detection and text recognition on the original product image. In some examples, the image parsing submodule first uses an object detection model to divide the main product image and packaging image into regions, locating the main product region, brand logo region, certification mark region, warning label region, etc., and outputting the corresponding bounding box coordinates. For the detected label region and packaging region containing text, the image parsing submodule further calls the image text recognition model to recognize the text information, obtaining image text fragments carrying spatial location information, such as recognizing text symbols such as "Suitable water volume 1.5L", "Suitable voltage 220V~", and "Indoor use only" from the label region on the side of the packaging. For partial images of product nameplates, image text fragments such as model number, production batch, and execution standard number, as well as the corresponding product appearance feature vectors, can also be extracted for subsequent multimodal alignment.
[0026] Furthermore, the sensitive field desensitization submodule is used to identify and process sensitive fields in the original data without destroying the semantic information of the product. In some examples, the sensitive field desensitization submodule, based on a pre-trained semantic model, performs semantic classification and entity recognition on the original product text, document text fragments, and image text fragments. It extracts fields marked as sensitive entities, such as supplier internal codes, distributor names, contact information, and internal cost prices, and replaces the corresponding text content with placeholders or performs partial masking, for example, retaining only the brand and model while hiding the internal material code; retaining only standardized specification information while hiding the internal purchase price. For fields such as internal QR codes and internal barcodes contained in image text fragments, the sensitive field desensitization submodule can prevent sensitive information from being leaked in subsequent processing steps by masking the image area corresponding to the coordinates of the masked area or replacing the recognized text. After sensitive field desensitization, desensitized product text information, desensitized document text fragments, and desensitized image text fragments are generated.
[0027] The multimodal indexing and association submodule is used to establish a one-to-one correspondence between the aforementioned anonymized multimodal information and the target cross-border goods. In some examples, the multimodal indexing and association submodule uses the internal identifier of the target cross-border goods as the index key, uniformly numbers the anonymized product text information, the anonymized product image and its spatial annotation information, and the anonymized document text fragments and their page layout information, and writes them into the multimodal index library.
[0028] For text from different sources of the same product, such as product titles in the business system, process descriptions in the production system, and technical parameter descriptions in the instruction manual, the multimodal index and associated submodules are distinguished by source tags and version tags, and aggregated and stored using internal identifiers as the primary key. For multiple images of the same product and their corresponding image text fragments, a hierarchical index structure is also established from product to image to local text using image number and region number.
[0029] Through the above processing, a multimodal data set corresponding one-to-one with the target cross-border commodity is formed. This multimodal data set includes multi-source multimodal data and its structured location information, including commodity text information, commodity image information, and commodity description document information.
[0030] In some embodiments, a multimodal dataset is input into a large multimodal model to obtain a multimodal feature representation for the target cross-border commodity, and candidate commodity attributes are extracted, including: The product text information, product image information, and product description document information in the multimodal dataset are respectively input into the text encoding substructure, image encoding substructure, and document encoding substructure in the multimodal large model to obtain the corresponding text feature representation, image feature representation, and document feature representation; Based on the cross-modal alignment structure, text feature representation, image feature representation and document feature representation are aligned and fused in a unified semantic space to generate a commodity-level multimodal feature representation corresponding to the target cross-border commodity; Under the constraint of a pre-defined set of attribute slots, the cross-modal attention aggregation mechanism guided by attribute slots is used to aggregate the corresponding attribute feature vectors for each attribute slot from the product-level multimodal feature representation. Based on the attribute discrimination substructure, the attribute category is determined and the attribute value is inferred for each attribute feature vector, and the output is a set of candidate product attributes containing candidate attribute names and candidate attribute values.
[0031] Specifically, the multimodal large model is deployed in the model service layer of the cross-border commodity data localization processing system. This multimodal large model includes a text encoding substructure, an image encoding substructure, a document structure encoding substructure, a cross-modal alignment structure, an attribute slot-guided cross-modal attention aggregation structure, and an attribute discrimination substructure. The aforementioned multimodal dataset is fed into this multimodal large model for unified processing through a service interface. This embodiment uses an electric kettle as an example to illustrate the process of extracting candidate commodity attributes.
[0032] In some examples, for the product text information in the aforementioned multimodal dataset, the text encoding substructure segments the product title, selling point description, parameter text, and document text fragments obtained from the instruction manual into sentences and words, maps each text sequence into word-level or sub-word-level vector sequences, and obtains the encoded text feature representation through a multi-layer self-attention network, such as obtaining text feature vectors representing semantic units such as "capacity 1.5L", "rated voltage 220V", and "food contact parts are stainless steel".
[0033] For product image information in a multimodal dataset, the image coding substructure extracts features from the main product image, packaging image, and partial image of the nameplate, divides the image into several image blocks, and generates an image feature vector for each image block. At the same time, it maps the spatial location information of the label area, certification mark area, and image text fragments obtained in the aforementioned image parsing stage to the corresponding image feature vector.
[0034] For product description documents, the document structure coding substructure treats paragraphs, headings, and table cells as document structure nodes, and performs joint coding on the text content, layout information, and row and column indexes of each node to obtain a document feature representation that characterizes the document hierarchy.
[0035] Furthermore, after obtaining text feature representations, image feature representations, and document feature representations, the cross-modal alignment structure maps these three types of features to a unified semantic space. In some examples, the cross-modal alignment structure first introduces a shared product semantic embedding space, performing linear transformations and normalization on the text feature representations, image feature representations, and document feature representations respectively, giving them a unified vector dimension and scale. Subsequently, through a cross-modal attention mechanism, using text fragments as queries, attention aggregation is performed on relevant image block features and document cell features to obtain aligned multimodal semantic units. For example, the text "capacity 1.5L" is aligned with the label area marked "1.5L" in the packaging image and the "rated capacity 1.5L" cell in the instruction manual parameter table within the semantic space.
[0036] In this way, the cross-modal alignment structure generates a product-level multimodal feature representation corresponding to the target cross-border product. This product-level multimodal feature representation not only includes the global product semantic vector, but also includes several aligned multimodal semantic units.
[0037] Building upon this, an attribute slot-guided cross-modal attention aggregation mechanism is used to aggregate corresponding attribute feature vectors from the product-level multimodal feature representation for a preset set of attribute slots. In this embodiment, the preset set of attribute slots includes attribute slots such as brand, model, product type, capacity, rated voltage, rated power, main material, applicable power frequency, applicable scenario, and certification mark. For each attribute slot, the system assigns a trainable attribute slot vector for initiating a query in the multimodal semantic space.
[0038] The attribute slot-guided cross-modal attention aggregation mechanism uses the attribute slot vector as the query vector and the multimodal semantic unit vectors in the product-level multimodal feature representation as keys and values. Through attention weight calculation, it selects the semantic units most relevant to the current attribute slot and performs weighted aggregation to obtain the attribute feature vector for that attribute slot. For example, for the "capacity" attribute slot, the attention aggregation process focuses on semantic units corresponding to text fragments related to capacity, packaging label image regions, and parameter table cells, thereby generating an attribute feature vector that comprehensively reflects capacity information.
[0039] Subsequently, the attribute discrimination substructure performs attribute category determination and attribute value inference on each attribute feature vector. In some examples, the attribute discrimination substructure includes an attribute category classification header and an attribute value inference header. The attribute category classification header is used to determine whether the current attribute slot is valid for the target product; for example, the "certification mark" attribute slot may not exist for some products. The attribute value inference header infers the attribute value based on the attribute slot type using either classification or regression. For discrete attributes such as brand, product type, and certification mark, multi-class classification is used to output candidate attribute values and calculate confidence scores. For numerical attributes such as capacity, rated voltage, and rated power, numerical results are output, which may include units or dimensions. The attribute discrimination substructure also records the multimodal evidence indication information corresponding to each candidate attribute value, such as recording which text segments, image regions, or document cells the value mainly originates from, for subsequent construction of neutral product semantic representations and consistency verification.
[0040] In some embodiments, based on a preset neutral product attribute pattern, the candidate product attributes are normalized in terms of attribute names, units, and values to generate a neutral product semantic representation that is independent of the platform and country, including: A neutral product attribute graph is constructed based on the neutral product attribute model. Neutral product attribute nodes are established for each neutral product attribute in the neutral product attribute graph, and the attribute synonym name nodes, attribute unit nodes, and attribute value range nodes corresponding to each neutral product attribute are associated. The attribute names and attribute values of candidate product attributes are mapped to the neutral product attribute graph. Based on graph embedding encoding and similarity measurement, the target neutral product attribute node and standard attribute name corresponding to each candidate product attribute are determined, and the attribute names of the candidate product attributes are normalized. Based on the unit conversion rules associated with each neutral commodity attribute node, the values of numerical attributes in the candidate commodity attributes are identified by dimension and converted by unit, and the physical quantities of the candidate commodity attributes are converted into the standard units specified in the neutral commodity attribute pattern. Based on the set of value specifications associated with each neutral product attribute node, the discrete and interval values of the candidate product attributes are merged and discretized to obtain standardized attribute values that conform to the value constraints of the neutral product attribute pattern. According to the field definitions in the neutral commodity attribute model, the standardized standard attribute name, standard unit, and standardized attribute value are associated with the evidence indication information of the candidate commodity attributes to generate a neutral commodity semantic representation for the target cross-border commodity.
[0041] Specifically, after extracting candidate product attributes, the cross-border commodity data localization processing system further standardizes the names, unifies units, and normalizes values of the candidate product attributes in the attribute normalization subsystem to generate a neutral product semantic representation independent of the platform and country. This attribute normalization subsystem includes a neutral product attribute graph management module, an attribute mapping and name normalization module, a unit conversion module, a value normalization and discretization module, and a neutral semantic representation generation module. These modules work collaboratively around a pre-defined neutral product attribute pattern.
[0042] In the neutral product attribute map management module, the system first constructs a neutral product attribute map based on a preset neutral product attribute pattern. The neutral product attribute pattern defines a set of standard attribute fields that are independent of the business platform and the country of sale, such as brand, neutral product type, rated capacity, rated voltage, rated power, main material, applicable age range, energy efficiency rating, hazard attribute label, and certification label set.
[0043] The system establishes neutral product attribute nodes in the graph, centered around each standard attribute field. Under each neutral product attribute node, multiple attribute synonym node, attribute unit node, and attribute value range node are associated. For example, for the "rated capacity" attribute node, associated attribute synonym nodes include "capacity," "nominal capacity," and "maximum capacity"; associated attribute unit nodes include "liter," "milliliters," and "ounce"; and associated attribute value range nodes describe the reasonable range of this attribute across different product types. The neutral product attribute graph can be gradually expanded through a combination of manual configuration and incremental learning. In the graph, relationships such as "synonyms," "unit convertibility," and "value constraints" are represented by edges.
[0044] Furthermore, in the attribute mapping and name normalization module, the system maps the attribute names and values in the candidate product attribute set to the neutral product attribute graph. Taking an electric kettle as an example, candidate product attributes may include different expressions such as "capacity: 1.5L" and "maximum water volume: 1.5 liters". The system first performs graph embedding encoding on the candidate attribute names and attribute values, mapping the attribute names and attribute values to the same embedding space as the neutral product attribute graph.
[0045] Then, based on a similarity measurement mechanism, the module calculates the similarity between the candidate attribute name and the synonymous name nodes of each attribute in the graph. Simultaneously, it combines the matching degree between the candidate attribute value and the attribute value category nodes to comprehensively determine the target neutral product attribute node corresponding to each candidate product attribute. For example, when the candidate attribute name is "maximum water volume" and the value is "1.5L", based on graph embedding similarity and category constraints, this candidate product attribute should be mapped to the "rated capacity" neutral product attribute node. After mapping, the system replaces the candidate attribute name with the standard attribute name corresponding to the target neutral product attribute node, thereby achieving attribute name standardization.
[0046] Furthermore, in the unit conversion module, the system performs dimensional identification and unit conversion for numerical attributes in candidate product attributes. Continuing with the example of "rated capacity," candidate attribute values may appear as "1.5L," "1500 ml," or "50.7 ounces." The system first identifies the physical quantity and unit type of the attribute value based on the text pattern, contextual information, and the matching relationship with the attribute unit node. Then, it queries the unit conversion rules associated with the corresponding neutral product attribute node, converting values of different units to the standard units specified in the neutral product attribute pattern, such as uniformly converting them to "liters" and standardizing them as the numerical value "1.5." For attributes such as "rated voltage" and "rated power," the system also follows the corresponding unit conversion rules, uniformly converting different expressions such as "110V" and "220 volts" into the numerical representations corresponding to the standard units.
[0047] Furthermore, in the value merging and discretization module, the system performs merging and discretization processing on the discrete and interval values of candidate product attributes based on the value specification set associated with each neutral product attribute node. For example, for the "Applicable Age" attribute, candidate values may be inconsistent expressions such as "3 years and above," "3-12 years old," and "For Children." The system merges these into standardized interval values such as "3-12 years old" or standardized labels "Infants and Children" using the mapping rules in the value specification set. Similarly, for multi-select attributes such as "Hazardous Attribute Identification," the system merges different texts such as "Flammable," "Combustible," and "Flammable Material" into a unified standardized label set according to the value specification set. Through the above processing, standardized attribute values that conform to the value constraints of the neutral product attribute pattern are obtained.
[0048] Finally, in the neutral semantic representation generation module, the system, according to the field definitions in the neutral product attribute pattern, associates the normalized standard attribute name, standard unit, and standardized attribute value with the multimodal evidence indication information retained in the previous steps for the candidate product attributes, thus constructing a neutral product semantic representation for the target cross-border product. This neutral product semantic representation records the standard name, standard unit, and standardized value of each neutral product attribute in a structured form, and saves the evidence index between it and text fragments, image regions, or document cells, so as to trace the attribute source and supporting evidence in the subsequent localization generation and consistency verification stages.
[0049] Through the processing flow of this embodiment, the system uniformly maps candidate product attributes from different text sources and modalities to a neutral product attribute graph, achieving standardization of attribute names, unification of units, and normalization of values. This constructs a neutral product semantic representation that is independent of platforms and countries and has clear value constraints. On the one hand, this reduces the complexity caused by inconsistencies in attribute definitions and unit standards in multi-country, multi-platform scenarios; on the other hand, it provides a unified and standardized semantic foundation for subsequent attribute mapping and localization generation based on localization strategy graphs, which is conducive to improving the consistency and maintainability of cross-border product attribute management.
[0050] In some embodiments, based on rule information from the target sales country and the target sales platform, a localization strategy graph is constructed that defines the target localized product attribute items and the dependencies and value constraints between the target localized product attribute items, including: Obtain rule information associated with the target sales country and target sales platform from the cross-border product localization knowledge base. The rule information includes target platform field definitions, mandatory field conditions, field dependencies, and field value constraints. The field definitions in the rule information are abstracted into localized attribute nodes, and the mandatory dependencies, conditional dependencies, and conflict constraints between fields are abstracted into constraint edges connecting localized attribute nodes to construct an initial rule relationship graph. Based on the preset rule template, the rule information is structured and encoded, and the text-based regulatory clauses and platform configurations are converted into computable dependency expressions and value constraint expressions, and associated with the corresponding localized attribute nodes and constraint edges. By using a graph structure learning model combined with historical localized samples, the effective conditions and priorities of each localized attribute node and constraint edge in the initial rule relationship graph are updated to determine the set of activated localized attribute nodes and their constraint relationships for the target cross-border goods, thus obtaining a localized strategy subgraph. The localization strategy subgraph is topologically sorted to determine the filling order of the target localized product attribute items. The target localized product attribute items, their corresponding dependencies, and value constraints are then solidified in a graph structure to form a localization strategy graph used to constrain the localization generation process.
[0051] Specifically, after constructing a neutral semantic representation of goods, the cross-border commodity data localization processing system uses a rule modeling subsystem to generate a localized strategy graph for the target sales country and target sales platform. The rule modeling subsystem includes a cross-border commodity localization knowledge base, a rule parsing and structured coding module, a rule relationship graph construction module, a graph structure learning module, and a strategy graph generation module. These modules operate collaboratively around the combination of the target sales country and the target sales platform. This embodiment uses the example of listing an electric kettle on a specific e-commerce platform in a European country to illustrate the process of constructing the localized strategy graph.
[0052] The cross-border commodity localization knowledge base pre-stores information such as regulatory clauses, platform field configurations, and previous localization configuration records for different countries and platforms. When the system receives the target sales country and target sales platform identifiers, the rule parsing and structured coding module first retrieves rule information matching that country and platform from the cross-border commodity localization knowledge base. For electric kettles, the rule information includes the target platform's field definitions for field names, field data types, and display formats, as well as mandatory field selection conditions, field dependencies, and field value constraints such as "rated voltage must be the local voltage standard field is required," "energy efficiency rating must be indicated when power exceeds a certain threshold," and "safety warnings must be added for children under 3 years old." These rules initially exist mostly in text form and may originate from platform access documents, regulatory provisions, and configuration records from maintenance personnel.
[0053] The rule relationship graph construction module abstracts the field definitions in the above rule information into localized attribute nodes, and abstracts the mandatory dependencies, conditional dependencies, and conflict constraints between fields into constraint edges connecting the localized attribute nodes, thus constructing an initial rule relationship graph. For example, "localized rated voltage," "localized rated power," "localized energy efficiency rating," "hazard attribute label," and "child safety warning" are abstracted into multiple localized attribute nodes; at the same time, constraints such as "rated voltage field must exist," "energy efficiency rating field must be filled when rated power is greater than a certain threshold," and "hazard attribute label of flammable must not appear simultaneously with certain applicable scenarios" are abstracted into different types of constraint edges, where mandatory dependencies are represented by mandatory existence edges, conditional dependencies by edges with condition labels, and conflict constraints by mutually exclusive edges. In this way, the system constructs an initial rule relationship graph covering the localized fields and their constraint relationships related to the target country and the target platform.
[0054] Subsequently, the rule parsing and structured coding module further structures and encodes the rule information based on preset rule templates, converting the regulatory clauses and platform configurations originally described in natural language into computable dependency expressions and value constraint expressions. In some examples, the system predefines multiple rule template types, such as "Field A is mandatory in any scenario," "Field C is mandatory when Field A exists and Field B meets a certain value range," and "The value combinations of Field A and Field B cannot appear simultaneously." After parsing the regulatory clauses sentence by sentence, it matches them to the corresponding templates, generates logical expressions for the condition parts, generates constraint expressions for the value parts, and associates these expressions with the corresponding localized attribute nodes and constraint edges. For example, when the knowledge base contains the clause "Heating appliances sold in this country must be labeled with the local power grid standard voltage and frequency," the system parses it into mandatory constraint expressions for the "localized rated voltage" and "localized rated frequency" nodes and attaches them to the corresponding constraint edges.
[0055] Building upon this foundation, the graph structure learning module, incorporating historical localized samples, updates the parameters for the activation conditions and priorities of each localized attribute node and constraint edge in the initial rule relationship graph. Historical localized samples include localized field filling records and compliance review results for similar products successfully listed on the platform in the target country. By modeling the graph structure of these samples, the module statistically analyzes the frequency and importance of each constraint edge being triggered under different product categories and regulatory environments. It also learns a set of weight parameters on the graph nodes and constraint edges to characterize the rule's activation tendency in specific product scenarios.
[0056] For example, for electric kettles, the graph structure learning module might prioritize constraints related to "rated voltage," "rated power," "energy efficiency rating," and "hazard identification," while constraints specific to seasonal promotional fields might be less important. Through parameter updates, the system can determine the set of localized attribute nodes and their constraints that actually need to be activated for the current target cross-border product, and then extract a localized strategy subgraph related to that product from the initial rule graph.
[0057] Furthermore, after obtaining the localized strategy subgraph, the strategy graph generation module performs topological sorting on the subgraph to determine the filling order of the target localized product attribute items. During the topological sorting process, the dependencies expressed by constraint edges are considered, prioritizing the filling of localized attribute items that serve as prerequisites. For example, basic fields such as "localized product type" and "localized hazard attribute identifier" are filled first, followed by fields that depend on these fields, such as "localized applicable scenarios" and "localized warning messages." After completing the topological sorting, the system solidifies the ordered list of target localized product attribute items, the dependencies between attribute items, and the value constraint expressions into a graph structure, generating a localized strategy graph used to constrain the localization generation process. This localized strategy graph retains references to the original rule clauses and platform configurations during storage for subsequent traceability.
[0058] Through the processing flow of this embodiment, the system can unify and abstract rule information scattered across regulatory clauses, platform configuration documents, and experience configuration records into a graph structure representation. It then adaptively adjusts the rule activation conditions and priorities based on historical localization samples, thereby generating a localization strategy graph that matches the target cross-border product category, target sales country, and target sales platform. This localization strategy graph not only clarifies the localized product attribute items that need to be filled in and their dependency order, but also solidifies the value constraints in a computable form. This provides a structured rule foundation for constraint decoding and result verification in the subsequent multimodal large-scale model localization generation process, which is beneficial for improving the consistency and adaptability of rule execution during localization generation.
[0059] In some embodiments, a neutral product semantic representation is input into a multimodal large model for localization generation. Following the attribute filling order and value constraints determined by the localization strategy graph, localized product attribute data and localized product description information conforming to the language of the target sales country and the field specifications of the target sales platform are obtained, including: Based on the localization strategy graph, the target localized product attribute items are topologically sorted to determine the filling order of each target localized product attribute item and the set of value constraints corresponding to each target localized product attribute item. The neutral product semantic representation, the localized attribute nodes and constraint expressions in the localization strategy diagram, and the historical localization samples for the target sales country and target sales platform are combined to form a localization generation context, which is then input into the localization generation substructure in the multimodal large model. In the localization generation substructure, for the target localized product attribute item to be filled, the attribute slot control mechanism is used to select the neutral product attribute associated with the target localized product attribute item from the neutral product semantic representation, and combine the corresponding multimodal feature representation and value constraint set to generate the candidate localized attribute value of the target localized product attribute item. Under the constrained decoding mechanism, the candidate localized attribute values are subject to word masking and numerical validity checks. Decoding paths that do not meet the constraints are eliminated based on the set of value constraints. The localized attribute values of each target localized product attribute item are then determined, forming localized product attribute data. Using localized product attribute data as structured input, the multimodal large model is driven to generate localized product description information that is expressed in the language of the target sales country and conforms to the field specifications of the target sales platform.
[0060] Specifically, after completing the construction of neutral product semantic representation and the generation of localization strategy graph, the cross-border commodity data localization processing system uses a localization generation subsystem to perform the localization generation process for the target sales country and target sales platform. The localization generation subsystem is deployed in service nodes associated with the multimodal large model and includes modules for determining the attribute filling order, constructing the localization generation context, controlling attribute slot generation, constrained decoding, and generating localization descriptions. These modules collaborate around the localization strategy graph to map neutral product semantics to localized product attribute data and localized product description information. This embodiment still uses the example of an electric kettle being listed on a specific e-commerce platform in a European country for illustration.
[0061] The attribute filling order determination module first performs a topological sorting of the target localized product attribute items based on the aforementioned localization strategy graph. In this localization strategy graph, localized attribute nodes represent localized fields that the platform requires to be filled in, such as localized product type, localized brand name, localized rated voltage, localized rated power, localized capacity, localized applicable power supply system, localized applicable scenario, localized energy efficiency rating, localized hazard attribute label, localized child safety warning, etc.; constraint edges represent the dependency or mutual exclusion relationships between fields, and are accompanied by value constraint expressions.
[0062] The attribute population order determination module performs topological sorting on the localized attribute nodes according to the dependency direction of the constraint edges, prioritizing fields that serve as prerequisites for population. For example, localized product type and localized hazard attribute identifiers are populated first, followed by localized applicable scenarios and localized warning messages that depend on these fields. The result of the topological sorting is an ordered list of attributes, and simultaneously, a corresponding set of value constraints is generated for each target localized product attribute item. The set of value constraints includes the numerical range, the enumerated set of values, and conditional constraints related to other fields.
[0063] The localization generation context building module provides a unified input context for the localization generation substructure of the multimodal large model. In some examples, this module combines a neutral product semantic representation built for the target cross-border product, localized attribute nodes and their constraint expressions in the localization strategy graph, and historical localization samples of the same product category in the target sales country and target sales platform to form a localization generation context. The neutral product semantic representation provides standard attribute names, standard units, and standardized values that are independent of the platform and country, such as "rated capacity: 1.5 liters", "rated voltage: 220 volts", "rated power: 1500 watts", "main material: stainless steel", etc. The localized attribute nodes and constraint expressions in the localization strategy graph provide the required fields, dependencies, and allowed value ranges for the current country and platform. The historical localization samples provide examples of localized field filling for similar electric kettles that have been approved on this platform, along with the corresponding localized description text, serving as a guide for the small sample size of the large model. After construction, the localization generation context is input into the localization generation substructure of the multimodal large model in a structured form.
[0064] Within the localization generation substructure, the attribute slot control generation module performs generation operations sequentially for each target localized commodity attribute item to be filled, according to the order determined by the topological sorting. For a target localized commodity attribute item, such as localized rated capacity, the localization generation substructure first selects the associated neutral commodity attribute from the neutral commodity semantic representation based on the mapping relationship between the attribute item and the neutral commodity attribute pattern, such as the "rated capacity" attribute node and its standardized value "1.5 liters". It also considers the multimodal evidence features associated with this neutral commodity attribute, such as image features from the capacity label area in the packaging image and document features from the corresponding cell in the instruction manual parameter table.
[0065] The attribute slot control mechanism assigns an attribute slot vector to each target localized product attribute item. This attribute slot vector, along with the neutral product attribute vector, multimodal feature representation, and the set of value constraints for that attribute item in the localization strategy graph, is input into the decoding network, ensuring that the decoding process always revolves around the semantic scope of the current attribute item. Based on this, the model generates several candidate localized attribute values for the target localized product attribute item. For example, for the localized rated capacity, it generates the target country's language expression and its different format variations corresponding to "1.5 liters".
[0066] Furthermore, the constrained decoding module is used to perform constraint control during the generation of candidate localized attribute values and the determination of generation results. In some examples, the constrained decoding module first masks the decoding vocabulary based on the set of value constraints for the attribute item in the localization strategy graph. For enumerated values, only allowed words or phrases in the value set are retained as optional outputs; for numeric values, only combinations of numbers and units that conform to a preset format are allowed to be output. Subsequently, after the candidate localized attribute values are generated, the module further performs a numeric validity check, eliminating candidate values whose numerical range exceeds the specified range or contradicts other fields.
[0067] For example, when local regulations stipulate that the rated capacity of a certain type of household electric kettle must not exceed a certain upper limit, the constrained decoding module will perform a range test on the candidate values. If the value exceeds the upper limit, the decoding path will be discarded. After word masking and numerical validity checks, the model determines one or more constrained localized attribute values for the current target localized product attribute item and records the confidence level and candidate path score. The system will perform this process sequentially on each target localized product attribute item in the topology sorting list, gradually filling in the complete localized product attribute data, including fields such as localized rated voltage, localized frequency, localized plug type, localized energy efficiency rating, and localized hazard attribute identification according to the European country's electrical standards.
[0068] Furthermore, after completing the population of localized product attribute data, the localization description generation module uses the localized product attribute data as structured input to drive the multimodal large model to generate localized product description information that is expressed in the language of the target sales country and conforms to the field specifications of the target sales platform. In this embodiment, the localization description generation module can generate localized descriptions of various granularities according to the platform's display templates, including concise titles for search ranking, short copy highlighting selling points for consumers, detailed descriptions with images and text, and functional usage descriptions for regulatory or customs clearance purposes. During the generation process, the localization description generation module embeds the localized product attribute data as a constraint into the decoding prompts, ensuring that key information in the generated description text, such as capacity, voltage, power, material, target audience, and safety warnings, is consistent with the localized product attribute data and does not contain any statements that conflict with the attribute fields.
[0069] Through the localization generation process in this embodiment, the system, under the constraints of the localization strategy graph, accurately maps neutral product semantic representations into localized product attribute data and localized product description information that conform to the language habits of the target sales country and the field specifications of the target sales platform. On the one hand, by combining the attribute slot control mechanism with the constrained decoding mechanism, the localization generation process is constrained by explicit rules based on the strong semantic capabilities of the large model, reducing non-compliant values and semantic shifts. On the other hand, using historical localization samples as the generation context helps the large model learn expression methods that conform to the style of the target platform. Thus, while ensuring the legality and consistency of field values, it improves the readability and adaptability of localized descriptions, reduces the workload of manual adjustments, and enhances the automation level and output quality of cross-border product data localization processing.
[0070] In some embodiments, consistency and rule checks are performed on the localized product attribute data, and localized product attribute data that does not meet preset conditions is corrected based on the check results to obtain localized product data, including: Based on the localized attribute nodes and constraint edges in the localization strategy graph, determine the set of target localized product attribute items that need to be validated. The localized product attribute data and the multimodal feature representation corresponding to the target cross-border product are input into the multimodal consistency discrimination substructure. The semantic consistency between the localized attribute values of each target localized product attribute item and the corresponding text fragments, image regions and document units is judged, and the consistency score and conflict marking information for each target localized product attribute item are output. Based on the dependencies and value constraint expressions between localized attribute nodes in the localization strategy graph, rule reasoning is performed on localized product attribute data to identify the rule violation type of each target localized product attribute item and generate rule verification results. Consistency scores, conflict flag information, and rule verification results are merged to form an error flag vector, which adds error types and adjustment priorities to localized product attribute data that do not meet preset conditions; The neutral product semantic representation, the current localized product attribute data, and the error label vector are used as the correction context. The correction generation substructure of the multimodal large model is input, and under the constraints of the localization strategy graph, the localized product attribute data marked as not meeting the preset conditions is regenerated or numerically adjusted to obtain the corrected localized product attribute data. The corrected localized product attribute data is merged with the localized product attribute data that has passed consistency and rule checks to form localized product data.
[0071] Specifically, after generating localized product attribute data, the cross-border commodity data localization processing system uses a verification and correction subsystem to perform consistency and rule verification on the localized product attribute data. Upon detecting anomalies, it drives a multimodal large model to perform targeted corrections. The verification and correction subsystem includes a target attribute determination module, a multimodal consistency discrimination module, a rule reasoning module, an error marker generation module, a correction generation module, and a result merging module. Each module collaboratively processes data based on the aforementioned localization strategy graph, neutral product semantic representation, and multimodal feature representation. This embodiment still uses the example of an electric kettle being listed on a platform in a European country for illustration.
[0072] The target attribute determination module first determines the set of target localized product attributes to be verified based on the localized attribute nodes and constraint edges in the localization strategy graph. For electric kettles, due to the large number of safety and compliance attributes involved, the target attribute set must include at least the following fields: localized rated voltage, localized rated power, localized capacity, localized hazard identification, localized energy efficiency rating, localized applicable age range, and localized child safety warnings. The target attribute determination module can assign higher verification priority to safety-related and regulatory-related fields based on the weight of constraint edges and the importance of nodes in the strategy graph, thus focusing on high-risk attributes while ensuring overall efficiency.
[0073] The multimodal consistency discrimination module is used to determine the consistency between localized product attribute data and multimodal evidence. In some examples, this module takes the current localized product attribute data and the multimodal feature representation corresponding to the target cross-border product as input, calls the multimodal consistency discrimination substructure, and determines the semantic consistency between the localized attribute value of each target localized product attribute item and the corresponding text fragment, image region, and document unit.
[0074] Taking the localized rated capacity field as an example, the system will retrieve the multimodal evidence indication information corresponding to this attribute in the neutral product semantic representation, locate the document cell of "rated capacity 1.5L" in the instruction manual parameter table and the "1.5L" marking area in the packaging label image, and input the feature vectors corresponding to these pieces of evidence and the semantic representation of the localized attribute value "1.5 liters" into the multimodal consistency discrimination substructure to calculate the consistency score.
[0075] When there is a significant contradiction between the localized attribute value and the evidence content, such as a localized attribute being "2.0 liters" while all evidence points to "1.5L", the multimodal consistency discrimination substructure will output a lower consistency score and mark the attribute with conflict information. Similarly, for fields such as localized rated voltage, localized applicable scenarios, and localized hazard attribute labels, the consistency between the localized values and the nameplate image, packaging warning label, and instruction manual paragraphs is also compared in this way.
[0076] Furthermore, the rule reasoning module is used to perform rule validation on localized product attribute data based on the dependencies and value constraint expressions between localized attribute nodes in the localization strategy graph. In some examples, the rule reasoning module treats the localized product attribute data as the result of assigning values to each localized attribute node in the localization strategy graph, and evaluates each rule by traversing the constraint edges in the strategy graph and their associated value constraint expressions.
[0077] For example, regarding the rule that "the rated voltage field must be within the local power grid standard range," the module checks whether the localized rated voltage is within the allowable range. Regarding the rule that "the energy efficiency rating field is mandatory when the rated power exceeds a specific threshold, and its value must be one of a predefined set," the module checks both the localized rated power and the localized energy efficiency rating fields to determine if they are missing or have invalid values. Regarding the rule that "when the hazard attribute is identified as flammable, the applicable scenario cannot be a children's room," the module checks whether the combination of the localized hazard attribute identifier and the localized applicable scenario field violates mutual exclusion constraints. Through the evaluation of rule expressions, the rule reasoning module identifies the rule violation type for each target localized product attribute item, such as missing, contradictory, out-of-bounds, or mutually exclusive conflicts, and generates rule verification results.
[0078] Furthermore, the error tag generation module is responsible for merging the multimodal consistency judgment results with the rule verification results to form a unified error tag vector. In some examples, each dimension of the error tag vector corresponds to a target localized product attribute item, used to record the consistency score, conflict tag information, and rule violation type of that attribute item, while also attaching the error type and adjustment priority.
[0079] For example, for the localized rated voltage field, if the multimodal conformance determination result shows that it is consistent with the nameplate image but does not conform to the local voltage standard rules, the error labeling vector will mark the field as a rule out-of-bounds error and adjust the priority to a higher level to ensure that compliance issues are corrected first; for the localized capacity field, if it is inconsistent with the multimodal evidence but does not violate the numerical range rules, the error labeling vector will mark it as a multimodal conflict error and adjust the priority to a relatively lower level.
[0080] Furthermore, the correction generation module utilizes the correction generation substructure of the multimodal large model to perform targeted corrections on localized product attribute data marked as not meeting preset conditions. In some examples, the correction generation module takes the neutral product semantic representation, the current localized product attribute data, and the error marker vector as correction context inputs to the correction generation substructure and decodes them under the constraints of the localization strategy graph. Unlike the initial localization generation, the correction generation substructure applies a freeze constraint to the localized attribute items that have passed the verification during decoding, allowing only the fields marked as needing correction in the error marker vector to be regenerated or have their values adjusted.
[0081] For example, if the current localized rated voltage is "110 volts," while the target national grid standard is "220-240 volts," and the rule reasoning module determines this to be an out-of-bounds error, the correction generation substructure, upon receiving this error information, will combine the original value of the rated voltage in the neutral product semantic representation (e.g., "220V"), relevant multimodal evidence (e.g., "220V~50Hz" in the nameplate image), and the value constraints in the localization strategy graph to regenerate the localized rated voltage field, obtaining "220 volts" that conforms to the grid standard requirements and is consistent with the evidence. For fields such as the localized applicable age range and localized child safety warnings, when the rule reasoning module finds that the former's value points to "suitable for children under 3 years old" while the latter is missing, the correction generation substructure will complete the corresponding safety warning text under constraints without modifying other error-free fields.
[0082] After the correction and generation process is complete, the merging module merges the corrected localized product attribute data with the original localized product attribute data that passed consistency and rule checks, forming the final localized product data. During the merging process, the module records the final value and a correction flag for each localized attribute item for subsequent auditing and backtracking. Once all high-priority errors have been corrected and no new violations are found after re-verification, the system submits the localized product data to the upper-level business system for generating product listing information or customs declaration documents.
[0083] Through the verification and correction process in this embodiment, the system can perform fine-grained consistency and rule verification on localized commodity attribute data under the dual constraints of the localization strategy graph and multimodal evidence. It also utilizes the correction generation capabilities of the multimodal large model to perform targeted corrections on problematic fields. On the one hand, this process significantly reduces the probability of attribute values in the localization results being inconsistent with multimodal evidence or violating regulatory rules, thus improving the reliability and compliance of localized commodity data. On the other hand, through error labeling vectors and field-level correction mechanisms, targeted adjustments to certain fields are achieved, reducing the workload of full regeneration and manual review, which is conducive to improving the automation level and overall processing efficiency of cross-border commodity data localization processing.
[0084] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0085] Figure 2 This is a schematic diagram of the structure of the cross-border commodity data localization processing device based on a multimodal large model provided in this application embodiment. Figure 2 As shown, the cross-border commodity data localization processing device based on a multimodal large model includes: The data acquisition module 201 is used to collect and preprocess the product text information, product image information and product description document information associated with the target cross-border product in a local computing environment corresponding to the country to which the cross-border product belongs, to obtain a multimodal data set; Extraction module 202 is used to input the multimodal dataset into the multimodal large model to obtain multimodal feature representations for the target cross-border commodities and extract candidate commodity attributes; The standardization module 203 is used to standardize the attribute names, unify the units, and normalize the values of candidate product attributes based on the preset neutral product attribute pattern, and generate a neutral product semantic representation that is independent of the platform and country. Module 204 is used to construct a localization strategy graph that limits the target localized product attribute items and the dependencies and value constraints between the target localized product attribute items, based on the rule information of the target sales country and the target sales platform. The generation module 205 is used to input the neutral product semantic representation into the multimodal large model for localization generation under the constraints of the localization strategy graph. According to the attribute filling order and value constraints determined by the localization strategy graph, it obtains localized product attribute data and localized product description information that conform to the language expression of the target sales country and the field specifications of the target sales platform. The verification module 206 is used to perform consistency verification and rule verification on localized product attribute data based on multimodal feature representation and localization strategy graph. Based on the verification results, localized product attribute data that does not meet the preset conditions is corrected to obtain localized product data.
[0086] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0087] Figure 3 This is a schematic diagram of the electronic device 3 provided in an embodiment of this application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various method embodiments described above. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the various device embodiments described above.
[0088] Electronic device 3 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 3 may include, but is not limited to, processor 301 and memory 302. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or different components.
[0089] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0090] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. The memory 302 can also include both internal and external storage units of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device.
[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0092] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a readable storage medium (e.g., a computer-readable storage medium). Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which may be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0093] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for localizing cross-border commodity data based on a multimodal large model, characterized in that, include: In the local computing environment corresponding to the country of origin of the cross-border goods, the product text information, product image information and product description document information associated with the target cross-border goods are collected and preprocessed to obtain a multimodal data set; The multimodal dataset is input into the text encoding substructure, image encoding substructure, document structure encoding substructure and cross-modal alignment structure of the multimodal large model to obtain multimodal feature representations for the target cross-border goods, and candidate goods attributes are extracted through the attribute discrimination substructure in the multimodal large model; Based on a preset neutral product attribute pattern, the candidate product attributes are normalized in terms of attribute names, units, and values, generating a neutral product semantic representation that is independent of the platform and country. Based on the rules of the target sales country and the target sales platform, construct a localization strategy graph that limits the target localized product attribute items and the dependencies and value constraints between the target localized product attribute items; Under the constraints of the localization strategy graph, the neutral product semantic representation is input into the localization generation substructure in the multimodal large model for localization generation. According to the attribute filling order and value constraints determined by the localization strategy graph, localized product attribute data and localized product description information that conform to the language expression of the target sales country and the field specifications of the target sales platform are obtained. Based on the multimodal feature representation and the localization strategy graph, consistency and rule checks are performed on the localized product attribute data. According to the check results, the localized product attribute data that does not meet the preset conditions is corrected to obtain localized product data. The step of constructing a localization strategy graph that defines the target localized product attribute items and the dependencies and value constraints between these items, based on the rule information of the target sales country and the target sales platform, includes: Obtain rule information associated with the target sales country and target sales platform from the cross-border product localization knowledge base. The rule information includes target platform field definitions, mandatory field conditions, field dependencies, and field value constraints. The fields defined in the rule information are abstracted into localized attribute nodes, and the mandatory dependencies, conditional dependencies, and conflict constraints between fields are abstracted into constraint edges connecting the localized attribute nodes, thus constructing an initial rule relationship graph. The rule information is structured and encoded based on a preset rule template, converting the text-based regulatory clauses and platform configurations into computable dependency expressions and value constraint expressions, and associating them with the corresponding localized attribute nodes and constraint edges. By using a graph structure learning model combined with historical localized samples, the effective conditions and priorities of each localized attribute node and constraint edge in the initial rule relationship graph are updated to determine the set of activated localized attribute nodes and their constraint relationships for the target cross-border goods, thus obtaining a localized strategy subgraph. The localization strategy subgraph is topologically sorted to determine the filling order of the target localized product attribute items. The target localized product attribute items, their corresponding dependencies, and value constraints are then solidified in a graph structure to form a localization strategy graph used to constrain the localization generation process.
2. The method according to claim 1, characterized in that, The process of collecting and preprocessing product text information, product image information, and product description document information associated with the target cross-border goods yields a multimodal data set, including: In the local computing environment, retrieve the original product text, original product images, and original product description electronic documents associated with the target cross-border product from business systems, production systems, and file storage systems; Perform layout structure analysis and optical character recognition on the original electronic product description document to extract paragraphs, titles, table cells and text, and generate document text fragments carrying layout position information; Perform product and label region detection and text recognition on the original product image to extract image text fragments carrying spatial location information and product appearance features; Sensitive fields are identified and desensitized based on a pre-trained semantic model for the original product text, document text fragments, and image text fragments. Based on the internal identifiers of the target cross-border goods, the anonymized product text information, product image information, and product description document information are uniformly numbered and associated with an index to construct a multimodal data set that corresponds one-to-one with the target cross-border goods.
3. The method according to claim 1, characterized in that, The process of inputting the multimodal dataset into a multimodal large model to obtain a multimodal feature representation for the target cross-border commodity and extracting candidate commodity attributes includes: The product text information, product image information, and product description document information in the multimodal dataset are respectively input into the text encoding substructure, image encoding substructure, and document structure encoding substructure in the multimodal large model to obtain the corresponding text feature representation, image feature representation, and document feature representation; Based on the cross-modal alignment structure, the text feature representation, image feature representation and document feature representation are aligned and fused in a unified semantic space to generate a product-level multimodal feature representation corresponding to the target cross-border product; Under the constraint of a preset set of attribute slots, the cross-modal attention aggregation mechanism guided by attribute slots is used to aggregate the corresponding attribute feature vectors for each attribute slot from the product-level multimodal feature representation. Based on the attribute discrimination substructure, the attribute category is determined and the attribute value is inferred for each attribute feature vector, and the output is a set of candidate product attributes containing candidate attribute names and candidate attribute values.
4. The method according to claim 1, characterized in that, The method, based on a preset neutral product attribute pattern, standardizes the attribute names, unifies units, and normalizes values of the candidate product attributes to generate a neutral product semantic representation independent of the platform and country, including: Based on the neutral product attribute pattern, a neutral product attribute graph is constructed. Neutral product attribute nodes are established for each neutral product attribute in the neutral product attribute graph, and attribute synonym name nodes, attribute unit nodes, and attribute value range nodes corresponding to each neutral product attribute are associated. The attribute names and attribute values of the candidate product attributes are mapped to the neutral product attribute graph. Based on graph embedding encoding and similarity measurement, the target neutral product attribute node and standard attribute name corresponding to each candidate product attribute are determined, and the attribute names of the candidate product attributes are normalized. Based on the unit conversion rules associated with each neutral product attribute node, dimensional identification and unit conversion are performed on the numerical attribute values in the candidate product attributes to convert the physical quantities of the candidate product attributes into the standard units specified in the neutral product attribute pattern. Based on the set of value specifications associated with each neutral product attribute node, the discrete and interval values of the candidate product attributes are merged and discretized to obtain standardized attribute values that conform to the neutral product attribute pattern value constraints. According to the field definitions in the neutral commodity attribute pattern, the standardized standard attribute name, standard unit, and standardized attribute value are associated with the evidence indication information of the candidate commodity attribute to generate a neutral commodity semantic representation for the target cross-border commodity.
5. The method according to claim 1, characterized in that, The process involves inputting the neutral product semantic representation into a multimodal large model for localization generation. Following the attribute filling order and value constraints determined by the localization strategy graph, localized product attribute data and localized product description information that conform to the target sales country's language and the target sales platform's field specifications are obtained, including: Based on the localization strategy graph, the target localized product attribute items are topologically sorted to determine the filling order of each target localized product attribute item and the set of value constraints corresponding to each target localized product attribute item. The neutral product semantic representation, the localized attribute nodes and constraint expressions in the localization strategy graph, and the historical localization samples for the target sales country and target sales platform are combined to form a localization generation context, which is then input into the localization generation substructure in the multimodal large model. In the localization generation substructure, for the target localized product attribute item to be filled, the attribute slot control mechanism is used to select the neutral product attribute associated with the target localized product attribute item from the neutral product semantic representation, and combine the corresponding multimodal feature representation and value constraint set to generate the candidate localized attribute value of the target localized product attribute item. Under the constrained decoding mechanism, the candidate localized attribute values are subjected to word masking and numerical validity checks. Decoding paths that do not meet the constraints are eliminated according to the value constraint set. The localized attribute values of each target localized product attribute item are determined to form the localized product attribute data. Using the localized product attribute data as structured input, the multimodal large model is driven to generate localized product description information that uses the language of the target sales country and conforms to the field specifications of the target sales platform.
6. The method according to claim 1, characterized in that, The process of performing consistency and rule checks on the localized product attribute data, and correcting localized product attribute data that does not meet preset conditions based on the check results, to obtain localized product data, includes: Based on the localized attribute nodes and constraint edges in the localization strategy graph, determine the set of target localized product attribute items that need to be validated. The localized product attribute data and the multimodal feature representation corresponding to the target cross-border product are input into the multimodal consistency discrimination substructure. The semantic consistency between the localized attribute values of each target localized product attribute item and the corresponding text fragments, image regions and document units is judged, and the consistency score and conflict marking information for each target localized product attribute item are output. Based on the dependencies and value constraint expressions between localized attribute nodes in the localization strategy graph, rule reasoning is performed on the localized product attribute data to identify the rule violation type of each target localized product attribute item and generate rule verification results. The consistency score, conflict marker information, and rule verification results are merged to form an error marker vector, and error types and adjustment priorities are added to localized product attribute data that do not meet the preset conditions. The neutral product semantic representation, the current localized product attribute data, and the error label vector are used as the correction context and input into the correction generation substructure in the multimodal large model. The correction generation substructure applies a freeze constraint to the localized product attribute data that has passed the verification, and performs regeneration or numerical adjustment for the localized product attribute data marked as not meeting the preset conditions under the constraint of the localization strategy graph, so as to obtain the corrected localized product attribute data. The corrected localized product attribute data is merged with the localized product attribute data that has passed consistency and rule checks to form the localized product data.
7. A cross-border commodity data localization processing device based on a multimodal large model, characterized in that, include: The data acquisition module is used to collect and preprocess product text information, product image information, and product description document information associated with the target cross-border product in a local computing environment corresponding to the country to which the cross-border product belongs, to obtain a multimodal data set; The extraction module is used to input the multimodal data set into the text encoding substructure, image encoding substructure, document structure encoding substructure and cross-modal alignment structure in the multimodal large model to obtain the multimodal feature representation for the target cross-border commodity, and extract candidate commodity attributes through the attribute discrimination substructure in the multimodal large model; The standardization module is used to standardize the attribute names, unify the units, and normalize the values of the candidate product attributes based on a preset neutral product attribute pattern, and generate a neutral product semantic representation that is independent of the platform and country. The building module is used to construct a localization strategy graph that defines the target localized product attribute items and the dependencies and value constraints between the target localized product attribute items, based on the rule information of the target sales country and the target sales platform. The generation module is used to input the neutral product semantic representation into the localization generation substructure in the multimodal large model for localization generation under the constraints of the localization strategy graph. According to the attribute filling order and value constraints determined by the localization strategy graph, it obtains localized product attribute data and localized product description information that conform to the language expression of the target sales country and the field specifications of the target sales platform. The verification module is used to perform consistency verification and rule verification on the localized product attribute data based on the multimodal feature representation and the localization strategy graph, and to correct the localized product attribute data that does not meet the preset conditions according to the verification results, so as to obtain localized product data. The construction module is used to obtain rule information associated with the target sales country and target sales platform from a cross-border commodity localization knowledge base. This rule information includes target platform field definitions, mandatory field conditions, field dependencies, and field value constraints. The module abstracts the field definitions in the rule information into localized attribute nodes, and abstracts the mandatory dependencies, conditional dependencies, and conflict constraints between fields into constraint edges connecting these localized attribute nodes, thus constructing an initial rule relationship graph. Based on a preset rule template, the module performs structured encoding on the rule information, converting textual regulatory clauses and platform configurations into computable dependency expressions and values. Constraint expressions are generated and associated with corresponding localized attribute nodes and constraint edges. Using a graph structure learning model combined with historical localization samples, the effective conditions and priorities of each localized attribute node and constraint edge in the initial rule relationship graph are updated to determine the set of activated localized attribute nodes and their constraint relationships for the target cross-border product, resulting in a localization strategy subgraph. The localization strategy subgraph is topologically sorted to determine the filling order of the target localized product attribute items, and the target localized product attribute items, their corresponding dependencies, and value constraints are solidified in a graph structure to form a localization strategy graph used to constrain the localization generation process.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.