Cross-border consumption behavior dynamic analysis method and device based on large language model
By analyzing multimodal data in cross-border consumption scenarios using a large language model, a unified semantic representation is generated and modal weights are dynamically allocated. This solves the problem of insufficient modal correlation in cross-border consumption behavior analysis, achieves more accurate consumer intent parsing, and supports refined operations in cross-border e-commerce.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-14
AI Technical Summary
Existing cross-border consumer behavior analysis technologies struggle to fully uncover the inherent semantic relationships between different modalities of data, such as text descriptions, product images, and promotional videos. Furthermore, they cannot adjust the focus of analysis according to the characteristics of different scenarios, resulting in insufficient accuracy in interpreting consumer behavior intentions and failing to support the refined operational needs of cross-border e-commerce.
A dynamic analysis method for cross-border consumer behavior based on a large language model is adopted. By acquiring multimodal raw data, feature extraction and fusion processing are performed to generate a unified semantic representation. The contribution weights of text modality and visual modality are dynamically allocated according to the scenario type, and finally, consumer behavior intent tags are generated.
It improves the accuracy and scenario adaptability of cross-border consumer behavior analysis, and can provide reliable and refined operational support for cross-border e-commerce.
Smart Images

Figure CN121479715B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, specifically to a method and apparatus for dynamic analysis of cross-border consumption behavior based on a large language model. Background Technology
[0002] In recent years, with the rapid development of global e-commerce, the cross-border consumption market has continued to expand. To accurately understand the preferences and intentions of consumers in different countries and regions, and thus conduct personalized product recommendations, advertising, and marketing strategy development, in-depth analysis of massive and diverse consumer-related data has become a key requirement for cross-border e-commerce operations.
[0003] Existing consumer behavior analysis technologies largely rely on single-modal analysis models such as natural language processing or computer vision, or simply concatenate multimodal data, for example, analyzing text reviews solely through sentiment analysis models or extracting product visual features solely through image recognition models. However, in complex cross-border consumption scenarios, these methods have significant limitations. On the one hand, they struggle to uncover the inherent semantic relationships between different modalities of data, such as text descriptions, product images, and promotional videos, resulting in an insufficient and in-depth understanding of the consumer's overall intent. On the other hand, the importance of textual and visual information in reflecting user intent varies across different consumption scenarios, but existing technologies cannot adjust their analytical focus according to the characteristics of each scenario, limiting the adaptability and accuracy of the analysis results. Therefore, existing technologies suffer from insufficient accuracy in analyzing consumer behavior intent in cross-border consumption scenarios, making it difficult to support the refined operational needs of cross-border e-commerce.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] This application provides a method and apparatus for dynamic analysis of cross-border consumer behavior based on a large language model, which can improve the adaptability and accuracy of consumer behavior intent analysis in cross-border consumption scenarios and meet the needs of refined operation of cross-border e-commerce.
[0006] In a first aspect, embodiments of this application provide a method for dynamic analysis of cross-border consumption behavior based on a large language model, including:
[0007] Acquire multimodal raw data from multiple data sources for a target cross-border consumption scenario, wherein the multimodal raw data includes at least textual and visual data related to the target product or service;
[0008] Feature extraction and fusion processing are performed on the original multimodal data to generate a unified semantic representation that includes textual and visual semantics;
[0009] Identify the scenario type to which the target cross-border consumption scenario belongs, and determine the contribution weights of the text modality and visual modality in the unified semantic representation based on the scenario type;
[0010] The weighted unified semantic representation is parsed based on the contribution weight to generate consumer behavior intent tags corresponding to the target cross-border consumption scenario.
[0011] Furthermore, in some embodiments of this application, the step of performing feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation containing textual and visual semantics includes:
[0012] The text data is parsed using a pre-trained text feature extraction model to obtain text feature vectors;
[0013] The visual data is used to extract features using a visual feature extraction model to obtain a visual feature vector;
[0014] The text feature vector and the visual feature vector are input into a cross-modal attention network;
[0015] The cross-modal attention network performs attention interaction on the text feature vector and the visual feature vector to generate a unified semantic representation that integrates cross-modal information.
[0016] Furthermore, in some embodiments of this application, the step of generating a unified semantic representation that fuses cross-modal information by performing attention interaction on the text feature vector and the visual feature vector through the cross-modal attention network includes:
[0017] In the cross-modal attention network, a multi-head attention mechanism is employed to calculate the attention weight distribution between each unit in the text feature vector and each region in the visual feature vector;
[0018] The visual feature vectors are weighted and converged according to the attention weight distribution to generate a visual context representation that is semantically aligned with the text feature vectors;
[0019] The text feature vector and the visual context representation are concatenated and projected to form the unified semantic representation.
[0020] Furthermore, in some embodiments of this application, the step of parsing the text data using a pre-trained text feature extraction model to obtain text feature vectors includes:
[0021] Identify the language category of the text data;
[0022] Based on the language category, the corresponding pre-trained text feature extraction model is invoked to perform preliminary semantic parsing on the text data to obtain basic semantic vectors;
[0023] Access a pre-built cultural symbol knowledge base to retrieve culturally specific semantic information associated with keywords in the text data;
[0024] The culturally specific semantic information and the basic semantic vector are fused to generate the text feature vector enhanced with cultural semantics.
[0025] Furthermore, in some embodiments of this application, the construction and updating of the cultural symbol knowledge base includes:
[0026] Automatically identify and extract expressions containing cultural metaphors or specific customs from unstructured text corpora from multiple languages and regions;
[0027] The expression fragments are cleaned and labeled to determine the cultural region and core cultural semantics to which they belong.
[0028] A contrastive learning algorithm is used to map the expression fragments and their core cultural semantics to a unified vector space, forming a structured mapping relationship and storing it in the cultural symbol knowledge base;
[0029] Based on new cross-border consumption data and user feedback, the mapping relationships in the cultural symbol knowledge base are dynamically calibrated and expanded.
[0030] Furthermore, in some embodiments of this application, the step of identifying the scenario type to which the target cross-border consumption scenario belongs, and determining the contribution weights corresponding to the textual and visual modalities in the unified semantic representation based on the scenario type, includes:
[0031] Scene feature analysis is performed on the unified semantic representation to identify the scene type to which the target cross-border consumption scene belongs;
[0032] The scene type is input into a pre-trained dynamic weight allocation model;
[0033] The dynamic weight allocation model calculates and outputs the text modality contribution coefficient and visual modality contribution coefficient, which are adapted to the scene type, in real time, as the contribution weight.
[0034] Furthermore, in some embodiments of this application, the training method of the dynamic weight allocation model includes:
[0035] Construct a training sample set, where each training sample contains a scenario type label, multimodal data, and corresponding real consumer behavior feedback;
[0036] Construct a dynamic weight allocation model to be trained, with scene type as input and modal contribution coefficient as output;
[0037] With the optimization goal of maximizing the accuracy of consumer behavior intention prediction or business conversion rate, the dynamic weight allocation model is iteratively trained using a reinforcement learning algorithm.
[0038] Furthermore, in some embodiments of this application, the step of parsing the weighted unified semantic representation according to the contribution weight to generate a consumer behavior intent tag corresponding to the target cross-border consumption scenario includes:
[0039] Based on the contribution weights, the text modal components and visual modal components in the unified semantic representation are weighted respectively to obtain a weighted semantic representation;
[0040] The weighted semantic representation is input into the trained multimodal large language model;
[0041] The multimodal large language model is used to infer the weighted semantic representation and output one or more consumer behavior intent labels and their corresponding confidence scores.
[0042] Furthermore, in some embodiments of this application, the method further includes:
[0043] The consumer behavior intent tags are classified and their intensity is quantified according to multiple predefined preference dimensions;
[0044] Based on the intensity quantification results of each preference dimension, an intent analysis map is generated in the form of a heatmap.
[0045] The intent analysis map and its corresponding scene types and cultural region information are integrated to generate a consumer intent visualization report.
[0046] Secondly, embodiments of this application provide a dynamic analysis device for cross-border consumption behavior based on a large language model, comprising:
[0047] The data acquisition module is used to acquire multimodal raw data of the target cross-border consumption scenario from multiple data sources. The multimodal raw data includes at least text data and visual data related to the target product or target service.
[0048] The data processing module is used to perform feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation that includes textual and visual semantics;
[0049] The dynamic weighting module is used to identify the scenario type to which the target cross-border consumption scenario belongs, and to determine the contribution weights of the text modality and visual modality in the unified semantic representation based on the scenario type.
[0050] The intent recognition module is used to parse the weighted unified semantic representation according to the contribution weight, and generate consumer behavior intent tags corresponding to the target cross-border consumption scenario.
[0051] This application provides a method and apparatus for dynamic analysis of cross-border consumption behavior based on a large language model.
[0052] First, multimodal raw data, including text and visual data, is acquired from the target cross-border consumption scenario. Feature extraction and fusion processing generate a unified semantic representation that considers both textual and visual semantics, effectively breaking down the isolation of different modal data and comprehensively integrating consumer behavior-related information, avoiding semantic gaps caused by single-modal or simple splicing processing. Then, by identifying scenario types and dynamically determining the contribution weights of textual and visual modalities, the analysis process accurately matches the information needs of different scenarios. Finally, based on the comprehensive and scenario-adaptive weighted unified semantic representation, parsing is performed to generate consumer behavior intent tags highly consistent with the target cross-border consumption scenario. Therefore, this application can improve the accuracy and scenario adaptability of cross-border consumer behavior intent parsing, effectively solving the problem of low accuracy in intent parsing caused by insufficient utilization of multimodal information or inadequate scenario adaptability in existing technologies, thus providing reliable support for the refined operation of cross-border e-commerce. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is an application environment diagram of the dynamic analysis method for cross-border consumption behavior based on a large language model provided in the embodiments of this application;
[0055] Figure 2 This is a flowchart illustrating the dynamic analysis method for cross-border consumption behavior based on a large language model provided in this application embodiment;
[0056] Figure 3 This is a schematic diagram of the structure of the cross-border consumption behavior dynamic analysis device based on a large language model provided in the embodiments of this application;
[0057] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0058] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of systems and methods consistent with those detailed in the appended claims or with some aspects of this application.
[0059] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover descriptions such as non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0060] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0061] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0062] To address the aforementioned technical problems and overcome the shortcomings of existing technologies, this application provides a method and apparatus for dynamic analysis of cross-border consumer behavior based on a large language model. This method and apparatus can improve the adaptability and accuracy of cross-border consumer behavior intent analysis and meet the needs of refined operations in cross-border e-commerce.
[0063] Figure 1 This is a diagram illustrating the application environment of a dynamic analysis method for cross-border consumer behavior based on a large language model, as shown in one embodiment. (Refer to...) Figure 1This dynamic analysis method for cross-border consumption behavior based on a large language model is applied to a dynamic analysis system for cross-border consumption behavior based on a large language model. The system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, tablet, or laptop. The server 120 can be a standalone server or a server cluster consisting of multiple servers. The server 120 is configured to execute the aforementioned dynamic analysis method for cross-border consumption behavior based on a large language model, including: acquiring multimodal raw data of a target cross-border consumption scenario from multiple data sources, wherein the multimodal raw data includes at least textual and visual data related to the target product or service; performing feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation containing textual and visual semantics; identifying the scenario type of the target cross-border consumption scenario and determining the contribution weights of the textual and visual modalities in the unified semantic representation based on the scenario type; and parsing the weighted unified semantic representation according to the contribution weights to generate consumer behavior intent tags corresponding to the target cross-border consumption scenario.
[0064] Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for dynamic analysis of cross-border consumption behavior based on a large language model, according to an embodiment of this application. This embodiment primarily uses the application of this method to computer equipment as an example. Specifically, the method for dynamic analysis of cross-border consumption behavior based on a large language model provided in this application may include the following steps:
[0065] S1. Obtain multimodal raw data from multiple data sources for the target cross-border consumption scenario. The multimodal raw data shall include at least textual and visual data related to the target product or service.
[0066] Specifically, step S1 primarily involves collecting comprehensive data related to cross-border consumption behavior to provide a foundation for subsequent analysis. Data sources can include product detail pages on global e-commerce platforms, user review sections, live stream replays, social media consumer sharing content, and interaction data from cross-border shopping apps. Data types must include at least two core categories: textual and visual data. Textual data can include product descriptions, user reviews, consultation dialogues, and promotional activity descriptions, such as Indonesian user reviews like "Barang berkualitas dan hargaterjangkau" (good quality and affordable price) or the "IPX7 waterproof rating" description on product detail pages. Visual data can include real product photos, videos of users using the product, product displays during live streams, and product packaging images, such as detailed fabric images of clothing, unboxing videos of electronic products, and live stream demonstrations of product usage. It's important to note that the collected data should focus on specific target cross-border consumption scenarios to ensure relevance to the analysis object; for example, collecting data for specific scenarios such as "cross-border retail of women's clothing in the Southeast Asian market" or "live-streaming sales of electronic products in the Middle East market."
[0067] S2. Perform feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation that includes textual and visual semantics;
[0068] Specifically, in step S2, semantic analysis is performed on the collected text data to extract features that reflect core information such as product attributes, user needs, and evaluation attitudes. For example, the semantic feature "strong battery life" is extracted from "This headphone has a battery life of up to 20 hours," and the emotional feature "dissatisfaction with the logistics speed" is extracted from "The logistics are too slow, and the experience is poor." Key information is extracted from the visual data to reveal features that reflect core information such as product appearance, usage status, and scenario adaptability. For example, the appearance features "full-screen design" and "metal body" are extracted from real-life photos taken with a mobile phone, and the usage feature "shoulder strap fit under heavy load" is extracted from outdoor backpack usage videos. The extracted text features and visual features are then deeply integrated to eliminate semantic barriers between different modalities, forming a unified semantic representation that includes both textual semantics (such as product functions and user evaluation tendencies) and visual semantics (such as the actual presentation effect of the product and usage scenarios). For example, the text feature "waterproof" and the visual feature "no penetration in rain test" are merged to form the unified semantic representation "The product has reliable waterproof function."
[0069] S3. Identify the scenario type of the target cross-border consumption scenario, and determine the contribution weights of the text modality and visual modality in the unified semantic representation based on the scenario type;
[0070] Specifically, for step S3, based on the collected multimodal data features and the corresponding consumption scenario background, the specific type of the target cross-border consumption scenario is determined. For example, if the data comes from a product details page and mainly consists of parameter descriptions and function descriptions, it is identified as a "product details page browsing scenario"; if the data comes from a live broadcast and includes host demonstrations and real-time interaction, it is identified as a "live-streaming e-commerce scenario"; and if the data mainly consists of user past purchase reviews, it is identified as a "user review analysis scenario". Then, based on the identified scenario type, the contribution weights of text and visual modalities in the unified semantic representation are flexibly adjusted to match the information needs of users in the scenario. For example, in the "product details page browsing scenario", users pay more attention to textual information such as product parameters and functional details, so the contribution weight of the text modal is increased; in the "live-streaming e-commerce scenario", users pay more attention to visual information such as the real-time display and usage effects of the product, so the contribution weight of the visual modal is increased; in the "user review analysis scenario", the weights of the two are dynamically balanced based on the importance of the textual description of the review content and the accompanying images / videos.
[0071] S4. Parse the weighted unified semantic representation according to the contribution weight to generate consumer behavior intent tags corresponding to the target cross-border consumption scenario;
[0072] Specifically, in step S4, the text modal components and visual modal components in the unified semantic representation are weighted and strengthened according to the determined contribution weights. For example, when the text modal weight is 70%, the text-related semantics in the unified semantic representation are emphasized; when the visual modal weight is 45%, the visual-related semantics in the unified semantic representation are strengthened. The weighted unified semantic representation is then analyzed in depth to uncover the underlying consumer core needs, purchasing tendencies, and key concerns, and these are extracted into clear tag forms. For example, in the "product details page browsing scenario," the weighted semantic representation highlights the core textual information of "reasonable price" and "large capacity," generating intent tags such as "focus on cost-effectiveness" and "emphasis on practical functions" after parsing. In the "live-streaming e-commerce scenario," the weighted semantic representation highlights the core visual information of "stylish appearance" and "easy to use," generating intent tags such as "focus on appearance design" and "preference for ease of use" after parsing. In the "user review analysis scenario," combining the weighted textual review semantics with visual supporting information generates intent tags such as "satisfied with quality" and "complaining about damaged packaging."
[0073] This embodiment systematically integrates raw textual and visual multimodal data from cross-border consumption scenarios, extracts and fuses features to form a unified representation that takes into account both semantics, and dynamically allocates modal contribution weights according to scenario type to accurately generate consumer behavior intent tags. This effectively solves the pain points of multimodal fragmentation and rigid weights in traditional technologies, significantly improves the accuracy and scenario adaptability of cross-border consumption intent analysis, and provides core support for the refined operation of cross-border e-commerce.
[0074] Furthermore, in some embodiments, step S2, "performing feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation containing textual and visual semantics," may specifically include:
[0075] S21. The text data is parsed using a pre-trained text feature extraction model to obtain text feature vectors;
[0076] Specifically, for step S21, the text data encompasses various textual information related to the target product or service in cross-border consumption scenarios, such as functional parameter descriptions on product detail pages, user-submitted shopping reviews, customer service conversations, and platform-released promotional activity descriptions. The text data may include different language types (e.g., English, Indonesian, Arabic, etc.). The pre-trained text feature extraction model, trained on a large amount of multilingual text data, possesses the ability to deeply analyze text semantics, identifying core information in the text, such as product attributes, user needs, and sentiment tendencies, and converting it into a computer-processable vector form. The collected text data is input into the pre-trained text feature extraction model. The model processes the text through word segmentation, semantic encoding, and other methods to uncover the deeper semantics behind the text, ultimately outputting a fixed-dimensional text feature vector.
[0077] For example, if the text data is an Indonesian user review “Produk tahan lama dan mudahdigunakan” (the product is durable and easy to use), the pre-trained text feature extraction model will parse out the core semantics of “durable” and “easy to use” and generate a text feature vector that can represent the semantics; if the text data is “wireless charging power 20W” from the product details page, the model will extract the key information of “wireless charging” and “20W power” and convert it into the corresponding text feature vector.
[0078] S22. Visual feature vectors are obtained by extracting features from visual data using a visual feature extraction model;
[0079] Specifically, for step S22, visual data includes various visual information related to the target product or service in cross-border consumption scenarios, such as real-life photos of the product (front view, detail view, usage scenario view), user-shared product usage videos, product display images during live streams, and images of product packaging and accessories. The visual feature extraction model has the ability to identify key elements in visual information, extracting core features such as the product's appearance attributes, usage status, and scenario adaptability from images or videos, and converting them into standardized vector forms. Visual data is input into the visual feature extraction model, which analyzes and processes image pixels and video frames to identify key features in the visual data, such as color, shape, structure, and motion, and performs quantization encoding, ultimately outputting a visual feature vector that matches the dimensions of the text feature vector.
[0080] For example, if the visual data is a real-life photo of an outdoor backpack, the visual feature extraction model will extract appearance and structural features such as "double shoulder design", "waterproof fabric texture", "multi-pocket structure" and "capacity size" to generate corresponding visual feature vectors; if the visual data is a video of a coffee machine in use, the model will extract features such as "operation steps", "coffee dispensing speed" and "machine material texture" from keyframes and convert them into visual feature vectors.
[0081] S23. Input the text feature vector and visual feature vector into a cross-modal attention network;
[0082] Specifically, step S23 is a crucial preliminary operation for achieving cross-modal information interaction, bridging the gap for deep fusion of features from the two modalities. Text feature vectors focus on semantic information (such as function, evaluation, and parameters), while visual feature vectors focus on concrete information (such as appearance and usage status). Inputting both simultaneously into the cross-modal attention network allows the network to acquire the core features of both modalities, providing a foundation for establishing semantic connections and achieving information complementarity. Following the input requirements of the cross-modal attention network, the text and visual feature vectors are standardized and adapted to ensure consistent dimensions and format. Then, both are synchronously input into the cross-modal attention network, awaiting subsequent attention interaction processing.
[0083] For example, the obtained text feature vector of "wireless charging 20W" and the obtained visual feature vector of "coffee machine body metal material, top charging panel" are simultaneously input into a cross-modal attention network to prepare for the subsequent establishment of the association between "wireless charging function" and "charging panel visual features".
[0084] S24. By using a cross-modal attention network to perform attention interaction between text feature vectors and visual feature vectors, a unified semantic representation that integrates cross-modal information is generated;
[0085] Specifically, in step S24, the cross-modal attention network can automatically identify the semantic correlation between text feature vectors and visual feature vectors, focusing on key information that corroborates and complements each other in the two modalities, while weakening irrelevant or conflicting secondary information, thus achieving accurate interaction and integration of information from the two modalities. The cross-modal attention network simultaneously analyzes the input text feature vectors and visual feature vectors, calculates the correlation strength between the two modal features, allocates attention resources according to the correlation strength, binds text semantics with corresponding visual features, and finally, through integration processing, generates a unified semantic representation that simultaneously contains the core semantics of the text and the core semantics of the vision and is semantically consistent.
[0086] For example, text feature vectors represent "waterproof function," while visual feature vectors represent "no water seepage and no wetness on the fabric during rain testing." Cross-modal attention networks identify the strong correlation between the two through attentional interactions, deeply integrating the textual semantics of "waterproof function" with the visual semantics of "no water seepage during rain testing" to generate a unified semantic representation of "the product has reliable waterproof function (verified by rain testing)." Similarly, text feature vectors represent "10-hour battery life," while visual feature vectors represent "the device still has power after 10 hours of continuous use in the video." Through interactive integration, the network generates a unified semantic representation of "the product has a 10-hour battery life (provided by actual usage video)."
[0087] This embodiment uses a pre-trained model to accurately extract the core feature vectors of text and visual modalities respectively, and then uses a cross-modal attention network to achieve deep interaction and organic fusion of the two types of features. This successfully breaks the limitation of isolated processing of data from different modalities and generates a unified semantic representation that is semantically complete and closely related, laying a high-quality data foundation for the subsequent accurate analysis of consumer intent.
[0088] Furthermore, in some embodiments, step S24, "to perform attention interaction between text feature vectors and visual feature vectors through a cross-modal attention network to generate a unified semantic representation that integrates cross-modal information," may specifically include:
[0089] S241. In a cross-modal attention network, a multi-head attention mechanism is used to calculate the attention weight distribution between each unit in the text feature vector and each region in the visual feature vector;
[0090] Specifically, in step S241, the multi-head attention mechanism deploys multiple attention heads in parallel to simultaneously capture the correlation between text and visual features from different semantic dimensions, avoiding the omission of correlations caused by single-dimensional analysis and making correlation recognition more comprehensive and accurate. Each unit of the text feature vector refers to the basic unit carrying specific semantics within the text feature vector, corresponding to a single semantic fragment after text segmentation, such as the vector unit corresponding to core semantics like "waterproof," "10-hour battery life," and "affordable price." Each region of the visual feature vector refers to a local area representing specific visual information within the visual feature vector, corresponding to visual elements with independent meaning in an image / video, such as the vector fragment corresponding to functional component areas of a product, key action areas in a usage scenario, and appearance detail areas. For each attention head, the semantic matching degree between each semantic unit in the text feature vector and each visual region in the visual feature vector is calculated. The higher the matching degree, the greater the corresponding attention weight, ultimately forming a two-dimensional attention weight matrix (i.e., weight distribution) covering all text units and visual regions.
[0091] For example, if the text feature vector contains two core semantic units, "waterproof" and "lightweight", and the visual feature vector comes from an outdoor backpack rain test video (containing three visual region features: "no water seepage on the backpack surface", "easy to carry", and "brightly colored backpack"), after calculation through a multi-head attention mechanism, the semantic unit "waterproof" has the highest attention weight with the visual region "no water seepage on the backpack surface", the semantic unit "lightweight" has the highest attention weight with the visual region "easy to carry", and the visual region "brightly colored backpack" has a lower weight with both text semantic units, forming a clear distribution of associated weights.
[0092] S242. Based on the attention weight distribution, the visual feature vectors are weighted and converged to generate a visual context representation that is semantically aligned with the text feature vectors;
[0093] Specifically, in step S242, based on the attention weight distribution, the features of each region in the visual feature vector are weighted, giving higher-weighted visual region features greater influence and weakening the interference of lower-weighted visual region features. Then, these weighted visual region features are integrated through convergence operations to form a structured set of visual information. The core objective of semantic alignment is to ensure that the generated visual context representation accurately corresponds to the core semantics of the text feature vector, avoiding a disconnect between visual features and text semantics, and ensuring that the information from both modalities complements each other around the same consumer-related theme.
[0094] For example, based on the attention weight distribution, the visual region feature of "no water seepage on the backpack surface" is given a high weight (e.g., 0.8), the visual region feature of "the backpack is lightweight and easy to carry" is given a relatively high weight (e.g., 0.7), and the visual region feature of "the backpack is brightly colored" is given a low weight (e.g., 0.2). After weighted aggregation, the generated visual context representation highlights the visual information related to "no water seepage" and "lightweight", which is precisely aligned with the semantics of "waterproof" and "lightweight" in the text feature vector, effectively filtering out visual interference information such as "brightly colored" that is irrelevant to the core semantics of the text.
[0095] S243. Concatenate and project the text feature vectors and visual context representations to form a unified semantic representation;
[0096] Specifically, in step S243, the semantically aligned text feature vector (carrying abstract semantics, such as functional descriptions and evaluation tendencies) and the visual context representation (carrying concrete verification information, such as usage effects and appearance features) are combined according to a preset dimensional order to form a joint feature vector containing complete bimodal information. Through linear or nonlinear projection operations, the combined feature vector is optimized for dimensional unification and semantic fusion, eliminating differences in data distribution and dimensional specifications between the two modal features. This allows for deep coupling of bimodal information, forming a logically consistent and standardized unified semantic representation, facilitating subsequent analysis and processing.
[0097] For example, the text feature vectors corresponding to "waterproof" and "lightweight" are concatenated with the generated visual context representations related to "no water leakage" and "lightweight and portable" to obtain joint features that include "textual semantics + aligned visual information". Then, through projection transformation, these features are transformed into vectors of fixed dimensions, ultimately forming a unified semantic representation that "the product has waterproof function (verified by rain test with no water leakage) and is lightweight and portable". This not only preserves the core semantics of the text but also incorporates empirical visual information, achieving the organic integration of multimodal information.
[0098] This embodiment employs a multi-head attention mechanism to deeply mine the local semantic associations between text feature units and visual feature regions. It achieves accurate alignment of bimodal semantics through weighted convergence, and then completes deep fusion through splicing and projection transformation. This not only strengthens the collaborative association between modalities but also avoids semantic disconnection, significantly improving the completeness, accuracy, and logic of the unified semantic representation.
[0099] Furthermore, in some embodiments, step S21, "parses the text data using a pre-trained text feature extraction model to obtain text feature vectors," may specifically include:
[0100] S211. Identify the language category of text data;
[0101] Specifically, for step S211, the text data in cross-border consumption scenarios originates from different countries and regions, encompassing a variety of languages. By identifying the language category, it ensures that a suitable parsing model is used subsequently, avoiding the problem of general models failing to accurately parse less common or specific languages. The text data covers various types of text information related to cross-border consumption, such as shopping reviews posted by users, descriptions on product detail pages, dialogues with customer service, and promotional information pushed by the platform. Languages may include Indonesian, Arabic, Thai, English, Vietnamese, and other commonly used cross-border languages. Through mature language recognition algorithms, the character features, grammatical structure, and common vocabulary of the text data are analyzed to automatically determine its language category.
[0102] For example, if the text data is “Barang ini sangat cocok untuk Lebaran”, it can be determined to be Indonesian through language recognition.
[0103] S212. Based on the language category, call the corresponding pre-trained text feature extraction model to perform preliminary semantic parsing on the text data to obtain basic semantic vectors;
[0104] Specifically, in step S212, the pre-trained model is a specially trained text feature extraction model for different language categories. These models have learned the grammatical rules, common semantic expressions, and lexical association logic of the corresponding language, enabling them to accurately capture the basic core meaning of the text. The text data with the identified language category is input into the corresponding model. The model processes the text through word segmentation, semantic encoding, and core information extraction to remove redundant information and focus on key content such as product attributes, user needs, and evaluation attitudes. It then transforms these basic semantics into a standardized vector form (i.e., basic semantic vectors).
[0105] For example, for the identified Indonesian text “Barang ini sangat cocok untuk Lebaran”, the Indonesian pre-trained text feature extraction model is called to parse it, extract the basic semantic meaning of “the product is suitable for use in a specific occasion”, and generate the corresponding basic semantic vector.
[0106] S213. Access a pre-built cultural symbol knowledge base to retrieve culturally specific semantic information associated with keywords in the text data;
[0107] Specifically, for step S213, the core content of the cultural symbol knowledge base is: the knowledge base stores culturally specific semantic information corresponding to different cultural regions and languages, including regional customs, religious metaphors, specific meanings of festivals, and culturally exclusive expressions, and this information forms an association mapping with the corresponding keywords and phrases.
[0108] Search logic: Extract core keywords (such as specific festival names, cultural symbol words, and custom-related expressions) from the preliminarily parsed text data, and use these keywords as the search basis to query the cultural symbol knowledge base to obtain the culturally specific semantic information associated with them.
[0109] For example, in the Indonesian text “Barang ini sangat cocok untuk Lebaran”, the keyword “Lebaran” (Ramadan) is extracted. After accessing the cultural symbol knowledge base, the culturally specific semantic information corresponding to the keyword is retrieved, such as “An important religious festival in Indonesia, during which consumers have a habit of concentrated shopping and purchasing gifts.”
[0110] S214. Integrate culture-specific semantic information and basic semantic vectors to generate text feature vectors enhanced with cultural semantics;
[0111] Specifically, in step S214, the vector fusion algorithm organically combines the obtained basic semantic vector (carrying the core literal meaning of the text) with the acquired culturally specific semantic information (carrying the cultural connotations behind the text). This allows the fused vector to retain the basic semantics while incorporating key cultural information, thus expanding and deepening the semantic dimension. The resulting text feature vector comprehensively represents the complete semantics of the text, including basic meanings such as "the product is suitable for use" and "the color is inappropriate," as well as culturally relevant semantics such as "suitable for Ramadan consumption scenarios" and "red involves cultural taboos." This provides more accurate text feature support for subsequent cross-modal fusion and intent parsing.
[0112] For example, by fusing the basic semantic vector of "the product is suitable for use in a specific occasion" with the culturally specific semantic information of "consumers concentrate on shopping and selecting gifts during Ramadan", the generated text feature vector can accurately represent the complete semantic meaning of "the product is suitable for selection as a gift during Ramadan in Indonesia". By fusing the basic semantic vector of "red is inappropriate" with the culturally specific semantic information of "red is associated with taboos in some parts of the Middle East", the generated text feature vector clearly conveys the core meaning of "the red design of the product does not conform to the cultural preferences of the Middle Eastern target market".
[0113] This embodiment first accurately identifies the language category of the text data and calls the appropriate pre-trained model to parse the basic semantics. Then, it integrates specific semantic information from the cultural symbol knowledge base to form a text feature vector enhanced with cultural semantics. This effectively overcomes the problems of incomplete semantic parsing of multilingual texts and insufficient cross-cultural adaptation, and significantly improves the depth and accuracy of text semantic parsing.
[0114] Furthermore, in some embodiments, the construction and updating of the cultural symbol knowledge base may specifically include:
[0115] S2131. Automatically identify and extract expression fragments containing cultural metaphors or specific customs from multilingual and multi-regional unstructured text corpora;
[0116] Specifically, for step S2131, the corpus sources cover unstructured texts related to cross-border consumption from different regions and languages worldwide, including but not limited to e-commerce user reviews, product promotion copy, social media discussions on consumption topics, introductions to regional consumption customs, and articles related to holiday shopping in various languages. The languages can cover multiple languages such as Indonesian, Arabic, Thai, and Vietnamese. Through semantic recognition algorithms, expressions containing cultural metaphors, regional customs, religious connotations, and specific holiday meanings are automatically filtered from the text corpus. These contents must be directly or indirectly related to cross-border consumption behavior, and ultimately extracted to form independent expression fragments.
[0117] For example, the expression “Belanja Lebaran untuk keluarga” (buying Ramadan-related goods for family) was extracted from Indonesian e-commerce reviews. This expression contains the cultural symbol of Ramadan (Lebaran) in Indonesia and the corresponding consumption customs.
[0118] S2132. Clean and label the expression fragments to determine the cultural region and core cultural semantics to which the expression fragments belong;
[0119] Specifically, for step S2132, redundant information (such as irrelevant modifiers, grammatically incorrect content, and repetitive expressions) and noisy data (such as interfering statements unrelated to cultural semantics) in the expression fragment are removed to ensure that the fragment's semantics are concise and its core information is highlighted. Two core annotation methods are used, either manually or automatically: first, identifying the cultural region to which the expression fragment belongs (e.g., Indonesia, a Middle Eastern country, Thailand, Vietnam, etc.); and second, extracting the core cultural semantics of the fragment, namely the cultural connotations, customs, and metaphorical meanings it carries.
[0120] For example, the extracted Indonesian expression fragment "Belanja Lebaran untuk keluarga" was cleaned, and irrelevant interjections were removed to retain the core expression; its cultural region was labeled as "Indonesia", and its core cultural meaning was "Ramadan is an important religious festival in Indonesia, during which there is a concentrated consumption custom of buying gifts for family members".
[0121] S2133. Using a contrastive learning algorithm, expressive fragments and their core cultural semantics are mapped to a unified vector space, forming a structured mapping relationship and storing it in a cultural symbol knowledge base;
[0122] Specifically, in step S2133, the contrastive learning algorithm performs feature learning on the expressive fragments (textual form) and core cultural semantics (semantic description), uncovers the inherent relationship between the two, and transforms them into vector representations in a unified vector space. This ensures that expressive fragments with the same or similar cultural semantics are located close to each other in the vector space, and that vectors of different cultural semantics are clearly distinguishable. The vectors of the expressive fragments, the vectors of the core cultural semantics, and the corresponding cultural region labels are bound together to form a three-dimensional structured mapping relationship of "expressive fragment - core cultural semantics - cultural region." This relationship has a standardized data format and can be quickly retrieved and accessed. All structured mapping relationships are organized and stored according to preset classification rules (such as by language type, cultural region, consumption scenario, etc.) to construct an efficiently accessible cultural symbol knowledge base.
[0123] For example, through a contrastive learning algorithm, the expression fragment “Belanja Lebaran untuk keluarga” and the core semantics of “Ramadan family gift consumption customs” are mapped into corresponding vectors, and the two vectors form a strong correlation. Then, the “Indonesia” cultural region label is bound to form a structured mapping relationship and stored. When “Lebaran” related expressions are retrieved later, the corresponding cultural semantics and regional information can be quickly matched. Similarly, the expression fragments and semantics related to “practical gifts for Eid al-Adha” will also form a correlation vector and be stored.
[0124] S2134. Based on new cross-border consumption data and user feedback, dynamically calibrate and expand the mapping relationships in the cultural symbol knowledge base;
[0125] Specifically, for step S2134, new cross-border consumption data (such as newly added multilingual user reviews and consumption copywriting from new regions) and user feedback (such as feedback on consumption disputes caused by cultural semantic misunderstandings and feedback on changes in regional consumption customs) are continuously collected. New expressive fragments containing cultural metaphors or customs are extracted from these data. The above process is repeated to supplement the knowledge base with new structured mapping relationships, expanding the coverage of the knowledge base. Existing mapping relationships in the knowledge base are periodically verified. If changes in cultural customs (such as a change in the semantic interpretation of a symbol in a certain region) or user feedback indicating semantic misjudgment (such as incorrect labeling of the core cultural semantics of an expressive fragment) are found, the corresponding vector mapping relationships, core semantic descriptions, or cultural region labels are adjusted promptly to ensure the accuracy of the knowledge base information.
[0126] For example, if a new Indonesian comment is subsequently collected, such as "Hadiah Lebaran kini lebih fokus padakecanggihan" (Ramadan gifts are now more focused on technology), then this expression fragment is extracted, labeled with the cultural region "Indonesia," and the core semantic "Indonesian Ramadan gift consumption customs have shifted from traditional to a preference for technology." After generating a correlation vector through a comparative learning algorithm, it is expanded into the knowledge base. If user feedback is received that the semantic interpretation of "blue packaging" in a certain Middle Eastern region has changed from "ordinary" to "auspicious," then the core cultural semantic mapping relationship of the relevant expression fragments of "blue packaging" in that region in the knowledge base is calibrated.
[0127] This embodiment collects a wide range of cross-cultural expressions from multiple languages and regions, cleans and annotates them, and constructs a structured cultural symbol knowledge base through comparative learning and mapping. It also relies on new data and user feedback to achieve dynamic calibration and expansion, creating a comprehensive, accurate, and timely cultural semantic support system that significantly reduces the risk of cross-cultural semantic misunderstandings.
[0128] Furthermore, in some embodiments, step S3, "identifying the scenario type to which the target cross-border consumption scenario belongs, and determining the contribution weights corresponding to the text modality and visual modality in the unified semantic representation based on the scenario type," may specifically include:
[0129] S31. Perform scene feature analysis on the unified semantic representation to identify the scene type to which the target cross-border consumption scene belongs;
[0130] Specifically, for step S31, the unified semantic representation contains key information directly related to the consumption scenario, including information types in textual semantics (such as product parameter descriptions, real-time interactive scripts, user feedback, etc.) and content forms in visual semantics (such as static product detail displays, dynamic usage demonstrations, and live-streaming explanations). These features collectively constitute the core basis for scenario recognition. By comprehensively analyzing the textual and visual scene features in the unified semantic representation and matching them with a pre-set scenario feature template library (containing typical feature combinations of different scenarios), the specific type of the target cross-border consumption scenario is determined. Common scenario types include, but are not limited to, product detail page browsing scenarios, live-streaming sales scenarios, user review analysis scenarios, and promotional activity scenarios.
[0131] For example, if the unified semantic representation uses textual semantics to describe product materials, functional parameters, specifications, and dimensions in detail, and visual semantics to focus on static real-life photos of the product from multiple angles, then it can be identified as a "product details page browsing scenario" through scene feature analysis. If the unified semantic representation uses textual semantics to include the host's real-time explanations and audience interaction questions, and visual semantics to focus on the host's demonstration of product usage and dynamic visuals showcasing the product's effects in real time, then it can be identified as a "live-streaming e-commerce scenario".
[0132] S32. Input the scene type into a pre-trained dynamic weight allocation model;
[0133] Specifically, for step S32, the pre-trained dynamic weight allocation model has been trained with a large amount of cross-border consumption scenario data. It possesses the ability to deeply learn the differences in text and visual modal information requirements across different scenario types, and can output a modal weight allocation scheme adapted to the input scenario type. The identified scenario type (such as "product details page browsing scenario" or "live-streaming e-commerce scenario") is standardized according to the model's required format and then input into the pre-trained dynamic weight allocation model, triggering the model's weight calculation process. For example, the identified "product details page browsing scenario" can be used as input into the pre-trained dynamic weight allocation model; or the "live-streaming e-commerce scenario" can be standardized and input into the model, waiting for the model to calculate the corresponding modal weights based on the scenario characteristics.
[0134] S33. Through a dynamic weight allocation model, the text modality contribution coefficient and visual modality contribution coefficient adapted to the scene type are calculated and output in real time as contribution weights;
[0135] Specifically, in step S33, after receiving the scene type input, the dynamic weight allocation model calls the scene weight mapping relationship formed by internal pre-training and, combined with the model's built-in optimization algorithm, quickly calculates the contribution of the text modality and visual modality to the consumer intent parsing in that scene, presenting it in the form of coefficients. The text modality contribution coefficient represents the importance of text semantic information in that scene, and the visual modality contribution coefficient represents the importance of visual semantic information. The sum of the two is usually 1 (or a fixed proportion sum) to ensure the rationality of the weight allocation.
[0136] For example, in the input scenario of "browsing product details page", the model calculates and outputs a text modality contribution coefficient of 0.7 and a visual modality contribution coefficient of 0.3 in real time based on the cognitive understanding formed during training (in this scenario, users pay more attention to the core information of the text description). This means that the weight of text information in this scenario is 70%, and the weight of visual information is 30%. In the scenario of "live-streaming e-commerce", the model determines that users rely more on the visual display of the product and outputs a text modality contribution coefficient of 0.55 and a visual modality contribution coefficient of 0.45. This means that the weight of visual information is increased to 45%, and the weight of text information is adjusted to 55%, achieving accurate adaptation to the needs of the scenario.
[0137] This embodiment accurately identifies scene types by analyzing scene features in the unified semantic representation, and then outputs modal contribution coefficients adapted to the scene in real time by a pre-trained dynamic weight allocation model. This realizes the upgrade of modal weights from fixed allocation to scene-based dynamic adaptation, effectively highlighting core modal information, weakening secondary interference, and improving the pertinence and effectiveness of intent parsing.
[0138] Furthermore, in some embodiments, the training method of the dynamic weight allocation model may specifically include:
[0139] S331. Construct a training sample set, where each training sample contains a scene type label, multimodal data, and corresponding real consumer behavior feedback;
[0140] Specifically, for step S331, each sample in the training sample set must fully cover the three core dimensions of "scenario-data-feedback" to form a closed-loop training basis. The scenario type label clearly defines the consumption scenario to which the data belongs, multimodal data provides the basic materials for model learning, and real consumer behavior feedback serves as the core basis for judging whether the weight allocation is reasonable. The scenario type label is based on the classification of actual cross-border consumption scenarios, clearly marking the scenario attributes corresponding to each sample, such as "product details page browsing scenario," "live streaming sales scenario," and "user review analysis scenario." The multimodal data corresponds to the scenario type label and includes the actual text data (such as product parameter descriptions, user consultation scripts, and anchor explanations) and visual data (such as real product photos, usage videos, and live streaming footage) generated in that scenario. Real consumer behavior feedback refers to the actual behavioral results generated by consumers in that scenario, which can directly reflect their intentions and needs, such as purchase behavior, favorite operations, return requests, negative reviews, and repeated browsing records.
[0141] For example, a training sample for a "live-streaming e-commerce scenario" is constructed, labeled as "live-streaming e-commerce scenario". The multimodal data includes text content of the host explaining "this skincare product has strong moisturizing power and is suitable for dry skin", as well as live-streaming footage of the host demonstrating the product's texture and application effect. Real consumer behavior feedback includes actual behavioral data such as "60 out of 100 users who watched the live stream placed an order" and "20 people left comments asking about the duration of moisturizing effect". Another example is a sample for a "product details page browsing scenario", labeled as the corresponding scenario. The multimodal data includes text descriptions of the product's "waterproof rating IPX8" and waterproof test images, with behavioral feedback such as "80 out of 500 users added the product to their shopping cart" and "30 people inquired about after-sales warranty policies".
[0142] S332. Construct a dynamic weight allocation model to be trained, setting the scene type as input and the modal contribution coefficient as output;
[0143] Specifically, for step S332, based on the modal weight allocation requirements of cross-border consumption scenarios, a model architecture with semantic understanding and coefficient calculation capabilities is designed to ensure that the model can receive scenario type information and output reasonable modal contribution coefficients through internal calculations. The model input is standardized scenario type information, which must be consistent with the scenario type label format in the training sample set (such as uniform text labels or encoding format) to ensure that the model can accurately identify and match. The model output consists of two quantified modal contribution coefficients, corresponding to the weight proportions of text modality and visual modality, respectively. The sum of the coefficients must conform to logic (such as a total of 1 or a fixed proportion range) to ensure the feasibility of weight allocation.
[0144] S333. With the optimization goal of maximizing the accuracy of consumer behavior intention prediction or business conversion rate, a reinforcement learning algorithm is used to iteratively train the dynamic weight allocation model;
[0145] Specifically, for step S333, maximizing the accuracy of consumer behavior intent prediction means that when the modal weights output by the model are applied to subsequent intent parsing, the parsing results match the consumer's true intent to the highest degree; maximizing the commercial conversion rate means that the weights output by the model can highlight the core modal information in the scenario, helping merchants accurately reach consumer needs, thereby improving the commercial conversion effect of product purchases, service subscriptions, etc.; the model training process is regarded as an interactive process of reinforcement learning, with the optimization objective as the reward signal. If the weights output by the model make the intent prediction accurate or the conversion rate improved, a positive reward is given; if the prediction deviation is large or the conversion rate is low, a negative feedback is given; the model continuously adjusts its internal parameters according to the reward signal; during the iterative training process, samples from the training sample set are continuously input into the model to obtain the modal contribution coefficients output by the model, and the reward value is calculated by combining the real consumer behavior feedback in the samples. The model parameters are adjusted based on the reward value, and this process is repeated until the weights output by the model are stable and the optimization objective reaches the preset standard.
[0146] This embodiment constructs a high-quality training sample set based on "scenario type - multimodal data - real consumer behavior feedback". With the optimization goal of consumer intent prediction accuracy or business conversion rate, the dynamic weight allocation model is iteratively trained through reinforcement learning algorithm so that the weights output by the model not only fit the core needs of the scenario, but also meet the business operation goals, providing reliable and efficient weight support for subsequent intent parsing.
[0147] Furthermore, in some embodiments, step S4, "parses the weighted unified semantic representation according to the contribution weight to generate consumer behavior intent tags corresponding to the target cross-border consumption scenario," may specifically include:
[0148] S41. Based on the contribution weights, the text modal components and visual modal components in the unified semantic representation are weighted separately to obtain the weighted semantic representation;
[0149] Specifically, in step S41, the text modality contribution weight and the visual modality contribution weight quantify the importance of the two modalities in the current cross-border consumption scenario, serving as the core basis for adjusting information influence. For the independent text modality components (carrying semantics such as product descriptions and user reviews) and visual modality components (carrying information such as product appearance and usage effects) in the unified semantic representation, each is multiplied by its corresponding contribution weight coefficient. The higher the weight coefficient, the stronger the influence of that modality component in the overall semantics; the lower the weight coefficient, the weaker the influence. Ultimately, this is integrated to form a weighted semantic representation that highlights the core information.
[0150] For example, if the target scenario is a "product details page browsing scenario", the corresponding text modal contribution weight is 0.7 and the visual modal contribution weight is 0.3. During weighted processing, the influence of text modal components such as "waterproof rating IPX7" and "20-hour battery life" in the unified semantic representation will be amplified, while the interference of visual modal components such as "product color is white" will be appropriately weakened. If the target scenario is a "live-streaming e-commerce scenario", the text modal contribution weight is 0.55 and the visual modal contribution weight is 0.45. Then, the influence of the visual modal component "smooth operation of the host demonstrating the product" will be strengthened, while the core information of the text modal component "easy to operate" will be retained, forming a weighted semantic representation adapted to the scenario.
[0151] S42. Input the weighted semantic representation into the trained multimodal large language model;
[0152] Specifically, for step S42, the multimodal large language model has been trained with a large amount of multimodal data related to cross-border consumption, and has the ability to simultaneously understand the semantic fusion of text and vision. It can accurately capture the core information reinforced in the weighted semantic representation, as well as the intrinsic relationship between the two modal information. The obtained weighted semantic representation is standardized according to the format required by the model (ensuring that the dimensions, data format and model input specifications are consistent), and then input into the trained multimodal large language model to trigger the model's semantic parsing and intent reasoning process.
[0153] For example, the weighted semantic representation that strengthens the core text information in the "product details page browsing scenario" or the weighted semantic representation that balances the core text and visual information in the "live streaming e-commerce scenario" can be standardized and input into a multimodal large language model. The model will then perform a deep interpretation of the input semantic information based on the cognition formed during training.
[0154] S43. Reason about the weighted semantic representation using a multimodal large language model, and output one or more consumer behavior intention labels and their corresponding confidence scores;
[0155] Specifically, in step S43, the multimodal large language model analyzes the weighted semantic representation of the input layer by layer, mining the underlying consumer's core needs, purchasing tendencies, and key concerns, etc., and extracts clear intent labels by combining the "semantic-intent" mapping relationship learned during model training. A label is a concise summary of the consumer's behavioral intent, such as "focusing on product cost-effectiveness," "emphasizing functional practicality," "caring about appearance design," or "worried about logistics speed." Confidence is a quantitative assessment of the model's accuracy of the intent label (usually represented by a value between 0 and 1). A higher value indicates a higher degree of match between the intent label and the consumer's true intent, and a more reliable result.
[0156] For example, after weighted semantic representation inference for the "product details page browsing scenario", the model outputs two intent labels: "focus on functional usability" (confidence 0.92) and "concern about price reasonableness" (confidence 0.85), indicating that the consumer's core intent is to focus on product functions and price. After weighted semantic representation inference for the "live-streaming e-commerce scenario", the model outputs intent labels: "preference for ease of use" (confidence 0.88) and "focus on appearance and texture" (confidence 0.83), accurately reflecting the consumer's core concerns in the live-streaming scenario.
[0157] This embodiment first strengthens the unified semantic representation based on the contribution weight of scenario adaptation, then inputs it into a trained multimodal large language model for deep reasoning, and finally outputs consumer behavior intent tags with quantified confidence. This not only achieves accurate and quantitative presentation of intent parsing, but also provides clear and reliable reference for cross-border e-commerce operation decisions.
[0158] Furthermore, in some embodiments, the method may further include:
[0159] S51. Classify and quantify the intensity of consumer behavioral intention tags according to multiple predefined preference dimensions;
[0160] Specifically, for step S51, based on the core needs of cross-border consumption scenarios, pre-defined preference dimensions covering key consumer decision-making factors are established. These dimensions must be universal and practical, comprehensively encompassing common user concerns in cross-border consumption. Common predefined preference dimensions include, but are not limited to, price sensitivity, functional requirement intensity, appearance design preference, service experience focus, and cultural compatibility requirements. The core connotation of each consumer behavioral intent tag is analyzed one by one, and based on the user focus it points to, it is categorized into the corresponding preference dimension, ensuring that each tag accurately matches a unique or most relevant dimension, avoiding ambiguity in classification. Based on the confidence level corresponding to the intent tag, combined with the intensity of preference reflected by the tag, standardized quantitative rules (such as a 0-10 scale, 0-1 range, etc.) are used to score the intensity of each tag on the corresponding dimension. The higher the confidence level and the stronger the preference expression, the higher the quantitative score, ultimately forming specific quantitative values for each dimension.
[0161] For example, if a consumer's behavioral intention is labeled "focus on the cost-effectiveness of the product" (confidence level 0.88), its core focus is on the balance between price and value, and it is categorized under the "price sensitivity" dimension, quantified as 8.5 points based on confidence level and preference intensity; if the label is "value for waterproof and long battery life features" (confidence level 0.92), it is categorized under the "functional demand intensity" dimension, quantified as 9.0 points; if the label is "believes that product packaging does not conform to local festival customs" (confidence level 0.80), it is categorized under the "cultural compatibility requirements" dimension, quantified as 7.8 points; and if the label is "likes a minimalist product appearance" (confidence level 0.75), it is categorized under the "appearance design preference" dimension, quantified as 7.2 points.
[0162] S52. Based on the intensity quantification results of each preference dimension, generate an intent analysis map presented in the form of a heatmap;
[0163] Specifically, for step S52, the heatmap uses predefined preference dimensions as the core coordinate axis. For example, the horizontal axis is set as dimensions such as "price sensitivity," "functional requirement intensity," and "appearance design preference." Quantitative scores are used as intensity indicators, and color gradients are used to represent the preference intensity of each dimension. The higher the quantitative score, the darker the corresponding area's color; for example, dark red represents high intensity, and light yellow represents low intensity, forming an intuitive visual contrast. The quantitative intensity values of each preference dimension are imported into the heatmap generation tool. According to the preset color mapping rules and chart layout, an intent analysis map containing elements such as dimension names, intensity color labels, and numerical scales is automatically generated, ensuring the map information is complete and visually clear.
[0164] For example, based on the quantitative results, "functional requirement intensity" (9.0 points) corresponds to the dark red area, "price sensitivity" (8.5 points) corresponds to the red area, "cultural compatibility requirements" (7.8 points) corresponds to the light red area, and "appearance design preference" (7.2 points) corresponds to the light yellow area. In the generated heat map, the color depth of each dimension is clearly distinguished, and users can quickly identify that in this cross-border consumption scenario, consumers pay the most attention to "functional requirements", followed by "price sensitivity", and pay relatively less attention to "appearance design".
[0165] S53. Integrate the intent analysis map and its corresponding scene types and cultural region information to generate a consumer intent visualization report;
[0166] Specifically, step S53 integrates three key components: first, the generated intent analysis heatmap (visually displaying the distribution of preference intensity); second, the specific type of the target cross-border consumption scenario (e.g., live-streaming e-commerce, product detail page browsing, etc., clarifying the analysis background); and third, the corresponding cultural and regional information (e.g., Indonesia, Arabic-speaking regions of the Middle East, Thailand, etc., adapting to the cultural attributes of the cross-border scenario). The visualization report should have a clear logical structure, typically including modules such as scenario and region overview, interpretation of the intent analysis heatmap, summary of core preferences, and decision-making suggestions. The report should support multilingual display (e.g., Arabic, Vietnamese, Indonesian, etc.) to meet the needs of operators in different regions, and may include supplementary content such as data tables and text descriptions to make the report more readable and instructive.
[0167] For example, by integrating the intent analysis heatmap of "Livestream E-commerce Scenarios in the Middle East," a description of the scenario type, and cultural background information of the Middle East, a visual report in Arabic was generated. In the report, the heatmap clearly shows "functional usability" (dark red) and "cultural compatibility" (red) as the core preference dimensions; the textual interpretation explains that "Middle Eastern consumers prioritize the practical functions and cultural compatibility of products (such as avoiding taboo elements) in livestream e-commerce scenarios"; and the decision-making recommendations suggest that "product promotion should focus on demonstrating practical functions, and packaging and copywriting should conform to Middle Eastern cultural customs," providing direct reference for operational decisions.
[0168] This embodiment categorizes and quantifies discrete consumer behavior intent tags according to predefined preference dimensions, generating an intuitive heatmap-style intent analysis map. It then integrates scenario type and cultural region information to form a structured and visualized report. This not only makes complex consumer intent data easy to understand and use, significantly reducing interpretation costs, but also significantly shortens the decision-making response time of cross-border e-commerce, adapting to the actual needs of global operations.
[0169] In summary, compared with existing technologies, the cross-border consumer behavior dynamic analysis method based on a large language model provided in this embodiment improves the accuracy and scenario adaptability of cross-border consumer behavior intent analysis by extracting and fusing features from textual and visual multimodal data in cross-border consumption scenarios to form a unified semantic representation, and dynamically assigning modal contribution weights based on scenario type to accurately generate consumer behavior intent tags. This effectively solves the problem of low intent analysis accuracy caused by insufficient utilization of multimodal information or insufficient scenario adaptability in existing technologies, thus providing reliable support for the refined operation of cross-border e-commerce.
[0170] To facilitate better implementation of the cross-border consumption behavior dynamic analysis method based on a large language model according to the embodiments of this application, this application also provides a cross-border consumption behavior dynamic analysis device based on a large language model, which is based on the aforementioned cross-border consumption behavior dynamic analysis method based on a large language model. The meanings of the terms used are the same as in the aforementioned cross-border consumption behavior dynamic analysis method based on a large language model, and specific implementation details can be found in the descriptions in the method embodiments.
[0171] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of the cross-border consumption behavior dynamic analysis device based on a large language model provided in an embodiment of this application. Specifically, the device may include a data acquisition module 201, a data processing module 202, a dynamic weighting module 203, and an intent recognition module 204, as follows:
[0172] The data acquisition module 201 is used to acquire multimodal raw data of the target cross-border consumption scenario from multiple data sources. The multimodal raw data includes at least text data and visual data related to the target product or target service.
[0173] The data processing module 202 is used to perform feature extraction and fusion processing on multimodal raw data to generate a unified semantic representation that includes textual and visual semantics;
[0174] The dynamic weight module 203 is used to identify the scenario type to which the target cross-border consumption scenario belongs, and to determine the contribution weights of the text modality and visual modality in the unified semantic representation based on the scenario type.
[0175] The intent recognition module 204 is used to parse the weighted unified semantic representation according to the contribution weight and generate consumer behavior intent tags corresponding to the target cross-border consumption scenario.
[0176] Furthermore, in some embodiments, the data processing module 202 is specifically used for:
[0177] The text data is parsed using a pre-trained text feature extraction model to obtain text feature vectors;
[0178] Visual feature vectors are obtained by extracting features from visual data using a visual feature extraction model.
[0179] Text feature vectors and visual feature vectors are input into a cross-modal attention network;
[0180] By using a cross-modal attention network to interact with textual and visual feature vectors, a unified semantic representation that integrates cross-modal information is generated.
[0181] Furthermore, in some embodiments, the data processing module 202 is specifically used for:
[0182] A multi-head attention mechanism is employed in the cross-modal attention network to calculate the attention weight distribution between each unit in the text feature vector and each region in the visual feature vector;
[0183] The visual feature vectors are weighted and converged according to the attention weight distribution to generate a visual context representation that is semantically aligned with the text feature vectors;
[0184] Text feature vectors and visual context representations are concatenated and projected to form a unified semantic representation.
[0185] Furthermore, in some embodiments, the data processing module 202 is specifically used for:
[0186] Identify the language category of the text data; based on the language category, call the corresponding pre-trained text feature extraction model to perform preliminary semantic parsing of the text data to obtain basic semantic vectors; access a pre-built cultural symbol knowledge base to retrieve culturally specific semantic information associated with keywords in the text data; fuse the culturally specific semantic information and the basic semantic vectors to generate a text feature vector enhanced with cultural semantics.
[0187] Furthermore, in some embodiments, the construction and updating of the cultural symbol knowledge base includes:
[0188] Automatically identify and extract expressions containing cultural metaphors or specific customs from unstructured text corpora from multiple languages and regions;
[0189] The expression fragments are cleaned and labeled to determine the cultural region and core cultural semantics to which they belong;
[0190] By employing a contrastive learning algorithm, expressive fragments and their core cultural semantics are mapped to a unified vector space, forming a structured mapping relationship and storing it in a cultural symbol knowledge base;
[0191] Based on new cross-border consumption data and user feedback, the mapping relationships in the cultural symbol knowledge base are dynamically calibrated and expanded.
[0192] Furthermore, in some embodiments, the dynamic weighting module 203 is specifically used for:
[0193] By performing scene feature analysis on the unified semantic representation, the scene type to which the target cross-border consumption scene belongs can be identified;
[0194] The scene type is input into a pre-trained dynamic weight allocation model;
[0195] The dynamic weight allocation model calculates and outputs text modality contribution coefficients and visual modality contribution coefficients that are adapted to the scene type in real time, which are then used as contribution weights.
[0196] Furthermore, in some embodiments, the training method of the dynamic weight allocation model includes:
[0197] Construct a training sample set, where each training sample contains a scenario type label, multimodal data, and corresponding real consumer behavior feedback;
[0198] Construct a dynamic weight allocation model to be trained, with scene type as input and modal contribution coefficient as output;
[0199] With the optimization goal of maximizing the accuracy of consumer behavior intention prediction or business conversion rate, a reinforcement learning algorithm is used to iteratively train the dynamic weight allocation model.
[0200] Furthermore, in some embodiments, the intent recognition module 204 is specifically used for:
[0201] Based on the contribution weights, the text modal components and visual modal components in the unified semantic representation are weighted separately to obtain a weighted semantic representation;
[0202] The weighted semantic representation is input into the trained multimodal large language model;
[0203] By reasoning about the weighted semantic representation using a multimodal large language model, one or more consumer behavioral intent labels and their corresponding confidence scores are output.
[0204] Furthermore, in some embodiments, the apparatus further includes a report generation module 205, specifically used for:
[0205] Consumer behavioral intent tags are categorized and their intensity quantified according to multiple predefined preference dimensions;
[0206] Based on the intensity quantification results of each preference dimension, an intent analysis map is generated in the form of a heatmap.
[0207] By integrating the intent analysis map and its corresponding scenario types and cultural region information, a consumer intent visualization report is generated.
[0208] For specific limitations regarding the cross-border consumption behavior dynamic analysis device based on large language models, please refer to the limitations of the cross-border consumption behavior dynamic analysis method based on large language models mentioned above, which will not be repeated here. Each module in the aforementioned cross-border consumption behavior dynamic analysis device based on large language models can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0209] The cross-border consumer behavior dynamic analysis device based on a large language model provided in this embodiment extracts and fuses features from textual and visual multimodal data in cross-border consumption scenarios to form a unified semantic representation. It also dynamically allocates modal contribution weights based on scenario type to accurately generate consumer behavior intent tags. This improves the accuracy and scenario adaptability of cross-border consumer behavior intent analysis and effectively solves the problem of low intent analysis accuracy caused by insufficient utilization of multimodal information or insufficient scenario adaptability in existing technologies. Thus, it provides reliable support for the refined operation of cross-border e-commerce.
[0210] Furthermore, embodiments of this application also provide an electronic device, such as... Figure 4 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:
[0211] The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0212] The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the electronic device. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 301.
[0213] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and a dynamic analysis method for cross-border consumer behavior based on a large language model by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.
[0214] The electronic device also includes a power supply 303 that supplies power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0215] The electronic device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0216] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302 to realize various functions, as follows:
[0217] Acquire multimodal raw data from multiple data sources for the target cross-border consumption scenario. The multimodal raw data includes at least textual and visual data related to the target product or service. Perform feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation containing textual and visual semantics. Identify the scenario type to which the target cross-border consumption scenario belongs, and determine the contribution weights of the textual and visual modalities in the unified semantic representation based on the scenario type. Parse the weighted unified semantic representation according to the contribution weights to generate consumer behavior intent tags corresponding to the target cross-border consumption scenario.
[0218] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0219] This application embodiment extracts and fuses features from textual and visual multimodal data in cross-border consumption scenarios to form a unified semantic representation. It also dynamically allocates modal contribution weights based on scenario type to accurately generate consumer behavior intent tags. This improves the accuracy and scenario adaptability of cross-border consumption behavior intent analysis and effectively solves the problem of low intent analysis accuracy caused by insufficient utilization of multimodal information or insufficient scenario adaptability in the prior art. This provides reliable support for the refined operation of cross-border e-commerce.
[0220] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0221] To this end, embodiments of this application provide a storage medium storing multiple instructions that can be loaded by a processor to execute steps in any of the cross-border consumption behavior dynamic analysis methods based on large language models provided in embodiments of this application. For example, the instructions can execute the following steps:
[0222] Acquire multimodal raw data from multiple data sources for the target cross-border consumption scenario. The multimodal raw data includes at least textual and visual data related to the target product or service. Perform feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation containing textual and visual semantics. Identify the scenario type to which the target cross-border consumption scenario belongs, and determine the contribution weights of the textual and visual modalities in the unified semantic representation based on the scenario type. Parse the weighted unified semantic representation according to the contribution weights to generate consumer behavior intent tags corresponding to the target cross-border consumption scenario.
[0223] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0224] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0225] Since the instructions stored in the storage medium can execute the steps in any of the cross-border consumption behavior dynamic analysis methods based on large language models provided in the embodiments of this application, the beneficial effects that any of the cross-border consumption behavior dynamic analysis methods based on large language models provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0226] The foregoing has provided a detailed description of a method and apparatus for dynamic analysis of cross-border consumption behavior based on a large language model, as provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for dynamic analysis of cross-border consumption behavior based on a large language model, characterized in that, include: Acquire multimodal raw data from multiple data sources for a target cross-border consumption scenario, wherein the multimodal raw data includes at least textual and visual data related to the target product or service; The process involves feature extraction and fusion of the multimodal raw data to generate a unified semantic representation encompassing both textual and visual semantics. This includes: identifying the language category of the text data; based on the language category, calling a corresponding pre-trained text feature extraction model to perform preliminary semantic parsing of the text data, obtaining a basic semantic vector; accessing a pre-constructed cultural symbol knowledge base to retrieve culturally specific semantic information associated with keywords in the text data; fusing the culturally specific semantic information and the basic semantic vector to generate a text feature vector enhanced with cultural semantics. The construction and updating of the cultural symbol knowledge base includes: automatically identifying and extracting expression fragments containing cultural metaphors or specific customs from unstructured text corpora from multiple languages and regions; cleaning and labeling the expression fragments to determine their cultural region and core cultural semantics; using a contrastive learning algorithm to map the expression fragments and their core cultural semantics to a unified vector space, forming a structured mapping relationship and storing it in the cultural symbol knowledge base; and dynamically calibrating and expanding the mapping relationship in the cultural symbol knowledge base based on new cross-border consumption data and user feedback. Identify the scenario type to which the target cross-border consumption scenario belongs, and determine the contribution weights of the text modality and visual modality in the unified semantic representation based on the scenario type; The weighted unified semantic representation is parsed based on the contribution weight to generate consumer behavior intent tags corresponding to the target cross-border consumption scenario.
2. The method for dynamic analysis of cross-border consumption behavior based on a large language model according to claim 1, characterized in that, The step of performing feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation containing textual and visual semantics includes: The text data is parsed using a pre-trained text feature extraction model to obtain text feature vectors; The visual data is used to extract features through a visual feature extraction model to obtain a visual feature vector. The text feature vector and the visual feature vector are input into a cross-modal attention network; The cross-modal attention network performs attention interaction on the text feature vector and the visual feature vector to generate a unified semantic representation that integrates cross-modal information.
3. The method for dynamic analysis of cross-border consumption behavior based on a large language model according to claim 2, characterized in that, The step of generating a unified semantic representation that integrates cross-modal information by performing attention interaction between the text feature vector and the visual feature vector through the cross-modal attention network includes: In the cross-modal attention network, a multi-head attention mechanism is employed to calculate the attention weight distribution between each unit in the text feature vector and each region in the visual feature vector; The visual feature vectors are weighted and converged according to the attention weight distribution to generate a visual context representation that is semantically aligned with the text feature vectors; The text feature vector and the visual context representation are concatenated and projected to form the unified semantic representation.
4. The method for dynamic analysis of cross-border consumption behavior based on a large language model according to claim 1, characterized in that, The step of identifying the scenario type to which the target cross-border consumption scenario belongs, and determining the contribution weights of the text modality and visual modality in the unified semantic representation based on the scenario type, includes: Scene feature analysis is performed on the unified semantic representation to identify the scene type to which the target cross-border consumption scene belongs; The scene type is input into a pre-trained dynamic weight allocation model; The dynamic weight allocation model calculates and outputs the text modality contribution coefficient and visual modality contribution coefficient, which are adapted to the scene type, in real time, as the contribution weight.
5. The method for dynamic analysis of cross-border consumption behavior based on a large language model according to claim 4, characterized in that, The training methods for the dynamic weight allocation model include: Construct a training sample set, where each training sample contains a scenario type label, multimodal data, and corresponding real consumer behavior feedback; Construct a dynamic weight allocation model to be trained, with scene type as input and modal contribution coefficient as output; With the optimization goal of maximizing the accuracy of consumer behavior intention prediction or business conversion rate, the dynamic weight allocation model is iteratively trained using a reinforcement learning algorithm.
6. The method for dynamic analysis of cross-border consumption behavior based on a large language model according to claim 1, characterized in that, The step of parsing the weighted unified semantic representation based on the contribution weight to generate consumer behavior intent tags corresponding to the target cross-border consumption scenario includes: Based on the contribution weights, the text modal components and visual modal components in the unified semantic representation are weighted respectively to obtain a weighted semantic representation; The weighted semantic representation is input into the trained multimodal large language model; The multimodal large language model is used to infer the weighted semantic representation and output one or more consumer behavior intent labels and their corresponding confidence scores.
7. The method for dynamic analysis of cross-border consumption behavior based on a large language model according to claim 1, characterized in that, The method further includes: The consumer behavior intent tags are classified and their intensity is quantified according to multiple predefined preference dimensions; Based on the intensity quantification results of each preference dimension, an intent analysis map is generated in the form of a heatmap. The intent analysis map and its corresponding scene types and cultural region information are integrated to generate a consumer intent visualization report.
8. A dynamic analysis device for cross-border consumption behavior based on a large language model, characterized in that, include: The data acquisition module is used to acquire multimodal raw data of the target cross-border consumption scenario from multiple data sources. The multimodal raw data includes at least text data and visual data related to the target product or target service. The data processing module is used to perform feature extraction and fusion processing on the multimodal raw data to generate a unified semantic representation containing textual and visual semantics. This includes: identifying the language category of the text data; based on the language category, calling the corresponding pre-trained text feature extraction model to perform preliminary semantic parsing on the text data to obtain a basic semantic vector; accessing a pre-built cultural symbol knowledge base to retrieve culturally specific semantic information associated with keywords in the text data; fusing the culturally specific semantic information and the basic semantic vector to generate the text feature vector enhanced with cultural semantics. The construction and updating of the cultural symbol knowledge base includes: automatically identifying and extracting expression fragments containing cultural metaphors or specific customs from unstructured text corpora from multiple languages and regions; cleaning and labeling the expression fragments to determine the cultural region and core cultural semantics to which the expression fragments belong; using a contrastive learning algorithm to map the expression fragments and their core cultural semantics to a unified vector space, forming a structured mapping relationship and storing it in the cultural symbol knowledge base; and dynamically calibrating and expanding the mapping relationship in the cultural symbol knowledge base based on new cross-border consumption data and user feedback. The dynamic weighting module is used to identify the scenario type to which the target cross-border consumption scenario belongs, and to determine the contribution weights of the text modality and visual modality in the unified semantic representation based on the scenario type. The intent recognition module is used to parse the weighted unified semantic representation according to the contribution weight, and generate consumer behavior intent tags corresponding to the target cross-border consumption scenario.
Citation Information
Patent Citations
Multi-modal language learning auxiliary system and method based on artificial intelligence
CN120688510A
User behavior prediction system and method based on multi-modal data fusion
CN120832498A