Live commodity explanation text generation method and device and computer readable medium

By generating compliant explanation scripts through multimodal data processing models and hybrid expert models, the problems of low efficiency, unstable quality, and lack of personalization in live explanations have been solved, enabling real-time personalized and compliant live explanations and improving conversion rates and efficiency.

CN121174019BActive Publication Date: 2026-05-01LINGKE (HANGZHOU) NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LINGKE (HANGZHOU) NETWORK TECH CO LTD
Filing Date
2025-11-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Currently, live-streaming product descriptions rely on manually written scripts, which suffers from low efficiency, unstable quality, lack of personalization, and poor real-time performance. Furthermore, existing automated script generation tools lack the ability to deeply integrate multi-source data and optimize in real time, resulting in limited improvements in live-streaming efficiency and conversion rates.

Method used

Employing a multimodal data processing model and a hybrid expert model, the system generates compliant explanation scripts by labeling product features on the training dataset and updating and optimizing the model in real time. Combined with live stream comments and performance data, it generates personalized and compliant real-time explanation scripts.

Benefits of technology

It significantly reduces the script preparation time for broadcasters, improves the efficiency and quality of live broadcast explanations, increases conversion rates, reduces the risk of violations, and continuously optimizes the generation effect of explanation content through a closed-loop training mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121174019B_ABST
    Figure CN121174019B_ABST
Patent Text Reader

Abstract

The application discloses a live commodity explanation text generation method and device and a computer readable medium. A historical live commodity multi-modal training data set is used to train a live commodity feature extraction layer of a preset multi-modal data processing model. The multi-modal training data set is labeled according to a preset commodity feature labeling rule. A hybrid expert model is used to generate a plurality of candidate explanation scripts according to the trained multi-modal data processing model, multi-modal production data of a current live commodity, and live room user labels. A compliant explanation script is generated according to the hybrid expert model and the plurality of candidate explanation scripts. The compliant explanation script is updated by using live room barrage corresponding to the compliant explanation script to generate a real-time explanation script. The hybrid expert model is optimized according to the compliant explanation script, live explanation video corresponding to the real-time explanation script, and live effect data. The live explanation efficiency, quality, conversion rate, and violation risk can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, equipment, and computer-readable media for generating product description text in live streaming Technical Field

[0001] This application relates to the field of e-commerce live streaming technology, specifically to a method for generating live product description text, an electronic device, and a computer-readable medium. Background Technology

[0002] Live-stream product explanations should proactively break down product information, address potential concerns, and answer questions. However, current live-stream product explanations primarily rely on manually written scripts and the host's improvisation, which presents four main problems: First, inefficiency, requiring hosts to spend significant time familiarizing themselves with product information and scripts; second, inconsistent quality, influenced by the host's experience and performance, potentially omitting key selling points or violating platform rules; third, lack of personalization, making it difficult to dynamically adjust the focus based on the characteristics of the live-stream audience; and fourth, poor real-time performance, unable to quickly respond to user questions in the chat.

[0003] Although there are some automated script generation tools available, these tools can only fill in information based on fixed templates and lack the ability to deeply integrate and optimize multi-source data in real time, so the improvement in live streaming efficiency and conversion effect is still limited. Summary of the Invention

[0004] This application aims to address one of the technical problems in related technologies to a certain extent. To this end, this application provides a method for generating live-stream product description text, an electronic device, and a computer-readable medium, which has the advantages of improving the efficiency, quality, and conversion rate of live-stream descriptions, reducing the risk of violations, and continuously improving the generation effect of description content for continuous optimization.

[0005] To achieve the above objectives, as the first aspect of this application, the following technical solution is adopted:

[0006] A method for generating product description text for live streaming, comprising:

[0007] A multimodal training dataset of historical live-streamed products is obtained, and the live-streamed product feature extraction layer of a preset multimodal data processing model is trained. The multimodal training dataset is labeled according to preset product feature labeling rules, and the product features include conversion rate, risk level, explanation logic, sales style, core selling points, user pain points, and compliance information.

[0008] Based on the trained multimodal data processing model, the multimodal production data of the current live-streamed products, the user tags in the live-streaming room, and the hybrid expert model, multiple candidate explanation scripts are generated.

[0009] Based on the hybrid expert model and the multiple candidate explanation scripts, a compliant explanation script is generated;

[0010] Collect the live stream comments corresponding to the compliance explanation script, update the compliance explanation script, and generate a real-time explanation script;

[0011] The hybrid expert model is optimized based on the compliant explanation script, the live explanation video corresponding to the real-time explanation script, and the live performance data.

[0012] Optionally, the step of generating multiple candidate explanation scripts based on the trained multimodal data processing model, the multimodal production data of the current live-streamed products, the user tags in the live-streaming room, and the hybrid expert model includes:

[0013] Based on the trained multimodal data processing model, product features are extracted from the multimodal production data of the current live-streamed products;

[0014] By matching the product features with the user tags in the live stream, a personalized explanation strategy is generated;

[0015] Based on the product characteristics, the personalized explanation strategy, and the hybrid expert model, multiple candidate explanation scripts are generated.

[0016] Optionally, the method further includes:

[0017] According to the preset product feature labeling rules, the multimodal production data of the current live-streamed product is labeled;

[0018] The scoring coefficient of the trained multimodal data processing model is determined based on the product features labeled in the multimodal production data and the product features extracted from the multimodal production data by the trained multimodal data processing model.

[0019] If the scoring coefficient does not meet the corresponding target threshold, multimodal training data including the multimodal production data is added to the multimodal training dataset, and the live-streaming product feature extraction layer of the trained multimodal data processing model is trained again.

[0020] Optionally, the scoring coefficients include the accuracy rate of extracting explanation logic, the accuracy rate of extracting core selling points, the matching degree of extracting user pain points, and the accuracy rate of extracting compliance information; the target thresholds for the accuracy rate of extracting explanation logic include 90%, the target thresholds for the accuracy rate of extracting core selling points include 95%, the target thresholds for the matching degree of extracting user pain points include 85%, and the target thresholds for the accuracy rate of extracting compliance information include 100%.

[0021] Optionally, generating a compliant explanation script based on the hybrid expert model and the multiple candidate explanation scripts includes:

[0022] Based on the hybrid expert model, multiple preset selection dimensions, and the weights corresponding to each selection dimension, a target explanation script is determined from the multiple candidate explanation scripts; wherein, the multiple selection dimensions include core selling point coverage, language fluency, user interest matching degree, and compliance risk value;

[0023] The target explanation script is revised based on the hybrid expert model to generate a compliant explanation script.

[0024] Optionally, the hybrid expert model includes a gating network, a content generation sub-model, a content scoring sub-model, and a compliance check sub-model.

[0025] Optionally, optimizing the hybrid expert model based on the compliant explanation script, the live explanation video corresponding to the real-time explanation script, and the live performance data includes:

[0026] Identify the longest segment of product feature explanation in the live-streamed video;

[0027] If the conversion rate of the product feature explanation segment is determined to be greater than the preset conversion rate threshold based on the live broadcast effect data, the parameters of the hybrid expert model are adjusted to increase the weight and time priority of the product feature corresponding to the product feature explanation segment in the candidate explanation script.

[0028] Based on the script segment corresponding to the product feature explanation segment in the real-time explanation script, the hybrid expert model is trained under positive supervision.

[0029] Update the collaboration rules of the hybrid expert model to improve the user interest matching score of the hybrid expert model for candidate explanation scripts that include product features corresponding to the product feature explanation segments;

[0030] If, based on the live explanation video and live effect data corresponding to the real-time explanation script, illegal content is identified in the compliant explanation script, the compliance rule base of the hybrid expert model is updated according to the identified illegal content.

[0031] Optionally, the multimodal training data includes brand introduction, product information, social media feedback, live-stream explanation videos, product tag information, merchant cooperation information, and merchant qualification documents. The merchant qualification documents include product certification certificates, import customs declarations, and quality inspection reports. The merchant qualification documents have been verified by OCR recognition and validity.

[0032] The method for generating live-stream product description text provided in this application embodiment involves obtaining a multimodal training dataset of historical live-stream products and training the live-stream product feature extraction layer of a preset multimodal data processing model. The multimodal training dataset is labeled according to preset product feature labeling rules, and the product features include conversion rate, risk level, explanation logic, sales style, core selling points, user pain points, and compliance information. Multiple candidate explanation scripts are generated based on the trained multimodal data processing model, the current live-stream product's multimodal production data, live-stream user tags, and a hybrid expert model. A compliant explanation script is generated based on the hybrid expert model and the multiple candidate explanation scripts. Live-stream comments corresponding to the compliant explanation script are collected, and the compliant explanation script is updated to generate a real-time explanation script. The hybrid expert model is optimized based on the compliant explanation script, the live-stream explanation video corresponding to the real-time explanation script, and live-stream effect data. By automating data collection and generation, the script preparation time for live streamers is significantly reduced, improving the efficiency of live stream presentations. The integration of multi-source data and historical experience through a hybrid expert model avoids the omission of selling points due to human error, thus improving the quality of live stream presentations. Generating candidate presentation scripts based on live stream user tags improves conversion rates. Generating compliant presentation scripts based on the hybrid expert model and multiple candidate scripts reduces the risk of violations. A closed-loop training mechanism that optimizes the hybrid expert model based on compliant presentation scripts, real-time presentation scripts, corresponding live stream presentation videos, and live stream performance data allows the model to iterate with live stream data, continuously improving the generation of presentation content and achieving continuous optimization.

[0033] As a second aspect of this application, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method for generating live-stream product description text as described in any of the preceding claims.

[0034] As a third aspect of this application, this application also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating live-stream product description text as described in any of the preceding claims.

[0035] These features and advantages of this application will be disclosed in detail in the following specific embodiments and accompanying drawings. The best embodiments or means of this application will be shown in detail in conjunction with the accompanying drawings, but are not intended to limit the technical solutions of this application. In addition, each of these features, elements and components appearing in the following text and drawings is multiple and is labeled with different symbols or numbers for convenience, but all represent parts with the same or similar structure or function. Attached Figure Description

[0036] The following description, in conjunction with the accompanying drawings, further illustrates this application:

[0037] Figure 1 is a flowchart of one embodiment of the method for generating live-stream product description text provided in this application;

[0038] Figure 2 is a flowchart of another implementation of the method for generating live-stream product description text provided in the embodiments of this application;

[0039] Figure 3 is a flowchart of another embodiment of the method for generating live product description text provided in this application;

[0040] Figure 4 is a flowchart of another embodiment of the method for generating live product description text provided in this application;

[0041] Figure 5 is a flowchart of another implementation of the method for generating live-stream product description text provided in the embodiments of this application;

[0042] Figure 6 is a block diagram of one embodiment of the electronic device provided in this application;

[0043] Figure 7 is a schematic diagram of a computer-readable medium provided in an embodiment of this application;

[0044] Explanation of reference numerals in the attached figures

[0045] 101: Processor; 102: Memory

[0046] 103: I / O interface; 104: Bus. Detailed Implementation

[0047] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described are intended to explain this application and should not be construed as limiting it.

[0048] The terms "an embodiment," "example," or "example" used in this specification refer to a particular feature, structure, or characteristic described in connection with the embodiment itself that may be included in at least one embodiment disclosed in this application. The phrase "in an embodiment" appearing in various places throughout the specification does not necessarily refer to the same embodiment.

[0049] Live-stream product explanations should proactively break down product information, address potential concerns, and answer questions. However, current live-stream product explanations primarily rely on manually written scripts and the host's improvisation, which presents four main problems: First, inefficiency, requiring hosts to spend significant time familiarizing themselves with product information and scripts; second, inconsistent quality, influenced by the host's experience and performance, potentially omitting key selling points or violating platform rules; third, lack of personalization, making it difficult to dynamically adjust the focus based on the characteristics of the live-stream audience; and fourth, poor real-time performance, unable to quickly respond to user questions in the chat.

[0050] Although there are some automated script generation tools available, these tools can only fill in information based on fixed templates and lack the ability to deeply integrate and optimize multi-source data in real time, so the improvement in live streaming efficiency and conversion effect is still limited.

[0051] As a first aspect of this application, a method for generating live-stream product description text is provided, as shown in Figure 1. The method includes:

[0052] In step S110, a multimodal training dataset of historical live-streamed products is obtained, and the live-streamed product feature extraction layer of the preset multimodal data processing model is trained; wherein, the multimodal training dataset is labeled according to preset product feature labeling rules, and the product features include conversion rate, risk level, explanation logic, sales style, core selling points, user pain points and compliance information;

[0053] In step S120, multiple candidate explanation scripts are generated based on the trained multimodal data processing model, the multimodal production data of the current live-streamed product, the user tags in the live-streaming room, and the hybrid expert model.

[0054] In step S130, a compliant explanation script is generated based on the hybrid expert model and the multiple candidate explanation scripts;

[0055] In step S140, the live stream comments corresponding to the compliance explanation script are collected, the compliance explanation script is updated, and a real-time explanation script is generated.

[0056] In step S150, the hybrid expert model is optimized based on the compliant explanation script, the live explanation video corresponding to the real-time explanation script, and the live effect data.

[0057] Multimodal training data can have multiple data sources and include various data formats. For example, multimodal training data for imported infant formula can include: video data, such as historical live-stream video clips; text data, such as text from the brand's official website about "a century of dairy history", text from e-commerce platforms about "milk source / docosahexaenoic acid (DHA) content" parameters, and social media posts recommending "babies love it / it doesn't cause heatiness"; image data, such as product packaging images and screenshots from live streams; and portable document format (PDF) data, such as import customs declarations and quality inspection reports.

[0058] The training data (i.e., the multimodal training dataset of historical live-streamed products) can be manually labeled by technical personnel according to preset product feature labeling rules. The goal is for the trained multimodal data processing model to automatically label (i.e., extract features) the production data (i.e., the multimodal production data of the current live-streamed products). It can be understood that labeling according to preset product feature labeling rules refers to extracting and labeling product features based on data, such as conversion rate, risk level, explanation logic, sales style, core selling points, user pain points, and compliance information.

[0059] The conversion rate and risk level are product characteristics annotated in the live-stream video clips. Conversion rate can be categorized as "high-conversion clips" or "low-conversion clips," while risk level can be categorized as "violation clips," "high-risk clips," or "risk-free clips." The explanation logic can include the order in which any of the following are presented: core selling points, user pain points, and compliance information. For example: milk source → quality inspection → preparation method. The speaking style can include the ratio of conversational to technical jargon, the speaking speed, etc. Core selling points and user pain points depend on the specific live-stream product. For example, the core selling point of imported infant formula might be "100% New Zealand milk source," and the user pain point might be "concern about causing heatiness." Compliance information is a product characteristic annotated with qualification documents such as import customs declarations and quality inspection reports. For example, the compliance information for imported infant formula might include "ingredient list location," "certification conclusions," and "milk source identification."

[0060] It is understandable that the preset multimodal data processing model needs to select a multimodal understanding and generation model based on the multimodal data processing requirements of the e-commerce live streaming scenario (which needs to process video, voice, text, images, PDFs, etc. simultaneously), such as the Tongyi Qianwen-Multimodal Collaboration Model (MMC).

[0061] It is understandable that when training a multimodal data processing model, the model can transform and summarize live-streamed video clips from the multimodal training data, converting audio and video information into structured explanatory text. Then, referring to the explanatory text and labeled product features, it learns and extracts product features from other data besides the live-streamed video clips. However, the multimodal production data processed by the trained multimodal data processing model does not include live-streamed video clips, unlike the multimodal training data.

[0062] When training a pre-defined multimodal data processing model using a multimodal training dataset, domain-adaptive fine-tuning can be performed. The underlying general feature extraction layer can be frozen, and only the upper-level live-streaming e-commerce-specific live-streaming product feature extraction layer can be trained, allowing the model to learn the feature extraction logic specific to the live-streaming e-commerce domain. When training the live-streaming product feature extraction layer, specific product feature extraction tasks and their corresponding training data can be set, enabling the model to learn the relevant feature extraction logic in a targeted manner. For example, for the "explanation logic extraction" task, annotated live-streaming explanation video clips (including speech-to-text and on-screen annotations) can be input, allowing the model to learn "high conversion rate; typical logic of the clip: security authentication → functional advantages → promotional activities." For the "core selling point extraction" task, image and text data from the product details page can be input, allowing the model to learn to associate "the map of the milk source in the image → the 'New Zealand milk source' selling point in the text."

[0063] Among them, the Mixture of Experts (MoE) model is an artificial intelligence architecture that improves model efficiency and performance through dynamic division of labor. Its core idea is to decompose complex tasks into multiple sub-tasks, which are handled by different "expert" sub-networks, and then the most suitable combination of experts is dynamically selected for computation through a "gating network".

[0064] The generated compliance explanation scripts are applied to the live-streaming presentations. During the actual presentation, live-streaming chat comments are collected, and the multimodal production data of the current live-streaming product is matched with the content of user questions in the chat comments to update the compliance explanation script in real time, generating a real-time explanation script. For example, in response to a chat comment asking "Does it contain sucrose?", the script generates an answer in real time: "It does not contain sucrose; the ingredient list is on page 3 of the quality inspection report," adding it to the compliance explanation script to form a real-time explanation script. Similarly, in response to a chat comment asking "Is it waterproof?", the script retrieves the product's waterproof parameters in real time and generates corresponding wording, adding it to the compliance explanation script to form a real-time explanation script.

[0065] It's understandable that, besides updating the compliance explanation script based on live stream comments, content operators and broadcasters can also refine the script, adjusting the language style to suit the broadcaster's personal characteristics. In other words, the real-time explanation script can be an updated and refined version of the compliance explanation script. Furthermore, the refined script and modification records can be stored in a database as incremental data for subsequent model training / optimization.

[0066] The method for generating live-stream product description text provided in this application embodiment involves obtaining a multimodal training dataset of historical live-stream products and training the live-stream product feature extraction layer of a preset multimodal data processing model. The multimodal training dataset is labeled according to preset product feature labeling rules, and the product features include conversion rate, risk level, explanation logic, sales style, core selling points, user pain points, and compliance information. Multiple candidate explanation scripts are generated based on the trained multimodal data processing model, the current live-stream product's multimodal production data, live-stream user tags, and a hybrid expert model. A compliant explanation script is generated based on the hybrid expert model and the multiple candidate explanation scripts. Live-stream comments corresponding to the compliant explanation script are collected, and the compliant explanation script is updated to generate a real-time explanation script. The hybrid expert model is optimized based on the compliant explanation script, the live-stream explanation video corresponding to the real-time explanation script, and live-stream effect data. By automating data collection and generation, the script preparation time for live streamers is significantly reduced, improving the efficiency of live stream presentations. The integration of multi-source data and historical experience through a hybrid expert model avoids the omission of selling points due to human error, thus improving the quality of live stream presentations. Generating candidate presentation scripts based on live stream user tags improves conversion rates. Generating compliant presentation scripts based on the hybrid expert model and multiple candidate scripts reduces the risk of violations. A closed-loop training mechanism that optimizes the hybrid expert model based on compliant presentation scripts, real-time presentation scripts, corresponding live stream presentation videos, and live stream performance data allows the model to iterate with live stream data, continuously improving the generation of presentation content and achieving continuous optimization.

[0067] Furthermore, the applicant of this application also proposes that the conversion rate can be determined by analyzing live-streamed video clips and their corresponding live-streaming performance data. The following example uses pseudocode:

[0068] / / 1. Define threshold parameters (can be adjusted according to actual business needs)

[0069] Set add-to-cart conversion rate threshold = Y%

[0070] Set click-through rate threshold = Z%

[0071] Set the online user threshold = W / / Exclude invalid data due to insufficient online users;

[0072] / / 2. Input Data Acquisition

[0073] Live broadcast video clips = {

[0074] Lecture Time Segment: [Start Time, End Time]

[0075] Product Information: {Product ID, Product Name, Product Category}

[0076] Time period data: {

[0077] Online users: Real-time online user list.

[0078] Click count: The number of product clicks during this period.

[0079] Add-to-cart count: The number of items added to the shopping cart during this period.

[0080] Transaction volume: The number of product orders completed during this period.

[0081] }

[0082] };

[0083] / / 3. Data Preprocessing

[0084] Calculate the average number of online users = the average number of online users during the lecture period.

[0085] Calculate the add-to-cart conversion rate = Number of transactions / Number of add-to-cart transactions

[0086] Calculate the click-through rate (CTR) = Number of transactions / Number of clicks;

[0087] / / 4. Logic for Judging the Degree of Conversion

[0088] If the average number of online users is greater than or equal to the online user base threshold: / / Ensure the data sample is valid

[0089] / / Key criterion: Both add-to-cart conversion rate and click-through rate are higher than the set threshold.

[0090] If (add-to-cart conversion rate > add-to-cart conversion rate threshold and click-to-buy rate > click-through conversion rate threshold):

[0091] This live stream video clip has been marked as a "high-conversion clip".

[0092] otherwise:

[0093] This live stream video clip has been marked as a "low-conversion clip".

[0094] otherwise:

[0095] The output is "Invalid data (insufficient number of online users), unable to determine".

[0096] Furthermore, the applicant of this application proposes that the level of risk can be determined by matching compliance rule bases and calling hybrid expert models. The following pseudocode provides an example:

[0097] / / 1. Initialize parameters and resources

[0098] Loading platform prohibited word library = [prohibited word 1, prohibited word 2, ..., prohibited word N] / / Contains sensitive words, illegal advertising words, etc.

[0099] Set the large model call parameters as follows: {Confidence threshold: 0.8, Analysis prompt: "Determine whether the following text actually contains prohibited words on the platform. If it is a false positive, please explain the reason"};

[0100] / / 2. Extract the explanatory text from the live video clips.

[0101] Explanation text = Text extracted from the host's speech in the live video clips / / Assuming speech-to-text processing has already been completed;

[0102] / / 3. Preliminary detection of prohibited words

[0103] Initialize the list of suspected prohibited words = []

[0104] Iterate through and explain each word / phrase in the text:

[0105] If the word / phrase exists in the platform's prohibited word database:

[0106] Add the words / phrases to the list of suspected prohibited words;

[0107] / / 4. Initial risk labeling (high-risk prediction)

[0108] If the list of suspected prohibited words is not empty:

[0109] The automatically labeled live-streamed video clips were classified as "high-risk clips".

[0110] else:

[0111] Output: "Risk-free segment"

[0112] The process is complete;

[0113] / / 5. Secondary validation of the large model (to determine if there are any misjudgments)

[0114] / / Analyze the large model to determine if suspected prohibited words are actually violations.

[0115] Large model input = {

[0116] Text to be analyzed: Explanation text,

[0117] Suspected prohibited words: List of suspected prohibited words

[0118] Prompt message: Large model call parameters. Analyze prompt words.

[0119] }

[0120] Large model output = calling the large model interface (large model input);

[0121] / / 6. Final violation determination

[0122] Analysis of large model output results:

[0123] If the large model determines that "a real prohibited word exists" and the confidence level is greater than or equal to the large model's parameters, then the confidence threshold is used.

[0124] The final video clip used for the live stream explanation was labeled as a "violation segment".

[0125] Record violation details = {Banned words: List of suspected prohibited words, Judgment basis: Output results of the large model. Reason}

[0126] else if the large model determines it as a "false positive" (e.g., homophones or words in normal context are mismatched):

[0127] The "high-risk segment" label has been removed and replaced with "risk-free segment".

[0128] Record the reason for misjudgment = Large model output result. Reason (used to optimize the prohibited word database)

[0129] else:

[0130] / / When the confidence level of the large model is insufficient, retain the high-risk labels and prompt for manual review.

[0131] Retain the "high-risk segments" label and output "requires manual review";

[0132] / / 7. Output the final result

[0133] Return to the live stream video clip ID, final risk level label, and details of violation / misjudgment.

[0134] In some embodiments, generating multiple candidate explanation scripts (i.e., those involved in step 120) based on the trained multimodal data processing model, the multimodal production data of the current live-streamed product, the user tags in the live-streaming room, and the hybrid expert model, as shown in Figure 2, may include:

[0135] In step S210, product features are extracted from the multimodal production data of the current live-streamed products based on the trained multimodal data processing model.

[0136] In step S220, a personalized explanation strategy is generated by matching the product features with the user tags in the live broadcast room.

[0137] In step S230, multiple candidate explanation scripts are generated based on the product characteristics, the personalized explanation strategy, and the hybrid expert model.

[0138] In this embodiment, the user tags for the live stream are not specifically limited. Different live stream products can be categorized using different criteria, such as age, region, and consumption preferences. It is understood that product characteristics and live stream user tags are semantically matched using historical live stream performance data. For example, if historical live stream performance data shows that younger audiences prioritize trendy features and mothers prioritize safety certifications, then when the live stream user tag is "25-35 year old mothers," the personalized explanation strategy might be "emphasizing safety certifications and taste."

[0139] In this embodiment of the application, the number of candidate explanation scripts is not specifically limited. For example, it may include 3-5.

[0140] Furthermore, the applicant of this application proposes that the pre-defined multimodal data processing model can be subjected to dual supervised training through indicator evaluation and business feedback to ensure that the model's processing results meet actual needs. Indicator evaluation refers to determining the scoring coefficient of the multimodal data processing model during the training process based on the multimodal training dataset. Training only ends when the scoring coefficient meets the corresponding target threshold. Business feedback refers to determining the scoring coefficient of the multimodal data processing model during actual inference after the model has been trained. If the scoring coefficient does not meet the corresponding target threshold, the live-streaming product feature extraction layer of the trained multimodal data processing model is retrained.

[0141] Accordingly, in some embodiments, as shown in FIG3, the method further includes:

[0142] In step S310, the multimodal production data of the current live-streamed product is labeled according to the preset product feature labeling rules;

[0143] In step S320, the scoring coefficient of the trained multimodal data processing model is determined based on the product features labeled in the multimodal production data and the product features extracted from the multimodal production data by the trained multimodal data processing model.

[0144] In step S330, if the scoring coefficient does not meet the corresponding target threshold, multimodal training data including the multimodal production data is added to the multimodal training dataset, and the live commodity feature extraction layer of the trained multimodal data processing model is trained again.

[0145] In some embodiments, the scoring coefficients include the accuracy rate of extracting explanation logic, the accuracy rate of extracting core selling points, the matching degree of extracting user pain points, and the accuracy rate of extracting compliance information; the target threshold corresponding to the accuracy rate of extracting explanation logic includes 90%, the target threshold corresponding to the accuracy rate of extracting core selling points includes 95%, the target threshold corresponding to the matching degree of extracting user pain points includes 85%, and the target threshold corresponding to the accuracy rate of extracting compliance information includes 100%.

[0146] Among them, the accuracy rate of the explanation logic extraction reflects the degree of overlap between the explanation logic extracted by the model and the explanation logic annotated by the human; the accuracy rate of the core selling point extraction reflects the scope of the core selling points extracted by the model that cover the core selling points annotated by the human (such as "DHA content" and "sugar-free"); the matching degree of the user pain point extraction reflects the degree of overlap between the user pain points extracted by the model and the user pain points annotated by the human (such as "whether it causes internal heat"); and the accuracy rate of the compliance information extraction reflects whether the compliance information extracted by the model (such as "import customs declaration number" and "quality inspection conclusion") is consistent with the human annotation on the original document.

[0147] For example, if the "accuracy rate of extracting explanation logic" drops to 85%, 1,000 newly labeled high-conversion video clips are added and the model is fine-tuned. If errors occur in the extraction of compliance information, the training data of the PDF parsing module in the multimodal data processing model is strengthened (such as adding fuzzy scanned documents and multilingual customs declaration samples).

[0148] In addition, the applicant of this application also proposed that the business feedback of the multimodal data processing model is not limited to this. If the model extracts a core selling point, and the core selling point also exists in the compliant explanation script, but the anchor does not explain it in the actual live broadcast, resulting in the barrage frequently asking about the selling point, then the selling point can also be supplemented and marked as "missed detection" in the multimodal production data and added to the next round of multimodal training dataset.

[0149] In some embodiments, generating a compliant explanation script (i.e., the step involved in step 130) based on the hybrid expert model and the multiple candidate explanation scripts, as shown in Figure 4, may include:

[0150] In step S410, the target explanation script is determined from the multiple candidate explanation scripts based on the hybrid expert model, multiple preset selection dimensions and the weights corresponding to each selection dimension; wherein, the multiple selection dimensions include core selling point coverage, language fluency, user interest matching degree and compliance risk value.

[0151] In step S420, the target explanation script is modified according to the hybrid expert model to generate a compliant explanation script.

[0152] In this process, multiple candidate explanation scripts are scored based on a hybrid expert model, multiple preset selection dimensions, and the weights corresponding to each selection dimension. The script with the highest score is then selected as the target explanation script. It is understood that this embodiment is not limited to this; selection can also be based on only some selection dimensions. For example, a script with a core selling point coverage rate exceeding 90% and a compliance risk value of 0 can be selected as the target explanation script.

[0153] In this embodiment of the application, the weights corresponding to each selected dimension are not specifically limited. For example, the weights corresponding to the core selling point coverage, language fluency, user interest matching degree and compliance risk value can be 30%, 20%, 30% and 20%, respectively.

[0154] First, the hybrid expert model can access a compliance rule base, which is built and dynamically updated based on laws and regulations (such as prohibited words in the Advertising Law) and e-commerce platform rules, covering clauses on absolute terms, false advertising, and category-specific restrictions. Then, the hybrid expert model performs multi-dimensional scanning and verification of the target explanation script based on the compliance rule base: keyword detection, such as identifying prohibited words (e.g., "top" or "best") and sensitive expressions (e.g., unfounded efficacy claims); logical verification, such as checking for ambiguity in expressions (e.g., "suitable for all people"); and qualification matching, such as verifying whether claims like "imported" or "organic" are supported by corresponding qualification documents (e.g., customs declarations, test reports). Finally, the hybrid expert model automatically corrects the target explanation script: replacement, such as replacing prohibited words with compliant expressions (e.g., "best" → "high-quality"); and adjustment, such as adding limiting conditions to eliminate ambiguity (e.g., "suitable for all babies" → "suitable for babies aged 3-12 months, see instructions"). Supplementation includes, for example, inserting key information from documents (such as "see page 2 of the quality inspection report") for content requiring supporting qualifications; and deletion includes, for example, removing uncorrectable violations (such as unsubstantiated claims of "therapeutic efficacy"). Furthermore, the hybrid expert model performs a secondary verification of the target explanation script to ensure zero compliance risk, outputting a compliant explanation script and correction records for content operations and broadcaster refinement, as well as model training / optimization.

[0155] In some embodiments, the hybrid expert model includes a gating network, a content generation sub-model, a content scoring sub-model, and a compliance check sub-model.

[0156] In this context, it can be understood that the gating network is used to dynamically select the most suitable combination of experts from the content generation sub-model, the content scoring sub-model, and the compliance check sub-model. The content generation sub-model is used to generate multiple candidate explanation scripts, the content scoring sub-model is used to score the multiple candidate explanation scripts respectively to determine the target explanation script, and the compliance check sub-model is used to revise the target explanation script to generate a compliant explanation script.

[0157] In some embodiments, optimizing the hybrid expert model based on the compliant explanation script, the live explanation video corresponding to the real-time explanation script, and the live effect data (i.e., the step involved in step 150), as shown in Figure 5, may include:

[0158] In step S510, the longest segment of product feature explanation in the live-stream explanation video is identified;

[0159] In step S520, if the conversion rate of the product feature explanation segment is determined to be greater than the preset conversion rate threshold based on the live broadcast effect data, the parameters of the hybrid expert model are adjusted to increase the weight and time priority of the product feature corresponding to the product feature explanation segment in the candidate explanation script.

[0160] In step S530, the hybrid expert model is subjected to positive supervised training based on the script segment corresponding to the product feature explanation segment in the real-time explanation script.

[0161] In step S540, the collaborative rules of the hybrid expert model are updated to improve the user interest matching score of the hybrid expert model for candidate explanation scripts that include product features corresponding to the product feature explanation segments.

[0162] In step S550, if it is found that there is illegal content in the compliant explanation script based on the live explanation video and live effect data corresponding to the real-time explanation script, the compliance rule base of the hybrid expert model is updated according to the identified illegal content.

[0163] The live streaming performance data includes performance metrics corresponding to different live streaming video segments, such as user dwell time (a relative value compared to other segments) and user behavior data during that time period (such as no-exit rate and number of questions asked in the chat). Based on the live streaming performance data, the conversion rate of the identified product feature explanation segments can be calculated.

[0164] If the conversion rate is greater than the preset conversion rate threshold, it can be considered that the identified product feature explanation segment has a high conversion rate. Therefore, the weight and time priority of the product features involved in the segment in the candidate explanation script generated by the model can be adjusted. For example, if "drinking method" is identified as a product feature involved in a segment with a high conversion rate, then the weight of the product features related to "drinking method" in the candidate explanation script can be increased (e.g., from the original basic weight of 15% to 25%-30%), the explanation time of the product features related to "drinking method" in the candidate explanation script can be increased (e.g., from the original proportion of 10% to 15%-20%), and "drinking method" can be arranged in the prime time of the live broadcast in the candidate explanation script (e.g., 10-20 minutes after the start of the broadcast, based on the high traffic time determined by historical data).

[0165] If the conversion rate exceeds a preset conversion rate threshold, it can be considered that the identified product feature explanation segments used high-quality language. Therefore, the hybrid expert model can be positively supervised and trained based on the corresponding script segments in the real-time explanation script, optimizing the language style of the candidate explanation scripts generated by the model. The script segments corresponding to the identified product feature explanation segments in the real-time explanation script include high-quality features such as the host's actual explanation style and explanation logic. These can serve as reference samples for the content generation sub-model within the hybrid expert model. The content generation sub-model will subsequently use similar language styles (such as step-by-step descriptions and interactive prompts) and explanation logic (such as milk source → quality inspection → preparation method) when generating candidate explanation scripts.

[0166] At the same time, the collaborative rules of the hybrid expert model (i.e., the gating network parameters in the hybrid expert model) can be updated so that the content scoring sub-model in the hybrid expert model can give higher "user interest matching" scores to candidate explanation scripts that include product features involved in the explanation segments of products with high conversion rates when scoring multiple candidate explanation scripts in the future.

[0167] At the same time, based on the live broadcast video and live broadcast effect data corresponding to the real-time explanation script, it can identify whether there are any violations in the compliance explanation script that were not checked and corrected by the compliance inspection sub-model in the previous hybrid expert model. This allows the compliance rule base of the hybrid expert model to be updated, so that the compliance inspection sub-model of the hybrid expert model will have the ability to check and correct the same violations in the future.

[0168] In some embodiments, the multimodal training data includes brand introduction, product information, social media feedback, live-stream explanation videos, product tag information, merchant cooperation information, and merchant qualification documents. The merchant qualification documents include product certification certificates, import customs declarations, and quality inspection reports. The merchant qualification documents have been verified by OCR recognition and validity.

[0169] Among these, brand story refers to information such as brand history, philosophy, and honors from the brand's official website; product information refers to data such as product parameters, prices, and sales volume obtained through the open interfaces of e-commerce platforms; social media feedback refers to user reviews and recommendations on social media platforms; live-streaming demonstration videos refer to demonstration videos from the product's own live-streaming room, external well-known live-streaming rooms, and brand live-streaming rooms; product label information refers to product labels such as ingredient lists and nutritional information; and merchant cooperation information refers to the list of cooperative SKUs, sales mechanisms (such as discounts, spending thresholds, and buy-two-get-one-free offers), and inventory quantities. For merchant qualification documents, after collection, OCR recognition and validity verification are performed manually or using algorithms.

[0170] It is understandable that multimodal training data is not limited to brand introductions, product information, social media feedback, live streaming videos, product tag information, merchant cooperation information, and merchant qualification documents. Any information that can reflect the product characteristics of the live streaming products can be included.

[0171] As a second aspect of this application, an electronic device is provided, as shown in FIG6, the electronic device comprising:

[0172] One or more processors 101;

[0173] The memory 102 stores one or more computer programs, which, when executed by the one or more processors 101, cause the one or more processors 101 to implement the method for generating live-stream product description text provided in the first aspect of the embodiments of this application.

[0174] The electronic device may also include one or more I / O interfaces 103 connected between the processor 101 and the memory 102, configured to enable information interaction between the processor 101 and the memory 102.

[0175] The processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) is connected between the processor and the memory, enabling information exchange between the processor and the memory, including but not limited to a data bus (Bus).

[0176] In some embodiments, the processor 101, memory 102, and I / O interface 103 are interconnected via bus 104, and thus connected to other components of the computing device.

[0177] As a third aspect of this application, as shown in FIG7, a computer-readable medium is provided having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for generating live-stream product description text provided in the first aspect of this application.

[0178] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. Accordingly, the computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can implement the methods of any of the above embodiments. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0179] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Those skilled in the art should understand that this application includes, but is not limited to, the contents described in the accompanying drawings and the specific embodiments above. Any modifications that do not depart from the functional and structural principles of this application will be included within the scope of the claims.

Claims

1. A method for generating live-stream product description text, characterized in that, include: A multimodal training dataset of historical live-streamed products is acquired, and the live-streamed product feature extraction layer of a pre-defined multimodal data processing model is trained. The multimodal training dataset is labeled according to pre-defined product feature labeling rules, and the product features include conversion rate, risk level, explanation logic, sales style, core selling points, user pain points, and compliance information. Multiple candidate explanation scripts are generated based on the trained multimodal data processing model, the current live-streamed product's multimodal production data, live-stream user tags, and a hybrid expert model. A compliant explanation script is generated based on the hybrid expert model and the multiple candidate explanation scripts. Live-stream comments corresponding to the compliant explanation script are collected, and the compliant explanation script is updated to generate a real-time explanation script. The hybrid expert model is optimized based on the compliant explanation script, the live-stream explanation video corresponding to the real-time explanation script, and live-stream effect data.

2. The method according to claim 1, characterized in that, The step of generating multiple candidate explanation scripts based on a trained multimodal data processing model, multimodal production data of the current live-streamed product, live-streaming user tags, and a hybrid expert model includes: extracting product features from the multimodal production data of the current live-streamed product using the trained multimodal data processing model; generating a personalized explanation strategy by matching the product features with live-streaming user tags; and generating multiple candidate explanation scripts based on the product features, the personalized explanation strategy, and the hybrid expert model.

3. The method according to claim 2, characterized in that, The method further includes: labeling the multimodal production data of the current live-streamed product according to the preset product feature labeling rules; determining the scoring coefficient of the trained multimodal data processing model based on the product features labeled in the multimodal production data and the product features extracted from the multimodal production data by the trained multimodal data processing model; and, if the scoring coefficient does not meet the corresponding target threshold, supplementing the multimodal training dataset with multimodal training data including the multimodal production data, and retraining the live-streamed product feature extraction layer of the trained multimodal data processing model.

4. The method according to claim 3, characterized in that, The scoring coefficients include the accuracy rate of extracting explanation logic, the accuracy rate of extracting core selling points, the matching degree of extracting user pain points, and the accuracy rate of extracting compliance information. The target thresholds for the accuracy rate of extracting explanation logic include 90%, the target thresholds for the accuracy rate of extracting core selling points include 95%, the target thresholds for the matching degree of extracting user pain points include 85%, and the target thresholds for the accuracy rate of extracting compliance information include 100%.

5. The method according to claim 1, characterized in that, The step of generating a compliant explanation script based on the hybrid expert model and the multiple candidate explanation scripts includes: determining a target explanation script from the multiple candidate explanation scripts based on the hybrid expert model, multiple preset selection dimensions and the weights corresponding to each selection dimension; wherein, the multiple selection dimensions include core selling point coverage, language fluency, user interest matching degree and compliance risk value; and modifying the target explanation script based on the hybrid expert model to generate a compliant explanation script.

6. The method according to claim 5, characterized in that, The hybrid expert model includes a gating network, a content generation sub-model, a content scoring sub-model, and a compliance inspection sub-model.

7. The method according to claim 1, characterized in that, The optimization of the hybrid expert model based on the compliant explanation script, the live explanation video corresponding to the real-time explanation script, and the live performance data includes: identifying the longest-running product feature explanation segment in the live explanation video; adjusting the parameters of the hybrid expert model when the conversion rate of the product feature explanation segment is greater than a preset conversion rate threshold, based on the live performance data, to increase the weight and time priority of the product feature corresponding to the product feature explanation segment in the candidate explanation script; performing positive supervised training on the hybrid expert model based on the script segment corresponding to the product feature explanation segment in the real-time explanation script; updating the collaboration rules of the hybrid expert model to improve the user interest matching score of the hybrid expert model for candidate explanation scripts containing the product feature corresponding to the product feature explanation segment; and updating the compliance rule library of the hybrid expert model based on the identified violations when violations are identified in the compliant explanation script based on the live explanation video and live performance data corresponding to the real-time explanation script.

8. The method according to any one of claims 1-7, characterized in that, The multimodal training data includes brand introductions, product information, social media feedback, live-streamed explanation videos, product tag information, merchant cooperation information, and merchant qualification documents. The merchant qualification documents include product certification certificates, import customs declarations, and quality inspection reports. The merchant qualification documents have been verified by OCR recognition and validity.

9. An electronic device, characterized in that, include: One or more processors; A memory having stored one or more computer programs thereon, which, when executed by one or more processors, cause the one or more processors to implement the method for generating live-stream product description text according to any one of claims 1 to 8.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for generating live-stream product description text as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Content generation method and device, readable medium, electronic equipment and program product

    CN118354107A

  • Commodity selling point extraction and generation method and system based on RAG and multi-modal model

    CN119940304A