Generation method and device of live broadcast commodity explanation text, and computer readable medium
By generating compliant explanation scripts through multimodal data processing models and hybrid expert models, the problems of low efficiency, unstable quality, lack of personalization, and poor real-time performance in live product explanations have been solved, achieving efficient, personalized, and real-time live explanation effects.
Patent Information
- Application Number
- CN202511695266.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-10-23
- Filing Date
- 2025-11-19
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2045-11-19
AI Technical Summary
Current live-streaming product descriptions rely on manually written scripts, which are inefficient, of inconsistent quality, lack personalization, and have poor real-time performance. Furthermore, automated script generation tools lack the ability to integrate multi-source data and perform real-time optimization, resulting in limited improvements in live-streaming efficiency and conversion rates.
Employing a multimodal data processing model and a hybrid expert model, the system annotates product features using a training dataset, generates compliant explanation scripts, and updates them in real time to adapt to live stream comments, thereby optimizing the model to improve explanation efficiency and quality.
It significantly reduces the script preparation time for broadcasters, improves the efficiency and quality of live broadcast explanations, enhances personalization and real-time performance, reduces the risk of violations, and achieves continuously optimized explanation content generation.
Smart Images

Figure CN121174019A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of e-commerce live broadcast, in particular to a live broadcast commodity explanation text generation method, an electronic device and a computer readable medium. BACKGROUND
[0002] The live broadcast commodity explanation content can actively disassemble commodity information, actively eliminate potential concerns of the audience, answer questions of the audience, etc. The current presentation of the live broadcast commodity explanation content mainly relies on manual script writing and improvisation of the anchor on the script, which has the following four problems: first, low efficiency, the anchor needs to spend a lot of time familiarizing with the commodity information and the script; second, unstable quality, affected by the experience and state of the anchor, the explanation content may miss the core selling point or violate the platform rules; third, insufficient personalization, it is difficult to dynamically adjust the explanation focus according to the fan characteristics of the live broadcast room; fourth, poor real-time performance, unable to quickly respond to user questions in the bullet screen.
[0003] Although there are some automatic script generation tools at present, these automatic script generation tools can only fill information based on fixed templates, lack depth fusion and real-time optimization capabilities for multi-source data, and the live broadcast efficiency and conversion effect are still limited. SUMMARY
[0004] The present application aims to solve one of the technical problems in the related art to some extent. To this end, the present application provides a live broadcast commodity explanation text generation method, an electronic device and a computer readable medium, which has the advantages of improving live broadcast explanation efficiency, live broadcast explanation quality and conversion rate, reducing the risk of violation, continuously improving the generation effect of explanation content, and achieving continuous optimization.
[0005] In order to achieve the above purpose, as a first aspect of the present application, the present application adopts the following technical solution: A live broadcast commodity explanation text generation method, comprising: Obtaining a multi-modal training data set of historical live broadcast commodities, and training a live broadcast commodity feature extraction layer of a pre-set multi-modal data processing model; wherein the multi-modal training data set is labeled according to a pre-set commodity feature labeling rule, and the commodity features include conversion degree, risk degree, explanation logic, rhetoric style, core selling point, user pain point and compliance information; Generating a plurality of candidate explanation scripts according to the trained multi-modal data processing model, multi-modal production data of a current live broadcast commodity, live broadcast room user labels and a hybrid expert model; Generating a compliant explanation script according to the hybrid expert model and the plurality of candidate explanation scripts; Collecting the bullet screen of the live broadcast room corresponding to the compliant explanation script, updating the compliant explanation script, and generating a real-time explanation script; According to the compliance explanation script, the live explanation video and live effect data corresponding to the real-time explanation script, the hybrid expert model is optimized.
[0006] Optionally, the generating a plurality of candidate explanation scripts according to the trained multi-modal data processing model, the multi-modal production data of the current live commodity, the user tags in the live room and the hybrid expert model comprises: According to the trained multi-modal data processing model, the product features are extracted from the multi-modal production data of the current live commodity; By matching the product features with the user tags in the live room, an individualized explanation strategy is generated; According to the product features, the individualized explanation strategy and the hybrid expert model, a plurality of candidate explanation scripts are generated.
[0007] Optionally, the method further comprises: According to the preset product feature annotation rule, the multi-modal production data of the current live commodity is annotated; According to the product features annotated on the multi-modal production data and the product features extracted from the multi-modal production data by the trained multi-modal data processing model, a scoring coefficient of the trained multi-modal data processing model is determined; In the case where the scoring coefficient does not satisfy a corresponding target threshold, multi-modal training data including the multi-modal production data are supplemented to the multi-modal training data set, and the live product feature extraction layer of the trained multi-modal data processing model is trained again.
[0008] Optionally, the scoring coefficient comprises an explanation logic extraction accuracy, a core selling point extraction accuracy, a user pain point extraction matching degree and a compliance information extraction accuracy; the target threshold corresponding to the explanation logic extraction accuracy comprises 90%, the target threshold corresponding to the core selling point extraction accuracy comprises 95%, the target threshold corresponding to the user pain point extraction matching degree comprises 85%, and the target threshold corresponding to the compliance information extraction accuracy comprises 100%.
[0009] Optionally, the generating a compliance explanation script according to the hybrid expert model and the plurality of candidate explanation scripts comprises: According to the hybrid expert model, a plurality of selection dimensions and weights corresponding to each selection dimension, a target explanation script is determined from the plurality of candidate explanation scripts; wherein the plurality of selection dimensions comprise a core selling point coverage, a language fluency, a user interest matching degree and a compliance risk value; According to the hybrid expert model, the target explanation script is corrected to generate a compliance explanation script.
[0010] Optionally, the hybrid expert model comprises a gating network, a content generation submodel, a content scoring submodel, and a compliance checking submodel.
[0011] Optionally, the optimization of the hybrid expert model according to the compliance explanation script, the live explanation video corresponding to the real-time explanation script, and the live effect data comprises: identifying a product feature explanation segment with the longest time in the live explanation video; in a case where it is determined according to the live effect data that a conversion rate of the product feature explanation segment is greater than a preset conversion rate threshold, adjusting parameters of the hybrid expert model to improve a weight and a time period priority of a product feature corresponding to the product feature explanation segment in a candidate explanation script; performing forward supervised training on the hybrid expert model according to a script segment corresponding to the product feature explanation segment in the real-time explanation script; updating a collaborative rule of the hybrid expert model to improve a user interest matching degree score of the hybrid expert model for a candidate explanation script comprising a product feature corresponding to the product feature explanation segment; in a case where it is identified according to the live explanation video corresponding to the real-time explanation script and the live effect data that there is illegal content in the compliance explanation script, updating a compliance rule base of the hybrid expert model according to the identified illegal content.
[0012] Optionally, the multi-modal training data comprises brand introduction, product information, social media feedback, live explanation video, product label information, merchant cooperation information, and merchant qualification documents, the merchant qualification documents comprising product certification certificates, import customs declaration forms, and quality inspection reports, and the merchant qualification documents have been identified by OCR and verified for validity.
[0013] The method for generating live commodity explanation text provided in the embodiments of the present application comprises the following steps: obtaining a multi-modal training data set of historical live commodities, and training a live commodity feature extraction layer of a preset multi-modal data processing model; wherein the multi-modal training data set is labeled according to a preset commodity feature labeling rule, and the commodity features include conversion degree, risk degree, explanation logic, sales pitch style, core selling point, user pain point and compliance information; generating a plurality of candidate explanation scripts according to the trained multi-modal data processing model, multi-modal production data of a current live commodity, live room user labels and a hybrid expert model; generating a compliant explanation script according to the hybrid expert model and the plurality of candidate explanation scripts; collecting live room bullet screen corresponding to the compliant explanation script, updating the compliant explanation script to generate a real-time explanation script; and optimizing the hybrid expert model according to the compliant explanation script, live explanation video corresponding to the real-time explanation script and live effect data. Through automatic collection and generation, the script preparation time of the host is significantly reduced, and the live explanation efficiency can be improved; through the fusion of multi-source data and historical experience by the hybrid expert model, the omission of selling points caused by human negligence can be avoided, and the live explanation quality can be improved; through the generation of candidate explanation scripts based on live room user labels, the conversion rate can be improved; through the generation of compliant explanation scripts according to the hybrid expert model and the plurality of candidate explanation scripts, the risk of violation can be reduced; and through the closed-loop training mechanism of optimizing the hybrid expert model according to the compliant explanation script, live explanation video corresponding to the real-time explanation script and live effect data, the model is iterated with live data, and the explanation content generation effect is continuously improved, achieving continuous optimization.
[0014] As a second aspect of the present application, the present application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the live commodity explanation text generation method according to any one of the preceding aspects when executing the computer program.
[0015] As a third aspect of the present application, the present application also provides a computer readable medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the live commodity explanation text generation method according to any one of the preceding aspects.
[0016] These features and advantages of the present application will be disclosed in detail in the following specific embodiments and drawings. The best implementation or means of the present application will be fully described in conjunction with the drawings, but it is not a limitation on the technical solutions of the present application. In addition, these features, elements and components appearing in each of the following text and drawings are multiple, and different symbols or numbers are marked for the convenience of representation, but all represent the same or similar structure or function parts. BRIEF DESCRIPTION OF DRAWINGS
[0017] The present application will be further described below with reference to the drawings. Figure 1 A flow chart of one embodiment of the method for generating live commodity explanation text provided by the embodiment of the present application; Figure 2 A flow chart of another embodiment of the method for generating live commodity explanation text provided by the embodiment of the present application; Figure 3 A flow chart of another embodiment of the method for generating live commodity explanation text provided by the embodiment of the present application; Figure 4 A flow chart of another embodiment of the method for generating live commodity explanation text provided by the embodiment of the present application; Figure 5 A flow chart of another embodiment of the method for generating live commodity explanation text provided by the embodiment of the present application; Figure 6 A module diagram of one embodiment of the electronic device provided by the embodiment of the present application; Figure 7 A schematic diagram of the computer readable medium provided by the embodiment of the present application; Explanation of reference signs 101: processor 102: memory 103: I / O interface 104: bus. DETAILED DESCRIPTION
[0018] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, in which the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. Based on the embodiments, it is intended to explain the present application, and cannot be understood as a limitation of the present application.
[0019] In the present specification, "one embodiment" or "an example" or "an example" means that a specific feature, structure or characteristic described in connection with the embodiment itself can be included in at least one embodiment of the present disclosure. The appearance of the phrase "in one embodiment" at various places in the specification does not necessarily refer to the same embodiment.
[0020] The live commodity explanation content can actively disassemble commodity information, actively eliminate potential concerns of the audience, answer questions of the audience, and the like. Currently, the presentation of the live commodity explanation content mainly relies on artificial script writing and improvisation of the anchor according to the script, and there are the following four problems: 1) low efficiency, the anchor needs to spend a lot of time familiarizing with the commodity information and the script; 2) unstable quality, affected by the experience and state of the anchor, the explanation content may miss the core selling point or violate the platform rules; 3) insufficient personalization, it is difficult to dynamically adjust the explanation focus according to the fan characteristics of the live room; and 4) poor real-time performance, unable to quickly respond to user questions in the bullet screen.
[0021] Although there are some automatic script generation tools at present, these automatic script generation tools can only fill information based on fixed templates, lack of deep fusion and real-time optimization capability of multi-source data, and the live efficiency and conversion effect are still limited.
[0022] As a first aspect of the present application, a method for generating a live commodity explanation text is provided, as shown in Figure 1 The method comprises: In step S110, a multi-modal training data set of historical live commodities is obtained, and a live commodity feature extraction layer of a pre-set multi-modal data processing model is trained; wherein the multi-modal training data set is labeled according to a pre-set commodity feature labeling rule, and the commodity features include conversion degree, risk degree, explanation logic, speaking style, core selling point, user pain point and compliance information; In step S120, a plurality of candidate explanation scripts are generated according to the trained multi-modal data processing model, multi-modal production data of a current live commodity, live room user labels and a hybrid expert model; In step S130, a compliant explanation script is generated according to the hybrid expert model and the plurality of candidate explanation scripts; In step S140, the compliant explanation script corresponding to the live room bullet screen is collected, the compliant explanation script is updated, and a real-time explanation script is generated; In step S150, the hybrid expert model is optimized according to the compliant explanation script, the live explanation video corresponding to the real-time explanation script and live effect data.
[0023] The multi-modal training data can have multiple data sources and can include various data formats. For example, the multi-modal training data of imported maternal milk powder can include: video data, such as historical live commentary video clips; text data, such as “100-year history of the dairy industry” text from the brand website, “milk source / Docosahexaenoic Acid (DHA) content” parameter text from the e-commerce platform, and “babies love it / no fire” social media notes; image data, such as product packaging images and live room screenshots; and Portable Document Format (PDF) data, such as import customs declaration and quality inspection reports.
[0024] The training data (i.e., the multi-modal training data set of historical live goods) can be manually annotated by technical personnel according to the pre-set product feature annotation rules, so that the trained multi-modal data processing model can automatically annotate (i.e., extract features) the production data (i.e., the multi-modal production data of the current live goods). It can be understood that annotation according to the pre-set product feature annotation rules means extracting and annotating product features belonging to conversion degree, risk degree, explanation logic, speech style, core selling point, user pain point, and compliance information, etc.
[0025] The conversion degree and the risk degree are product features for annotating live commentary video clips. The conversion degree can include “high conversion clip” or “low conversion clip”, and the risk degree can include “illegal clip”, “high-risk clip” or “no-risk clip”. The explanation logic can include the explanation order between any of the core selling point, user pain point, and compliance information, for example: milk source → quality inspection → brewing method. The speech style can include the proportion between colloquial explanation and professional term explanation, speech speed, etc. The core selling point and the user pain point depend on the specific live goods. For example, the core selling point of imported maternal milk powder can be “100% New Zealand milk source”, and the user pain point of imported maternal milk powder can be “fear of fire”. The compliance information is a product feature for annotating qualification documents such as import customs declaration and quality inspection reports of live goods, for example, the compliance information of imported maternal milk powder can be “ingredient table location”, “certification conclusion”, “milk source identification”, etc.
[0026] It can be understood that the pre-set multi-modal data processing model needs to select a multi-modal understanding and generation model according to the multi-modal data processing needs of e-commerce live scenes (need to process video, voice, text, picture, PDF, etc.), such as the Multi-Modal Collaboration Model (MMC) model.
[0027] It can be understood that, when training the multi-modal data processing model, the multi-modal data processing model can transfer and induct the live commentary video segment in the multi-modal training data, convert the voice and picture information into structured commentary text, and then refer to the commentary text and the labeled product features to learn and extract product features from other data other than the live commentary video segment. The multi-modal production data processed by the trained multi-modal data processing model does not include the live commentary video segment compared with the multi-modal training data.
[0028] When training the preset multi-modal data processing model using the multi-modal training data set, domain adaptation fine-tuning can be performed, the general feature extraction layer at the bottom of the model is frozen, and only the live e-commerce dedicated live product feature extraction layer at the top is trained, so that the model learns the feature extraction logic in the live e-commerce field. When training the live product feature extraction layer, a specific product feature extraction task and its corresponding training data can be set to enable the model to learn the corresponding feature extraction logic. For example, for the "explanation logic extraction" task, the labeled live commentary video segment (including voice-to-text and picture annotation) is input to enable the model to learn "high conversion; typical logic of the segment: safety certification → functional advantage → preferential activity"; for the "core selling point extraction" task, the image-text data of the product detail page is input to enable the model to learn the association between "the picture of the milk source map" and "the 'New Zealand milk source' selling point in the text".
[0029] The mixture of experts (MoE) is an artificial intelligence architecture that improves model efficiency and performance through dynamic division of labor. The core idea is to decompose complex tasks into multiple subtasks, which are processed by different "expert" subnetworks, and then dynamically select the most suitable combination of experts through a "gating network" for calculation.
[0030] The generated compliance explanation script is applied to live room explanation. In the actual explanation process, the live room barrage is collected, the multi-modal production data of the current live product is matched according to the live room user question content in the live room barrage, and the compliance explanation script is updated in real time to generate a real-time explanation script. For example, in the live broadcast, the question "whether it contains sucrose" in the barrage generates the answer "does not contain sucrose, the ingredient table is on page 3 of the quality inspection report" in real time, which is supplemented into the compliance explanation script to form a real-time explanation script; for example, in the live broadcast, the question "whether it is waterproof" in the barrage generates the corresponding language and supplements it into the compliance explanation script to form a real-time explanation script.
[0031] Among them, it can be understood that in addition to updating the compliance explanation script according to the live room bullet screen, the content operation and the host can also refine the compliance explanation script, adjust the language style to adapt to the personal characteristics of the host, that is, the real-time explanation script can be the version after the compliance explanation script is updated and refined. In addition, the refined script and the modification record can also be stored in the database as incremental data for subsequent model training / optimization.
[0032] The live commodity explanation text generation method provided by the embodiment of the application, acquires a historical live commodity multi-modal training data set, trains a live commodity feature extraction layer of a preset multi-modal data processing model; wherein the multi-modal training data set is labeled according to a preset commodity feature labeling rule, and the commodity features include conversion degree, risk degree, explanation logic, language style, core selling point, user pain point and compliance information; according to the trained multi-modal data processing model, the multi-modal production data of the current live commodity and the live room user label and the hybrid expert model, a plurality of candidate explanation scripts are generated; according to the hybrid expert model and the plurality of candidate explanation scripts, a compliance explanation script is generated; the live room bullet screen corresponding to the compliance explanation script is collected, the compliance explanation script is updated, and a real-time explanation script is generated; according to the compliance explanation script, the live explanation video and the live effect data corresponding to the real-time explanation script, the hybrid expert model is optimized. Through automatic collection and generation, the script preparation time of the host is significantly reduced, and the live explanation efficiency can be improved; through the fusion of multi-source data and historical experience by the hybrid expert model, the omission of selling points caused by human negligence is avoided, and the live explanation quality can be improved; through the generation of candidate explanation scripts based on live room user labels, the conversion rate can be improved; through the generation of compliance explanation scripts based on the hybrid expert model and the plurality of candidate explanation scripts, the risk of violation can be reduced; through the closed-loop training mechanism of optimizing the hybrid expert model according to the live explanation video and the live effect data corresponding to the compliance explanation script and the real-time explanation script, the model is iterated with live data, and the explanation content generation effect is continuously improved, achieving continuous optimization.
[0033] Further, the applicant of the present application also proposes that the conversion degree can be judged through the live explanation video segment and the live effect data corresponding thereto, and the following pseudo code is used for illustration: / / 1. Define threshold parameters (which can be adjusted according to actual business) Set add-to-cart conversion rate threshold = Y% Set click conversion rate threshold = Z% Set online user base threshold = W / / exclude invalid data with too few online users / / 2. Input data acquisition Live commentary video segment = { Commentary period: [start time, end time], Product information: {product ID, product name, product category}, Period data: { Online number: Real-time online number list, Clicks: The number of clicks on the product in this period, Add to cart number: The number of products added to the shopping cart in this period, Transaction number: The number of product transactions in this period } }; / / 3. Data preprocessing Calculate average online number = average value of online number in commentary period Calculate add-to-cart conversion rate = transaction number / add-to-cart number Calculate click-to-buy rate = transaction number / clicks / / 4. Conversion degree judgment logic If average online number >= online number base threshold value: / / Ensure data sample is valid / / Core judgment: add-to-cart conversion rate and click-to-buy rate are higher than the set threshold If (add-to-cart conversion rate > add-to-cart conversion rate threshold and click-to-buy rate > click conversion rate threshold): Mark this live commentary video segment as "high conversion segment" Else: Mark this live commentary video segment as "low conversion segment" Else: Output "Invalid data (insufficient online number), cannot determine"
[0034] Further, the applicant of the present application also proposes that the risk degree can be judged by matching regulation rule library and calling hybrid expert model, the following uses pseudo code to illustrate: / / 1. Initialize parameters and resources Load platform violation word library = [violation word1, violation word2,..., violation wordN] / / Contains sensitive words, violation promotion words, etc. Set large model calling parameters = {confidence threshold: 0.8, analysis prompt word: "Judge whether the following text contains platform violation words, if it is a false positive, please explain the reason"} / / 2. Extract the commentary text of the live commentary video segment Explanation text = Extract the anchor voice-to-text content from the live commentary video segment / / Assume that voice-to-text processing has been completed; / / 3. Preliminary banned word detection Initialize suspected banned word list = [] Traverse each word / phrase in the explanation text: If the word / phrase exists in the platform banned word library: Add the word / phrase to the suspected banned word list; / / 4. First risk labeling (high-risk prediction) if suspected banned word list is not empty: Automatically label the live commentary video segment as "high-risk segment" else: Output result: "no-risk segment" End of process; / / 5. Large model secondary verification (determine whether it is a false positive) / / Call the large model to analyze whether the suspected banned word is a real violation Large model input = { Text to be analyzed: explanation text, Suspected banned word: suspected banned word list, Prompt information: large model call parameters. Analysis prompt word } Large model output result = call large model interface (large model input); / / 6. Final violation determination Parse the large model output result: if the large model determines "there is a real banned word" and the confidence >= large model call parameters. Confidence threshold: Finally label the live commentary video segment as "violation segment" Record violation details = {banned word: suspected banned word list, judgment basis: large model output result. Reason} else if the large model determines "it is a false positive" (such as homophonic words, normal context words being mistakenly matched): Revoke the "high-risk segment" label and update it to "no-risk segment" Record the reason for the false positive = large model output result. Reason (used to optimize the banned word library) else: / / When the large model confidence is insufficient, keep the high-risk label and output "manual review is required" Keep the "high-risk segment" label and output "manual review is required"; / / 7. Output final result Return live commentary video clip ID, final risk level label, violation / misjudgment details.
[0035] In some embodiments, the generating of the plurality of candidate commentary scripts (i.e., involved in step 120) according to the trained multi-modal data processing model, the multi-modal production data of the current live commodity and the live room user label and the hybrid expert model can include: Figure 2 In step S210, the product features are extracted from the multi-modal production data of the current live commodity according to the trained multi-modal data processing model; In step S220, the personalized commentary strategy is generated by matching the product features with the live room user label; In step S230, the plurality of candidate commentary scripts are generated according to the product features, the personalized commentary strategy and the hybrid expert model.
[0036] Wherein, the live room user label is not specifically limited in the embodiments of the present application, and different live room user labels can be set according to different live commodities. For example, the live room user label can be divided according to age, region and consumption preference. It can be understood that the product features and the live room user label are semantically matched through historical live effect data. For example, historical live effect data shows that young groups focus on trendy functions and mom groups focus on safety certification. Therefore, when the live room user label is "25-35 year-old mom", the personalized commentary strategy can be "focus on safety certification and taste".
[0037] Wherein, the number of candidate commentary scripts is not specifically limited in the embodiments of the present application, and for example, it can include 3-5.
[0038] Further, the applicant of the present application also proposes that the preset multi-modal data processing model can be supervised and trained twice through index evaluation and business feedback to ensure that the model processing result meets the actual demand. The index evaluation refers to determining the scoring coefficient of the multi-modal data processing model during the training of the preset multi-modal data processing model according to the multi-modal training data set. Only when the scoring coefficient meets the corresponding target threshold, the training is ended. The business feedback refers to determining the scoring coefficient of the multi-modal data processing model during the actual inference process after the multi-modal data processing model is trained. If the scoring coefficient does not meet the corresponding target threshold, the live commodity feature extraction layer of the trained multi-modal data processing model is trained again.
[0039] Correspondingly, in some embodiments, as shown in Figure 3 As shown, the method further comprises: In step S310, the multi-modal production data of the current live broadcast commodity is labeled according to the preset commodity feature labeling rule. In step S320, the scoring coefficient of the trained multi-modal data processing model is determined according to the labeled commodity features of the multi-modal production data and the commodity features extracted from the multi-modal production data by the trained multi-modal data processing model. In step S330, if the scoring coefficient does not satisfy the corresponding target threshold, multi-modal training data including the multi-modal production data is supplemented to the multi-modal training data set, and the live broadcast commodity feature extraction layer of the trained multi-modal data processing model is retrained.
[0040] In some embodiments, the scoring coefficient includes explanation logic extraction accuracy, core selling point extraction accuracy, user pain point extraction matching degree, and compliance information extraction accuracy; the target threshold corresponding to the explanation logic extraction accuracy includes 90%, the target threshold corresponding to the core selling point extraction accuracy includes 95%, the target threshold corresponding to the user pain point extraction matching degree includes 85%, and the target threshold corresponding to the compliance information extraction accuracy includes 100%.
[0041] Among them, the explanation logic extraction accuracy reflects the coincidence degree of the model extracted explanation logic and the artificially labeled explanation logic; the core selling point extraction accuracy reflects the range of the model extracted core selling point covering the artificially labeled core selling point (such as "DHA content" "sucrose-free"); the user pain point extraction matching degree reflects the coincidence degree of the model extracted user pain point and the artificially labeled user pain point (such as "whether it is hot"); and the compliance information extraction accuracy reflects whether the model extracted compliance information (such as "import customs declaration number" "quality inspection conclusion") is consistent with the artificially labeled one on the original.
[0042] For example, the performance of the model on each index is counted every week, if the "explanation logic extraction accuracy" decreases to 85%, 1000 newly labeled high conversion video clips are supplemented to fine-tune the model again; if there is an error in the compliance information extraction, the training data of the PDF analysis module in the multi-modal data processing model is strengthened (such as adding fuzzy scanned copies and multi-language customs declaration sample).
[0043] In addition, the applicant also proposes that the business feedback of the multi-modal data processing model is not limited to this, if the model extracts a certain core selling point, the core selling point also exists in the compliance explanation script, but the anchor does not explain it in the actual live broadcast, resulting in high-frequency inquiries about the selling point in the barrage, at this time, the selling point can also be supplemented and labeled as "missed" in the multi-modal production data and added to the next round of multi-modal training data set.
[0044] In some embodiments, the generating the compliance script (i.e. involved in step 130) according to the hybrid expert model and the plurality of candidate scripts can include: Figure 4 In step S410, a target script is determined from the plurality of candidate scripts according to the hybrid expert model, a plurality of preset selection dimensions and weights corresponding to each selection dimension; the plurality of selection dimensions include core selling point coverage, language fluency, user interest matching degree and compliance risk value. In step S420, the target script is corrected according to the hybrid expert model to generate a compliance script.
[0045] In some embodiments, the hybrid expert model, the plurality of preset selection dimensions and the weights corresponding to each selection dimension are used to score the plurality of candidate scripts, and then the script with the highest score is determined from the plurality of candidate scripts as the target script. It can be understood that the embodiments of the present application are not limited to this, but can also be selected according to only part of the selection dimensions, for example, the script with core selling point coverage exceeding 90% and compliance risk value of 0 is selected as the target script.
[0046] In some embodiments, the weights corresponding to each selection dimension are not specifically limited by the embodiments of the present application, for example, the weights corresponding to core selling point coverage, language fluency, user interest matching degree and compliance risk value can respectively include 30%, 20%, 30% and 20%.
[0047] Firstly, the hybrid expert model can call a compliance rule library, which can be constructed and dynamically updated based on laws and regulations (such as the "Advertising Law" forbidden words) and e-commerce platform rules, covering absolute language, false propaganda, category-specific restrictions, and other provisions. Further, the hybrid expert model performs multi-dimensional scanning and verification on the target explanation script according to the compliance rule library: keyword detection, such as identifying forbidden words (such as "top" and "best") and sensitive expressions (such as unfounded efficacy claims) in the script; logic verification, such as checking whether the expression has ambiguity (such as "suitable for all people"); and qualification matching, such as verifying whether claims such as "imported" and "organic" are supported by corresponding qualification documents (such as customs declarations and test reports). Further, the hybrid expert model automatically corrects the target explanation script: replacement, such as replacing forbidden words with compliant expressions (e.g., "best" to "quality"); adjustment, such as supplementing qualifying conditions to eliminate ambiguity (e.g., "suitable for all babies" to "suitable for 3-12 month babies, see instructions for details"); supplementation, such as inserting file key information for content requiring qualification support (e.g., "see page 2 of the quality inspection report"); and deletion, such as removing uncorrectable violations (e.g., unfounded "treatment efficacy"). Further, the hybrid expert model performs a second verification on the target explanation script to ensure that the compliance risk is zero, outputs a compliant explanation script and a correction record, and provides content operations and anchors for fine-tuning and model training / optimization.
[0048] In some embodiments, the hybrid expert model includes a gating network, a content generation sub-model, a content scoring sub-model, and a compliance checking sub-model.
[0049] Among them, it can be understood that the gating network is used to dynamically select the most suitable expert combination from the content generation sub-model, the content scoring sub-model, and the compliance checking sub-model, the content generation sub-model is used to generate multiple candidate explanation scripts, the content scoring sub-model is used to score the multiple candidate explanation scripts respectively to determine the target explanation script, and the compliance checking sub-model is used to correct the target explanation script to generate a compliant explanation script.
[0050] In some embodiments, the hybrid expert model is optimized (i.e., involved in step 150) based on the compliant explanation script, the live explanation video corresponding to the real-time explanation script, and live effect data, as shown in Figure 5 may include: In step S510, the longest product feature explanation segment in the live explanation video is identified; In step S520, in a case where it is determined according to the live effect data that the conversion rate of the product feature explanation segment is greater than a preset conversion rate threshold, the parameters of the hybrid expert model are adjusted to improve the weight and time period priority of the product feature corresponding to the product feature explanation segment in the candidate explanation script. In step S530, the hybrid expert model is positively supervised and trained according to the script segment corresponding to the product feature explanation segment in the real-time explanation script. In step S540, the collaborative rules of the hybrid expert model are updated to improve the user interest matching degree score of the hybrid expert model for the candidate explanation script including the product feature corresponding to the product feature explanation segment. In step S550, in a case where it is identified according to the live effect data and the live explanation video corresponding to the real-time explanation script that there is illegal content in the compliant explanation script, the compliance rule library of the hybrid expert model is updated according to the identified illegal content.
[0051] Among them, the live effect data includes effect indicators corresponding to different live explanation video segments, such as user dwell time (relative value compared with other segments), user behavior data in this period (such as no exit rate, amount of barrage questions). According to the live effect data, the conversion rate of the identified product feature explanation segment can be calculated.
[0052] If the conversion rate is greater than the preset conversion rate threshold, it can be considered that the conversion rate of the identified product feature explanation segment is high, so that the weight and time period priority of the product feature involved in the segment in the candidate explanation script generated by the model can be adjusted. For example, it is identified that “brewing method” is a product feature involved in a segment with high conversion rate, so that the weight of the “brewing method” related product feature in the candidate explanation script is increased (such as from the original basic weight of 15% to 25%-30%), the explanation time of the “brewing method” related product feature in the candidate explanation script is increased (such as from the original proportion of 10% to 15%-20%), and the “brewing method” is arranged in the live golden time period in the candidate explanation script (such as 10-20 minutes after the start of the broadcast, combined with historical data to determine the high traffic period).
[0053] If the conversion rate is greater than the preset conversion rate threshold, it can also be considered that the identified product feature explanation segment uses high-quality language, so that the hybrid expert model is positively supervised and trained according to the corresponding script segment of the identified product feature explanation segment in the real-time explanation script, and the language style of the candidate explanation script generated by the model is optimized. The script segment corresponding to the identified product feature explanation segment in the real-time explanation script includes high-quality features such as the language style and explanation logic of the actual explanation of the anchor, which can be used as a reference sample for the content generation sub-model in the hybrid expert model. The content generation sub-model will use similar language style (such as step-by-step description and interactive prompt) and explanation logic (such as milk source, quality inspection, and brewing method) when generating candidate explanation scripts in the future.
[0054] At the same time, the collaboration rules of the hybrid expert model (i.e., the gating network parameters in the hybrid expert model) can also be updated, so that the content scoring sub-model in the hybrid expert model gives a higher "user interest matching degree" score to the candidate explanation script that includes the product feature involved in the identified product feature explanation segment with a high conversion rate when scoring multiple candidate explanation scripts in the future.
[0055] At the same time, the compliance rule library of the hybrid expert model can also be updated according to the live streaming explanation video and live streaming effect data corresponding to the real-time explanation script, to identify whether there is any rule violation content in the compliance explanation script that has not been checked and corrected by the compliance checking sub-model in the hybrid expert model in the past, so that the compliance checking sub-model in the hybrid expert model has the ability to check and correct the same rule violation content in the future.
[0056] In some embodiments, the multi-modal training data includes brand introduction, product information, social media feedback, live streaming explanation video, product label information, merchant cooperation information, and merchant qualification documents, and the merchant qualification documents include product certification, import customs declaration, and quality inspection report; the merchant qualification documents have been identified by OCR and verified for validity.
[0057] Among them, the brand story refers to the information of brand history, concept, honor, etc. from the brand's official website; the product information refers to the product parameters, price, sales, etc. data obtained through the open interface of the e-commerce platform; the social media feedback refers to the user's grass-roots notes and recommended content on the product on social media; the live streaming explanation video refers to the explanation video of the self-owned live streaming room, external well-known live streaming room, and brand live streaming room selling the product; the product label information refers to the ingredient list, nutritional component list, etc. of the product; the merchant cooperation information refers to the cooperation SKU list, sales mechanism (such as discount, full reduction, buy two get one free), inventory quantity, etc. For the merchant qualification documents, OCR identification and validity verification are performed by manual or algorithm after collection.
[0058] It can be understood that the multi-modal training data is not limited to brand introduction, commodity information, social media feedback, live broadcast explanation video, commodity label information, merchant cooperation information and merchant qualification file, and any information capable of reflecting the commodity characteristics of the live broadcast commodity can be included.
[0059] As a second aspect of the present application, an electronic device is provided, such as Figure 6 As shown in the figure, the electronic device comprises: one or more processors 101; a memory 102, having one or more computer programs stored thereon, when the one or more computer programs are executed by the one or more processors 101, so that the one or more processors 101 implement the live broadcast commodity explanation text generation method provided by the first aspect of the present application.
[0060] The electronic device can further comprise one or more I / O interfaces 103 connected between the processor 101 and the memory 102, configured to realize the information interaction of the processor 101 and the memory 102.
[0061] Among them, the processor 101 is a device with data processing capability, including but not limited to central processing unit (CPU) and the like; the memory 102 is a device with data storage capability, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH); the I / O interface (read-write interface) is connected between the processor and the memory, and can realize the information interaction of the processor and the memory, including but not limited to data bus (Bus) and the like.
[0062] In some embodiments, the processor 101, the memory 102 and the I / O interface 103 are connected with each other through the bus 104, and further connected with other components of the computing device.
[0063] As a third aspect of the present application, as shown in the figure, a computer readable medium is provided, having a computer program stored thereon, when the computer program is executed by the processor, the live broadcast commodity explanation text generation method provided by the first aspect of the present application is realized. Figure 7
[0064] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. Accordingly, the computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the method of any one of the above embodiments can be implemented. In the embodiments provided in the present application, any reference to memory, storage, database or other medium can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM), etc.
[0065] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Those skilled in the art should understand that the present application includes but is not limited to the contents described in the above specific embodiments and the accompanying drawings. Any modification that does not deviate from the functional and structural principles of the present application will be included in the scope of the claims.
Claims
1. A method for generating live-stream product description text, characterized in that, include: A multimodal training dataset of historical live-streamed products is obtained, and the live-streamed product feature extraction layer of a preset multimodal data processing model is trained. The multimodal training dataset is labeled according to preset product feature labeling rules, and the product features include conversion rate, risk level, explanation logic, sales style, core selling points, user pain points, and compliance information. Based on the trained multimodal data processing model, the multimodal production data of the current live-streamed products, the user tags in the live-streaming room, and the hybrid expert model, multiple candidate explanation scripts are generated. Based on the hybrid expert model and the multiple candidate explanation scripts, a compliant explanation script is generated; Collect the live stream comments corresponding to the compliance explanation script, update the compliance explanation script, and generate a real-time explanation script; The hybrid expert model is optimized based on the compliant explanation script, the live explanation video corresponding to the real-time explanation script, and the live performance data.
2. The method according to claim 1, characterized in that, The process generates multiple candidate explanation scripts based on a trained multimodal data processing model, multimodal production data of the current live-streamed products, user tags in the live-streaming room, and a hybrid expert model, including: Based on the trained multimodal data processing model, product features are extracted from the multimodal production data of the current live-streamed products; By matching the product features with the user tags in the live stream, a personalized explanation strategy is generated; Based on the product characteristics, the personalized explanation strategy, and the hybrid expert model, multiple candidate explanation scripts are generated.
3. The method according to claim 2, characterized in that, The method further includes: According to the preset product feature labeling rules, the multimodal production data of the current live-streamed product is labeled; The scoring coefficient of the trained multimodal data processing model is determined based on the product features labeled in the multimodal production data and the product features extracted from the multimodal production data by the trained multimodal data processing model. If the scoring coefficient does not meet the corresponding target threshold, multimodal training data including the multimodal production data is added to the multimodal training dataset, and the live-streaming product feature extraction layer of the trained multimodal data processing model is trained again.
4. The method according to claim 3, characterized in that, The scoring coefficients include the accuracy rate of extracting explanation logic, the accuracy rate of extracting core selling points, the matching degree of extracting user pain points, and the accuracy rate of extracting compliance information. The target thresholds for the accuracy rate of extracting explanation logic include 90%, the target thresholds for the accuracy rate of extracting core selling points include 95%, the target thresholds for the matching degree of extracting user pain points include 85%, and the target thresholds for the accuracy rate of extracting compliance information include 100%.
5. The method according to claim 1, characterized in that, The step of generating a compliant explanation script based on the hybrid expert model and the multiple candidate explanation scripts includes: Based on the hybrid expert model, multiple preset selection dimensions, and the weights corresponding to each selection dimension, a target explanation script is determined from the multiple candidate explanation scripts; wherein, the multiple selection dimensions include core selling point coverage, language fluency, user interest matching degree, and compliance risk value; The target explanation script is revised based on the hybrid expert model to generate a compliant explanation script.
6. The method according to claim 5, characterized in that, The hybrid expert model includes a gating network, a content generation sub-model, a content scoring sub-model, and a compliance inspection sub-model.
7. The method according to claim 1, characterized in that, The optimization of the hybrid expert model based on the compliant explanation script, the live explanation video corresponding to the real-time explanation script, and the live performance data includes: Identify the longest segment of product feature explanation in the live-streamed video; If the conversion rate of the product feature explanation segment is determined to be greater than the preset conversion rate threshold based on the live broadcast effect data, the parameters of the hybrid expert model are adjusted to increase the weight and time priority of the product feature corresponding to the product feature explanation segment in the candidate explanation script. Based on the script segment corresponding to the product feature explanation segment in the real-time explanation script, the hybrid expert model is trained under positive supervision. Update the collaboration rules of the hybrid expert model to improve the user interest matching score of the hybrid expert model for candidate explanation scripts that include product features corresponding to the product feature explanation segments; If, based on the live explanation video and live effect data corresponding to the real-time explanation script, illegal content is identified in the compliant explanation script, the compliance rule base of the hybrid expert model is updated according to the identified illegal content.
8. The method according to any one of claims 1-7, characterized in that, The multimodal training data includes brand introductions, product information, social media feedback, live-streamed explanation videos, product tag information, merchant cooperation information, and merchant qualification documents. The merchant qualification documents include product certification certificates, import customs declarations, and quality inspection reports. The merchant qualification documents have been verified by OCR recognition and validity.
9. An electronic device, characterized in that, include: One or more processors; A memory having stored one or more computer programs thereon, which, when executed by one or more processors, cause the one or more processors to implement the method for generating live-stream product description text according to any one of claims 1 to 8.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for generating live-stream product description text as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Content generation method and device, readable medium, electronic equipment and program product
CN118354107A
Method, system and device for generating commodity recommendation content based on real-time bullet screen information and storage medium
CN118828124A
Commodity selling point extraction and generation method and system based on RAG and multi-modal model
CN119940304A
Method and device for generating real-time interpretation of a video
US20200007947A1
Video recommendation with multi-gate mixture of experts soft actor critic
US20220019878A1