Information processing method, electronic equipment, storage medium and product

By analyzing e-commerce live stream videos, audio, and audience comments, and combining this with historical user behavior data, product recommendations are dynamically adjusted. This solves the problem of mismatch between recommended products and audience interests in e-commerce live streams, resulting in higher user purchase intent and better recommendation effectiveness.

CN121486618APending Publication Date: 2026-02-06MIGU VIDEO TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511555724.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

In existing technologies, e-commerce live streaming lacks real-time correlation between product recommendations and viewers' current interests, resulting in poor recommendation performance and reduced user purchase intentions.

Method used

By acquiring video, audio, and text information, as well as audience comments, semantic analysis is performed using the BERT model to extract keywords and sentiment trends. Combined with users' historical behavior data, product recommendations are dynamically adjusted.

Benefits of technology

It improved the real-time nature and accuracy of product recommendations, significantly enhancing user purchase intention and recommendation effectiveness, especially in terms of click-through rate and conversion rate at crucial moments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486618A_ABST
    Figure CN121486618A_ABST
Patent Text Reader

Abstract

The invention discloses an information processing method, electronic equipment, a storage medium and a product. The method comprises the following steps: acquiring first video information, and determining first text information corresponding to audio in the first video information; determining first information based on the first text information, wherein the first information represents keywords for recommending commodities; obtaining second text information; the second text information represents the comment content of the audience on the first video information; determining second information based on the second text information, wherein the second information represents the emotional tendency of the audience to the first video information; and determining commodity recommendation information based on the first information and the second information. Therefore, when the user watches the video, the keyword of the related commodity contained in the audio can be determined according to the audio in the video; determining the emotional tendency of the audience to the video content according to the comments of the audience to the video content; and determining a commodity to be recommended to the user according to the keyword determined from the video and the emotional tendency determined from the audience comments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an information processing method, electronic device, storage medium and product. Background Technology

[0002] With the popularization of e-commerce live streaming, accurate product recommendations have become an important means to attract users and promote purchases. Related technologies primarily rely on users' historical behavior, generating recommended products by analyzing their browsing and purchase records. However, these technologies lack real-time correlation with viewers' current interests when recommending products, leading to a mismatch between recommended products and viewers' current interests, thus reducing users' willingness to buy and the effectiveness of the recommendations. Summary of the Invention

[0003] In view of this, embodiments of this application provide an information processing method, electronic device, storage medium, and product, which aim to make recommended products more compatible with the current interests of the audience, thereby improving the user's willingness to purchase and the recommendation effect.

[0004] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide an information processing method, the method comprising: Obtain first video information and determine the first text information corresponding to the audio in the first video information; First information is determined based on the first text information, and the first information represents keywords used to recommend products; Obtain second text information; the second text information represents the audience's comments on the first video information; Based on the second text information, second information is determined, which represents the viewer's emotional tendency toward the first video information; Product recommendation information is determined based on the first information and the second information.

[0005] In the above scheme, determining product recommendation information based on the first information and the second information includes: Obtain third information, which includes: the user's historical product browsing history, the user's historical purchase history, the user's interaction behavior records with the first video information, and product information sold by the e-commerce platform; The product feature vector is determined based on the product information sold on the e-commerce platform; A first feature vector is determined based on the first information, the second information, and the third information; Candidate product information is determined based on the first feature vector and the product feature vector; The recommended product information is determined based on the candidate product information.

[0006] In the above scheme, determining candidate product information based on the first feature vector and the product feature vector includes: For each product, a first product score is determined based on the first feature vector and the product feature vector; The similarity information between the first feature vector and the product feature vector is determined based on the first product score; The candidate product information is determined based on the similarity information.

[0007] In the above scheme, determining the product recommendation information based on the candidate product information includes: A second feature vector is determined based on the candidate product information; The score of the second product is determined based on the first text information, the second information, and the second feature vector. The product recommendation information is determined based on the second product score.

[0008] The method in the above scheme further includes: The first video information and the product recommendation information are displayed.

[0009] The method in the above scheme further includes: Obtain a third feature vector, which includes features of products sold in live streaming rooms on e-commerce platforms; The recommended products for sale in the live broadcast room are determined based on the product recommendation information and the third feature vector; Display third information related to the live stream.

[0010] The above solution is applied to an electronic device, which includes a first display screen and a second display screen, or the electronic device includes a first electronic device having a first display screen and a second electronic device having a second display screen. The method further includes: The system controls the first display screen to display the first video information and controls the second display screen to display the product recommendation information and the third information.

[0011] In a second aspect, embodiments of this application provide an electronic device, including a processor and a memory for storing a computer program capable of running on the processor, wherein the processor executes the computer program to implement the steps of the method described in the first aspect.

[0012] Thirdly, embodiments of this application provide a computer storage medium storing a computer program, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0013] Fourthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0014] The technical solution provided in this application embodiment obtains first video information and determines first text information corresponding to the audio in the first video information; determines first information based on the first text information, the first information representing keywords used to recommend products; obtains second text information; the second text information represents the comments of the audience on the first video information; determines second information based on the second text information, the second information representing the audience's emotional tendency towards the first video information; and determines product recommendation information based on the first information and the second information.

[0015] In this way, when users watch videos, the system can determine the specific text content of the audio based on the video's audio, and then identify keywords related to products contained in the audio; it can also obtain viewers' comments on the video content (such as bullet comments), and then determine the viewers' sentiment towards the video content; based on the keywords identified from the video and the sentiment determined from the viewers' comments, it can determine the products to recommend to the user. The solution provided in this application can more accurately capture the real-time needs of viewers during video viewing, making the recommended products more closely match the viewers' current interests, thereby increasing users' willingness to purchase and improving the recommendation effect. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the information processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the information processing device provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0017] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.

[0019] With the popularization of e-commerce live streaming, accurate product recommendations have become an important means to attract users and promote purchases. Related technologies primarily rely on users' historical behavior, generating recommended products by analyzing their browsing and purchase records. However, these technologies lack real-time correlation with viewers' current interests when recommending products, leading to a mismatch between recommended products and viewers' current interests, thus reducing users' willingness to buy and the effectiveness of the recommendations.

[0020] Taking live sports broadcasts as an example, the technology used to recommend products lacks real-time correlation with the live sports content that viewers are watching. It usually ignores the interactive data and emotional tendencies of viewers during the live broadcast and cannot dynamically adjust the product recommendation content, thus missing many potential purchase opportunities.

[0021] In various embodiments of this application, when a user watches a video, the system can determine the specific text content of the audio based on the audio in the video, and then determine the keywords related to the product contained in the audio; it can also obtain viewers' comments on the video content (such as bullet comments), and then determine the viewers' emotional inclination towards the video content; based on the keywords determined from the video and the emotional inclination determined from the viewers' comments, it can determine the products to be recommended to the user. The solution provided by this application can more accurately capture the real-time needs of viewers during the video viewing process, making the recommended products more compatible with the viewers' current interests, improving users' willingness to purchase and the recommendation effect.

[0022] This application provides an information processing method, such as... Figure 1 As shown, the method includes: Step 101: Obtain the first video information and determine the first text information corresponding to the audio in the first video information; Here, taking a live sports event as an example, the raw data V(t) of the live sports event video stream is obtained, where V(t) represents the live sports event video frame at time point t. Video frame decoding is performed on the live sports event video frames at each time point, thereby converting the live sports event video into a continuous live sports event image sequence F(t). The audio stream A(t) corresponding to each time point t is extracted using an audio separation module. Furthermore, speech recognition processing is performed on the audio data in the live sports event audio stream A(t) to obtain the live sports event commentary text Td(t), i.e., the first text information, where Td(t) represents the text content of the live sports event commentary at time point t. The methods for separating the video into images and audio, and for recognizing the audio into text, can be understood by referring to relevant technologies and will not be elaborated upon here.

[0023] Step 102: Determine first information based on the first text information, wherein the first information represents keywords used to recommend products; For example, the live commentary text Tp(t) of the event is input into the BERT model (a neural network model that excels at natural language processing) based on the Hugging Face Transformers library (a Python library that allows users to call pre-trained models), and the text content is semantically analyzed to identify keywords related to product recommendations.

[0024] Step 103: Obtain the second text information; the second text information represents the audience's comments on the first video information; Taking live sports broadcasts as an example, the text of audience comments C(t) is obtained from the interactive data interface of the live sports broadcast platform. This is the second text information. C(t) represents the specific text content of the audience comments at time point t, such as the content of the bullet comments posted by the audience.

[0025] In some embodiments, the live broadcast commentary text Td(t) and the live broadcast viewer commentary text C(t) are preprocessed, including removing stop words, word segmentation, and syntactic analysis, to obtain the preprocessed live broadcast commentary text. and comment text from viewers during the live event Here, the preprocessing tools can use neural network models from related technologies, which will not be elaborated further.

[0026] Step 104: Determine the second information based on the second text information. The second information represents the viewer's emotional tendency towards the first video information. For example, the text of viewers' comments during the live broadcast of the event. Inputting the text into a BERT model based on the Hugging Face Transformers library, semantic analysis is performed on the text content to identify the audience's emotional inclination towards the event; Specifically, the text of the live commentary for the event and comment text from viewers during the live event As input data, the token sequence X of the live broadcast commentary text is obtained by semantic analysis of the pre-trained BERT model. T (t) and the token sequence X of the live event viewer comment text C (t),X T (t) and X C (t) is represented as: (1) (2) Among them, X T (t) and X C (t) represents the token sequence of the live broadcast commentary text and the live broadcast audience commentary text at time point t, respectively.

[0027] Then, the pre-trained BERT model is used on X. T (t) and X C (t) performs semantic encoding, converting the text data into a corresponding high-dimensional vector representation. This represents a high-dimensional vector representation of the live commentary text for a sporting event. A high-dimensional vector representation of the text of viewers' comments during a live sports event.

[0028] (3) (4) High-dimensional vector representation of live sports commentary text High-dimensional vector representation of live event viewer comments Semantic analysis is performed, and the BERT model's self-attention mechanism is used to identify a set of keywords related to product recommendations. That is, the first piece of information and the sentiment vector. That is, the second piece of information: (5) (6) in, , , These represent the high-dimensional vectors of the live commentary text for the event. The query matrix, key matrix, and value matrix. , , These represent the high-dimensional vectors of the audience comments text during the live broadcast of the event. The query matrix, key matrix, and value matrix. The scaling factor is the vector dimension; the Softmax function is used to calculate attention weights, thereby extracting keywords relevant to product recommendations. and sentiment information Here, the data processing procedures of the BERT model and self-attention mechanism can be understood by referring to relevant technologies.

[0029] In practical applications, it can be The keywords in the text are sorted by weight to obtain the top m keywords. , denoted as: (7) in, Between 0 and 1, for example, extracting the top 3 keywords by weight from a segment of match commentary at a certain point in time. This can be represented as: ("Player's name", 0.42), (Jersey, 0.35), (Limited Edition, 0.18); Sentiment information. It can be represented as (0.87, 0.10, 0.03), where 0.87 represents the probability of positive emotion, 0.10 represents the probability of neutral emotion, and 0.03 represents the probability of negative emotion.

[0030] Step 105: Determine product recommendation information based on the first information and the second information.

[0031] Here, recommending products based on extracted keywords and audience sentiment can be understood as follows: while a user is watching a video, the system can determine the specific text content of the audio, and thus identify product-related keywords contained within the audio; it can also obtain audience comments on the video content (such as bullet comments), thereby determining the audience's sentiment towards the video content; and based on the keywords determined from the video and the sentiment determined from the audience comments, it determines the products to recommend to the user. The solution provided in this application can more accurately capture the audience's immediate needs during video viewing, making the recommended products more closely match the audience's current interests, thereby increasing the user's purchase intention and recommendation effectiveness.

[0032] It should be noted that the first video information in this application embodiment is a live video of a sports event, but the first video information can also be a recorded video, or a video of a movie, TV series, etc. The specific content of the first video information is not limited in this application embodiment.

[0033] In some embodiments, determining product recommendation information based on the first information and the second information includes: Obtain third information, which includes: the user's historical product browsing history, the user's historical purchase history, the user's interaction behavior records with the first video information, and product information sold by the e-commerce platform; The product feature vector is determined based on the product information sold on the e-commerce platform; A first feature vector is determined based on the first information, the second information, and the third information; Candidate product information is determined based on the first feature vector and the product feature vector; The recommended product information is determined based on the candidate product information.

[0034] Here, multiple source datasets (third information) are obtained, including: historical browsing behavior dataset Db, which is the user's historical product browsing records; purchase record dataset Dp, which is the user's historical purchase records; real-time interaction dataset Di, which is the user's interaction behavior records with the first video information; and product dataset Ds, which is the product information sold by the e-commerce platform.

[0035] The user's historical product browsing history, historical purchase history, and product information sold on the e-commerce platform can be obtained from the data interface provided by the e-commerce platform. The user's interaction behavior records with the first video information can be obtained from the data interface provided by the live streaming platform. The product dataset Ds can include: product inventory units, prices, inventory and sales data, etc.; the user's historical product browsing history can include: user behaviors such as clicking, swiping, adding to favorites, adding to cart, placing orders, sharing, commenting, etc. when browsing the e-commerce platform on the terminal (such as mobile phone, computer, etc.).

[0036] In some embodiments, if a user does not log in to the live streaming platform, or does not authorize the data interface provided by the live streaming platform to transmit their interactive behavior records, or the user does not engage in any interaction (such as liking, commenting, etc.), then the aforementioned user's interactive behavior records on the first video information can be replaced with the interactive behavior records of other viewers watching the first video information (such as 300 likes and 150 comments on the live stream), that is, records of other viewers' clicks, likes, or comments on event-related information. If a user logs in to the live streaming platform, or authorizes the data interface to transmit their interactive behavior records, or engages in interaction while watching the match, then the real-time interaction dataset Di is replaced with the user's own interactive behavior records on the first video information, or the real-time interaction dataset Di includes both the user's own interactive behavior records and the interactive behavior records of other viewers.

[0037] Based on the above example, the historical browsing behavior dataset Db, purchase record dataset Dp, real-time interaction dataset Di, and product dataset Ds obtained from the multi-source dataset are combined with the keyword set. The data is merged with the sentiment vector Se(t) to generate a fused feature vector. That is, the first eigenvector.

[0038] Specifically, features are extracted from the historical browsing behavior dataset Db and the purchase record dataset Dp to generate corresponding historical browsing behavior feature vectors. and purchase record feature vector Feature extraction is performed on the real-time interactive dataset Di to generate a real-time interactive feature vector. Feature extraction is performed on the product dataset Ds to generate product feature vectors. Combine the keyword set Kw(t) and the sentiment vector. As semantic features, generate semantic feature vectors. and sentiment feature vector The generated feature vectors are weighted and fused to produce the final fused feature vector. Formula (8) is used to calculate the fusion feature vector. The formula: (8) in, , , , , and These represent the weighting coefficients for historical browsing behavior, purchase records, real-time interactions, product data, keywords, and sentiment feature vectors, respectively. These coefficients are used to adjust the weights of each feature during the fusion process, resulting in a fused feature vector. This reflects audience preferences and the relevance of live event content to products. Here, , , , , and The specific value can be determined according to the actual situation, and the embodiments of this application do not limit it.

[0039] In some embodiments, determining candidate product information based on the first feature vector and the product feature vector includes: For each product, a first product score is determined based on the first feature vector and the product feature vector; The similarity information between the first feature vector and the product feature vector is determined based on the first product score; The candidate product information is determined based on the similarity information.

[0040] Based on the above example, the feature vectors will be fused. With product feature vector Input the product recommendation generation module to perform matching and calculate the matching score of each product, i.e., the score of the first product. Formula (9) is the formula for calculating the score of the first product: (9) Where M(i,t) represents the matching score of product i with the fused feature vector at time t, and d represents the feature dimension, which can be 256 in practical applications. and These represent the components of the fused feature vector and the product feature vector in the j-th dimension, respectively. The weights for dimension j are used to reflect the importance of each dimension; For bias terms, M(i,t) represents the standard deviation. A larger value for M(i,t) indicates that the product aligns better with the audience's current preferences and the context of the event. With weight Bias All of these are learnable parameters, which are automatically optimized through backpropagation during the training phase.

[0041] Furthermore, normalizing M(i,t) and adding an attention term yields the comprehensive similarity, i.e., the similarity information S(i,t), which can be expressed as: (10) in, represents normalization, and Attn represents the attention term.

[0042] Specifically, formula (11) is the detailed formula for calculating similarity information S(i,t).

[0043] (11) Where S(i,t) represents the comprehensive similarity score between product i and the fused feature vector at time point t. In formula (11), the first term (the term before the plus sign) is Gaussian similarity, and the second term (the term after the plus sign) is similarity supplement based on attention mechanism. These are the weighting coefficients in the attention mechanism. , , Let be the query matrix, key matrix, and value matrix of the kth attention head, respectively, where k' represents the number of attention heads; is the scaling factor for the vector dimension.

[0044] Each product is sorted according to its comprehensive similarity score S(i,t), and the top N products with the highest comprehensive similarity scores are selected to generate a product recommendation candidate set based on formula (12). This refers to candidate product information.

[0045] (12) Where Threshold(t) is the threshold value of the candidate recommendation set at time point t. The threshold value can be dynamically adjusted according to the actual situation to achieve a balance between the number of candidate products and the relevance of the recommendations. This application does not limit the specific value of the candidate recommendation set threshold or the number of products in the candidate product set (i.e., the specific value of N).

[0046] In some embodiments, determining the product recommendation information based on the candidate product information includes: A second feature vector is determined based on the candidate product information; The score of the second product is determined based on the first text information, the second information, and the second feature vector. The product recommendation information is determined based on the second product score.

[0047] For example, the product recommendation candidate set Input a pre-trained convolutional neural network model, and use the pre-trained convolutional neural network model to process the product candidate set. Feature extraction is performed to generate a deep feature vector for each product. ; Here, the convolutional neural network model can be ResNet-50 or EfficientNet, etc., which includes multi-layer convolution and non-linear transformation, and can upgrade low-level edges / textures in product features to high-level semantics (brand logo, equipment shape, jersey pattern, material and texture distribution, color scheme and style, logo / number, etc.).

[0048] Next, the high-dimensional vector of the live commentary text will be... Vector of audience emotional inclination By using a multilayer perceptron model, the vector is compared with the deep feature vector of the product. The data is then fused to generate a matching score vector for the product. That is, the score of the second item, and formula (13) is used to calculate the matching score vector. The formula.

[0049] (13) Where ReLU is the activation function; The weight matrix is ​​the linear mapping parameter of the multilayer perceptron, used to... Emotional Vector With product depth feature vector The concatenation is used for projection and weighted fusion; ⊕ represents the vector concatenation operation. For bias terms; The matrix is ​​trainable and optimized through backpropagation.

[0050] Next, the matching score vector of the product is... The sorting module reorders the products, selects the top Nc products with the highest matching scores, and generates the final product recommendation list. This refers to product recommendation information. Here, this embodiment of the application does not limit the number of products in the product recommendation list (i.e., the specific value of Nc).

[0051] Specifically, Features can include semantics, visual characteristics, click-through rate, conversion rate, price / inventory, etc., for The sorting formula can be formula (14), which sorts the products by comparing the size of U(i,t).

[0052] (14) Where w is the weight, λ is the penalty weight, and penalty(i,t) is the penalty term. For example, a product may have a high negative review rate and low inventory. In this case, calculating U(i,t) requires correspondingly adding penalty(i,t). It can be understood that the higher the negative review rate and the lower the inventory, the higher the value of penalty(i,t). Here, the specific values ​​of w, λ, and penalty(i,t) can be determined according to the actual situation. This application embodiment does not limit the specific values ​​of w, λ, and penalty(i,t).

[0053] For example, if =(0.9,0.3,0.1), =(0.7,0.5,0.2), set the weights w=(0.5,0.3,0.2), and calculate... The score is 0.9×0.5+0.3×0.3+0.1×0.2=0.56; The score is 0.7×0.5+0.5×0.3+0.2×0.2=0.54; The corresponding product is listed first. The corresponding products are listed below.

[0054] In some embodiments, if preset keywords, such as "a player's name" or "a brand's name," are detected in the live sports commentary text, they can be displayed in the product recommendation list. The system prioritizes displaying products that match preset keywords, such as a player's jersey or shoes from a specific brand.

[0055] In some embodiments, the method further includes: The first video information and the product recommendation information are displayed.

[0056] Here, the first video information (i.e., the video the user wants to watch) and the product recommendation information (i.e., the list of products recommended to the user) are displayed to the user, so that the user can view the products that appear in the video or are strongly related to the video at any time while watching videos such as competitions, thereby increasing the user's willingness to buy and the recommendation effect.

[0057] In some embodiments, the specific page displaying the product recommendation information can be an HTML5 (a web technology standard) page, with the e-commerce platform providing the URL (Uniform Resource Locator) and embedding it on the display screen of an electronic device in the form of WebView / mini-program, so that users can view the product recommendation information without closing the live streaming platform software.

[0058] In some embodiments, the method further includes: Obtain a third feature vector, which includes features of products sold in live streaming rooms on e-commerce platforms; The recommended products for sale in the live broadcast room are determined based on the product recommendation information and the third feature vector; Display third information related to the live stream.

[0059] For example, the association analysis module is used to match the products in the product recommendation list with the products in the live-streaming sales window of the e-commerce platform. In the association analysis module, formula (15) is the formula for calculating the matching score.

[0060] (15) in, Na represents the correlation between product i in the product recommendation list and product j in the live-streaming sales showcase at time point t, where Na is the number of correlation features. The weight of the k-th feature. Let be the weight matrix of product i and product j on the k-th feature. and Let represent the depth feature vector of product i and the feature vector of product j in the live-streaming e-commerce showcase, respectively. ⊕ denotes the concatenation operation. This is the bias term. The association analysis module specifically includes a learnable multi-feature fusion network, which can be an MLP or FM / CrossNet, etc.

[0061] Then, each product can be associated with By comparing with a preset relevance threshold, products exceeding the threshold are displayed to the user along with live stream information (third information). This information can take the form of a small window playing the live stream video, or text or icons displayed on the page to indicate the existence of a live stream. This embodiment does not limit the specific form of the third information. Since the live stream sales window is a collection of salable products for users watching events, linking the product recommendation list to the live stream sales window provides users with products that can be sold immediately, avoiding situations where users encounter inconsistent prices, unsold inventory, or unverified merchant qualifications when they want to purchase a product.

[0062] It's understandable that the more attention users pay to a product or the higher its popularity, the more likely that product should be displayed.

[0063] Based on this, in some embodiments, when displaying a product recommendation list to a user, the displayed product recommendation list is dynamically adjusted based on the user's interaction with the product recommendation information and the sales data of the e-commerce platform to optimize the display order of the products.

[0064] Specifically, formula (16) is the formula for calculating the exposure rate of a product.

[0065] (16) in, This represents the visibility of product i at time point t. Let T(t) represent the attractiveness coefficient of product i at time point t, and let T(t) represent the total exposure time of all products in the product recommendation list at time point t.

[0066] For example, a product card occupies approximately 80% of the screen area of ​​a user's terminal (such as a mobile phone, computer, etc.), and the user stays on the product card for 2.5 seconds (exposure time of 2.5 seconds), resulting in... ≈2.0; The e-commerce platform statistics show that the product's click-through rate is 0.12 and conversion rate is 0.04. Taking into account the context, a weighted average of 0.1 is calculated. =0.26; T(t) is the total exposure time of all products. The exposure time of all products is recalculated every 30 seconds. ,according to The items are displayed to the user in order of size. Here, all items are recalculated. The time interval can be adjusted according to the actual situation, and this application embodiment does not limit it.

[0067] Taking live sports broadcasts as an example, if a player's jersey's E (exposure rating) rises to first place, the jersey's card will be moved to first place and its size increased. If a product's E (exposure rating) remains below a preset exposure threshold, that product will be removed from the display list or replaced with another product. Here, context-weighted filtering means that if preset keywords, such as "a player's name" or "a brand's name," are detected in the live sports broadcast commentary text, they can be added to the product recommendation list. The system prioritizes displaying products that correspond to preset keywords, such as a player's jersey or shoes from a certain brand. The higher the relevance between the description and a product, the higher the context-weighted value. Products not mentioned in the description will have a lower context-weighted value.

[0068] In some embodiments, the present application is applied to an electronic device, the electronic device including a first display screen and a second display screen, or the electronic device including a first electronic device having a first display screen and a second electronic device having a second display screen, the method further comprising: The system controls the first display screen to display the first video information and controls the second display screen to display the product recommendation information and the third information.

[0069] Specifically, in this embodiment, a main screen (first display screen) and a secondary screen (second display screen) are used to display the live broadcast of the event and product recommendation information, respectively. The secondary screen can be an external display connected to the electronic device on which the main screen is located. For example, if a computer has two displays, one display shows the live broadcast of the event and the other displays the product recommendation information.

[0070] The secondary screen can also belong to a different electronic device than the main screen. For example, the main screen is used to play the event, and the first electronic device where the main screen is located can be a TV, computer, or tablet, etc. The second electronic device where the secondary screen is located can be a user's mobile phone, computer, tablet, etc. The first electronic device and the second electronic device are different. In practical applications, the TV screen is the main screen and the mobile phone screen is the secondary screen. Users can log in to the live sports broadcasting platform with the same account on both the TV and the mobile phone to achieve data communication between the TV and the mobile phone. Alternatively, when the main screen is playing the event, it can generate an S-ID (Session ID) and an event-ID. The S-ID indicates that the user is using the live sports broadcasting platform and watching a specific event, and the event-ID indicates that the user wants to establish a connection between the second electronic device and the first electronic device. At this time, the main screen displays a QR code including the S-ID and event-ID. The user scans the code with the second electronic device (such as a mobile phone) to bind the relationship with the first electronic device. The first electronic device and the second electronic device are connected via WebSocket or SSE, that is, data is transmitted through a cloud server (e.g., the mobile phone can provide the cloud server with historical browsing behavior datasets, purchase record datasets, and product datasets recorded by the e-commerce platform, etc., and the TV can provide the cloud server with live sports broadcast video streams, etc.). It can be understood that the solution provided in this application embodiment can send relevant data to the first electronic device from the cloud server and execute it based on the processor of the first electronic device; it can also send relevant data to the second electronic device from the cloud server and run it based on the processor of the second electronic device; or it can perform calculations by the cloud server and send product recommendation information to the second electronic device.

[0071] In practical applications, users can watch live sports events on the first electronic device and display interactive data interfaces provided by the live sports event platform on the second electronic device. This means that users can view comments posted by other viewers (such as bullet comments) on the second electronic device, and can also perform actions such as clicking, liking, and commenting. At the same time, the second electronic device displays a product recommendation information page embedded in the form of WebView / mini-program.

[0072] In this way, the first electronic device and the second electronic device can run different operating systems without affecting the implementation of the solution in this application.

[0073] The technical solution provided in this application embodiment obtains first video information and determines first text information corresponding to the audio in the first video information; determines first information based on the first text information, the first information representing keywords used to recommend products; obtains second text information; the second text information represents the comments of the audience on the first video information; determines second information based on the second text information, the second information representing the audience's emotional tendency towards the first video information; and determines product recommendation information based on the first information and the second information.

[0074] In this way, when users watch videos, the system can determine the specific text content of the audio based on the video's audio, and then identify keywords related to products contained in the audio; it can also obtain viewers' comments on the video content (such as bullet comments), and then determine the viewers' sentiment towards the video content; based on the keywords identified from the video and the sentiment determined from the viewers' comments, it can determine the products to recommend to the user. The solution provided in this application can more accurately capture the real-time needs of viewers during video viewing, making the recommended products more closely match the viewers' current interests, thereby increasing users' willingness to purchase and improving the recommendation effect.

[0075] The solution of this application will be further described below with reference to application examples.

[0076] During a football match, e-commerce platforms collaborated with the live streaming platform to recommend relevant products to viewers through a dual-screen interactive method provided in this application. During the live broadcast, viewers could not only watch the match on television but also view the recommended products in real-time on a secondary screen on their mobile devices. After the match began, the commentator described a classic goal by a player from the home team and mentioned his limited-edition jersey. At this point, the commentary text was extracted from the live stream, and semantic analysis was performed using a deep learning model to identify the keywords "player's name" and "limited-edition jersey." Simultaneously, real-time comments from viewers on the secondary screen were captured, such as "I want to buy this jersey" and "That player's jersey is so cool."

[0077] Based on historical behavioral data provided by the e-commerce platform, it was discovered that viewers had repeatedly clicked on jersey-related products, especially those related to the player, in their browsing history over the past week. Real-time interaction data from the live streaming platform was also obtained, including likes, comments, and shares, particularly during the period when the commentator mentioned the player's goal. Product data from the e-commerce platform was also retrieved, such as the limited-edition jersey having 500 units in stock and a price of 120 yuan. Therefore, the player's jersey was prominently displayed on a secondary screen with a "Buy Now" button. Within 5 minutes of the recommended product being displayed, the click-through rate reached 15.6%; within 10 minutes, the purchase conversion rate was 4.2%, and 21 limited-edition jerseys were sold.

[0078] To train and validate the effectiveness of the system, the model of this application embodiment was trained using a comprehensive dataset including narration text, audience comments, historical browsing data, purchase records, and product information. The following are some training samples: Sample 1: Commentary: A player scored a brilliant goal in the 20th minute of the match, and his action and jersey once again became the focus; Audience comment: I want to buy this player's jersey; Historical behavior data: The user has viewed related products for this player multiple times in the past 7 days; Real-time interaction data: Number of likes: 300, Number of comments: 150; Recommended product: Limited edition jersey of a certain player, 500 pieces in stock; Experimental results: Click-through rate 18%, purchase conversion rate 5%, 25 pieces sold.

[0079] Sample 2: Commentary text: The commentator mentioned that the ball used in this match had a unique design, which the players praised highly; Audience comments: This match ball is really good-looking, where can I buy it; Historical behavior data: The user has not recently browsed football-related products; Real-time interaction data: Number of likes: 100, Number of comments: 60; Recommended product: Match ball, 1000 units in stock; Experimental results: Click-through rate 5%, Purchase conversion rate 1%, 10 units sold.

[0080] To verify the effectiveness of this invention, the performance of the traditional method and the method of this invention was compared during the live broadcast of this event. The specific data is shown in Table 1 below: Table 1

[0081] Comparative experiments show that the solution provided in this application significantly improves the real-time performance and accuracy of product recommendations. Especially during key events, by analyzing commentary texts and audience comments in real time, the system can quickly respond to audience needs, increasing user click-through rates and purchase conversion rates. Compared with traditional methods, the method of this invention increases click-through rates by about 5 times and purchase conversion rates by about 6 times during key moments. In addition, through real-time data feedback and dynamic adjustments, this invention can ensure that product recommendations are highly relevant to users' current interests, significantly improving the sales efficiency of e-commerce platforms.

[0082] In order to implement the method of the embodiments of this application, the embodiments of this application also provide an information processing device, which corresponds to the above-described information processing method. The steps in the embodiments of the above-described information processing method are also fully applicable to the information processing device embodiments.

[0083] like Figure 2 As shown, the information processing device includes: a first acquisition module 201, a first determination module 202, a second determination module 203, a third determination module 204, and a fourth determination module 205.

[0084] The first acquisition module 201 is used to acquire first video information, and the first determination module 202 is used to determine the first text information corresponding to the audio in the first video information. The second determining module 203 is used to determine first information based on the first text information, wherein the first information represents keywords used to recommend products; The first acquisition module 201 is also used to acquire second text information; the second text information represents the audience's comments on the first video information; The third determining module 204 is used to determine second information based on the second text information, wherein the second information represents the viewer's emotional tendency toward the first video information; The fourth determining module 205 is used to determine product recommendation information based on the first information and the second information.

[0085] In some embodiments, the fourth determining module 205 is specifically used for: Obtain third information, which includes: the user's historical product browsing history, the user's historical purchase history, the user's interaction behavior records with the first video information, and product information sold by the e-commerce platform; The product feature vector is determined based on the product information sold on the e-commerce platform; A first feature vector is determined based on the first information, the second information, and the third information; Candidate product information is determined based on the first feature vector and the product feature vector; The recommended product information is determined based on the candidate product information.

[0086] In some embodiments, the fourth determining module 205 is specifically used for: For each product, a first product score is determined based on the first feature vector and the product feature vector; The similarity information between the first feature vector and the product feature vector is determined based on the first product score; The candidate product information is determined based on the similarity information.

[0087] In some embodiments, the fourth determining module 205 is specifically used for: A second feature vector is determined based on the candidate product information; The score of the second product is determined based on the first text information, the second information, and the second feature vector. The product recommendation information is determined based on the second product score.

[0088] In some embodiments, the device further includes a display module 206, the display module 206 being configured to: The first video information and the product recommendation information are displayed.

[0089] In some embodiments, the apparatus further includes a second acquisition module 207 and a fifth determination module 208, wherein the second acquisition module 207 is configured to: Obtain a third feature vector, which includes features of products sold in live streaming rooms on e-commerce platforms; The fifth determining module 208 is used to: determine recommended products for sale in the live broadcast room based on the product recommendation information and the third feature vector; The display module 206 is also used to display third information related to the live broadcast room.

[0090] In some embodiments, the device is applied to an electronic device, the electronic device including a first display screen and a second display screen, or the electronic device including a first electronic device having a first display screen and a second electronic device having a second display screen, the device further including a control module 209, the control module 209 being configured to: The system controls the first display screen to display the first video information and controls the second display screen to display the product recommendation information and the third information.

[0091] It should be noted that the information processing device provided in the above embodiments is only illustrated by the division of the above program modules. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the information processing device and the information processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0092] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiments of this application, the embodiments of this application also provide an electronic device. Figure 3 The diagram shows only an exemplary structure of the electronic device, not the entire structure; implementation is possible as needed. Figure 3 The structure shown may be part or all of the structure.

[0093] like Figure 3 As shown, the electronic device 300 provided in this application embodiment includes at least one processor 301, a memory 302, and a user interface 303. The various components in the electronic device 300 are coupled together via a bus system 304. It can be understood that the bus system 304 is used to implement communication between these components. In addition to a data bus, the bus system 304 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 3 The general designated all buses as Bus System 304.

[0094] The user interface 303 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.

[0095] The memory 302 in this embodiment is used to store various types of data to support the operation of the electronic device 300. Examples of such data include any computer program used to operate on the electronic device 300.

[0096] The information processing method disclosed in this application embodiment can be applied to or implemented by the processor 301. The processor 301 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the information processing method can be completed by the integrated logic circuit of the hardware in the processor 301 or by instructions in the form of software. The processor 301 mentioned above may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 301 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the memory 302. The processor 301 reads the information in the memory 302 and, in conjunction with its hardware, completes the steps of the information processing method provided in the embodiments of this application.

[0097] In an exemplary embodiment, the electronic device 300 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.

[0098] It is understood that memory 302 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.

[0099] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 302 that stores a computer program. The computer program can be executed by the processor 301 of the electronic device 300 to complete the steps described in the method of this application embodiment. The computer-readable storage medium can be a ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.

[0100] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by a processor 301 of an electronic device 300 to perform the steps described in the method of this application embodiment.

[0101] It should be noted that terms such as "first" and "second" are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0102] It should be understood that the phrase "some embodiments" throughout the specification means that a particular feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "some embodiments" appearing throughout the specification does not necessarily refer to the same embodiment.

[0103] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.

[0104] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An information processing method, characterized in that, The method includes: Obtain first video information and determine the first text information corresponding to the audio in the first video information; First information is determined based on the first text information, and the first information represents keywords used to recommend products; Obtain second text information; the second text information represents the audience's comments on the first video information; Based on the second text information, second information is determined, which represents the viewer's emotional tendency toward the first video information; Product recommendation information is determined based on the first information and the second information.

2. The method according to claim 1, characterized in that, The step of determining product recommendation information based on the first information and the second information includes: Obtain third information, which includes: the user's historical product browsing history, the user's historical purchase history, the user's interaction behavior records with the first video information, and product information sold by the e-commerce platform; The product feature vector is determined based on the product information sold on the e-commerce platform; A first feature vector is determined based on the first information, the second information, and the third information; Candidate product information is determined based on the first feature vector and the product feature vector; The recommended product information is determined based on the candidate product information.

3. The method according to claim 2, characterized in that, The step of determining candidate product information based on the first feature vector and the product feature vector includes: For each product, a first product score is determined based on the first feature vector and the product feature vector; The similarity information between the first feature vector and the product feature vector is determined based on the first product score; The candidate product information is determined based on the similarity information.

4. The method according to claim 2, characterized in that, The step of determining the product recommendation information based on the candidate product information includes: A second feature vector is determined based on the candidate product information; The score of the second product is determined based on the first text information, the second information, and the second feature vector. The product recommendation information is determined based on the second product score.

5. The method according to claim 1, characterized in that, The method further includes: The first video information and the product recommendation information are displayed.

6. The method according to claim 5, characterized in that, The method further includes: Obtain a third feature vector, which includes features of products sold in live streaming rooms on e-commerce platforms; The recommended products for sale in the live broadcast room are determined based on the product recommendation information and the third feature vector; Display third information related to the live stream.

7. The method according to claim 6, characterized in that, The method is applied to an electronic device, which includes a first display screen and a second display screen, or the electronic device includes a first electronic device having a first display screen and a second electronic device having a second display screen, and the method further includes: The system controls the first display screen to display the first video information and controls the second display screen to display the product recommendation information and the third information.

8. An electronic device comprising a processor and a memory for storing a computer program capable of running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

9. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.