Human-computer interaction method based on e-commerce platform

Through multi-source information collection and multimodal data analysis, combined with user behavior tag generation and scene intelligent recognition, immersive interactive display and real-time personalized recommendations of e-commerce platforms are realized, which solves the problem of single human-computer interaction mode of existing e-commerce platforms and improves user shopping experience and platform efficiency.

CN120832020APending Publication Date: 2025-10-24GUICHANG TECH (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510934480.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

The human-computer interaction methods of existing e-commerce platforms are single and static, and cannot adapt to complex user behavior characteristics and diverse shopping scenarios. This leads to inaccurate product recommendations, delayed interactive responses, and fragmented information push, affecting user shopping efficiency and platform satisfaction.

Method used

Through multi-source information collection, multimodal data analysis, user behavior label generation, scene intelligent identification, immersive interactive display, real-time personalized recommendation, proactive interactive feedback and intelligent customer service linkage, a highly responsive and adaptable human-computer interaction service system is built to achieve an organic combination of multimodal interaction and scene intelligent identification.

Benefits of technology

It significantly improves the freedom and convenience of human-computer interaction on e-commerce platforms, realizes scenario-based and personalized product display and precise recommendations, solves the problems of slow response and fragmented customer service systems on traditional e-commerce platforms, and improves users' shopping decision-making efficiency and the platform's intelligence level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832020A_ABST
    Figure CN120832020A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of man-machine interaction equipment, and discloses a man-machine interaction method based on an e-commerce platform, and the method comprises the following specific steps: S1, collecting the multi-source demand information of a user; s2, multi-modal data analysis is carried out; s3, generating a user behavior tag; s4, intelligently identifying a user scene; s5, performing immersive interactive display; s6, performing real-time personalized recommendation; s7, active interaction feedback is carried out; s8, intelligent customer service linkage; and S9, interactive data closed-loop optimization. Through organic combination of multi-mode interactive input and scene intelligent recognition, a user can search for commodities and purchase requests in a more natural mode, multiple interactive modes such as voice input, image recognition and text retrieval are supported, the dependence of an existing e-commerce platform on single page click and keyword input is broken through, and the user experience is improved. According to the method, the man-machine interaction freedom degree and the input convenience of the e-commerce platform are remarkably improved, and the interaction requirements of different users in different scenes are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of human-computer interaction devices, and particularly relates to a human-computer interaction method based on an e-commerce platform. BACKGROUND

[0002] With the rapid development of e-commerce, the e-commerce platform has become an important channel for modern consumers to obtain goods and services, and users have higher requirements for the interactive experience of the e-commerce platform. The existing e-commerce platform generally adopts a traditional human-computer interaction mode based on page clicking and text input. Although this mode can meet the basic shopping needs to a certain extent, in actual application, this single and static interaction mode cannot adapt to the increasingly complex user behavior characteristics and diversified shopping scenarios.

[0003] Users expect to quickly and conveniently complete the goods search, comparison and ordering process through a more intelligent, personalized and multi-modal human-computer interaction mode, and obtain an immersive shopping experience in the process. However, there are still many problems in the human-computer interaction process of most e-commerce platforms at present, such as inaccurate goods recommendation, delayed interaction response, fragmented information push, single customer service response and the like, which seriously affect the shopping efficiency and platform satisfaction of users. Although some e-commerce platforms have introduced voice assistants, image search, intelligent customer service and the like, they mostly stay in the simple superposition of existing mature modules, lack systematic design and optimization of the human-computer interaction logic in the e-commerce scenario, and cannot realize substantial improvement of the interactive experience. SUMMARY

[0004] The application aims to provide a human-computer interaction method based on an e-commerce platform to solve the problems in the background.

[0005] To achieve the above-mentioned purpose, the application provides the following technical solution: a human-computer interaction method based on an e-commerce platform, and the specific steps of the human-computer interaction method based on the e-commerce platform are as follows:

[0006] S1, user multi-source demand information collection: multi-source input data of users on the e-commerce platform are collected, including text search words, voice instructions, image uploads, browsing tracks and historical purchase records, a user current behavior data set is constructed, and the set is used as a basic input for subsequent data processing;

[0007] S2, multi-modal data analysis: modal recognition and data analysis are performed on the multi-source input data of the user multi-source demand information collection, the specific types and contents of voice input, image input and text input are distinguished, and a unified data label of multi-modal input is established;

[0008] S3, user behavior label generation: based on the input data of multi-modal data analysis, the preset user behavior modeling algorithm is called to generate multi-dimensional user behavior labels including interest preference, browsing habit, price sensitivity, etc. for accurate matching of subsequent scenes and recommendations;

[0009] S4, user scene intelligent identification: through the user behavior label generated by the user behavior label, the preset scene model library is matched, the current consumption scene of the user (such as home, office, and baby) is automatically identified, and the scene identification parameter is outputted for subsequent display and recommendation;

[0010] S5, immersive interactive display: based on the scene identification parameter of the user scene intelligent identification, the layout structure, information level and display mode of the current platform page are dynamically adjusted, the text, video or 3D product model is loaded according to the scene priority, and the immersive and multi-dimensional interactive experience is realized;

[0011] S6, real-time personalized recommendation: according to the user click path and stay node of the immersive interactive display page content, combined with the user behavior label generation, the real-time recommendation algorithm is used to dynamically generate personalized product recommendation results, and the current page is pushed and updated in real time;

[0012] S7, active interactive feedback: when the user stays at a node of the real-time personalized recommendation page for more than a preset time or frequently switches the browsing path, the pop-up prompt, preferential information or customer service inquiry entrance based on the user scene and the recommended content are actively triggered to realize timely interactive feedback;

[0013] S8, intelligent customer service linkage: when the user requests customer service through the interface of active interactive feedback, the system synchronizes the real-time personalized recommendation link and the user active interactive feedback information to the intelligent customer service system in real time, and the intelligent customer service automatically calls the corresponding knowledge base content based on the current recommendation link node for quick response;

[0014] S9, interactive data closed loop optimization: the platform collects the whole process interactive data of S1 to S8 in real time, dynamically adjusts the recommendation algorithm weight, page display priority and feedback trigger threshold based on the interactive effect and user feedback, and continuously optimizes the subsequent user human-computer interaction experience.

[0015] Preferably, the specific steps of the user demand collection in S1 are as follows:

[0016] S11, multi-source data collection and behavior set construction: by collecting the active input of the user in the e-commerce platform such as text search words, voice instructions and uploaded images, combined with passive behavior data such as browsing track and historical purchase record, the system can integrate to form a comprehensive user current behavior data set. The convergence of these multi-dimensional information provides a basis and rich input materials for subsequent data processing;

[0017] The user behavior data set construction expression formula is:

[0018] D u ={T u ,V u ,I u ,B u ,P u}

[0019] In the formula, D u : user full behavior data set, T u : user text input set (such as keywords), V u : user voice input set, I u : user image input set, B u : user historical browsing behavior set, P u : user historical purchase record set

[0020] Comprehensive collection of user multi-source input data, including text, voice, image and other input methods, combined with user's historical browsing and purchase records, can accurately restore the user's current shopping demand state.

[0021] By constructing the user full data set, high-quality input basis is provided for subsequent interaction process, avoiding the scene recognition deviation caused by single input, and realizing more comprehensive human-computer interaction data modeling.

[0022] S12, data fusion lays the foundation for processing: the scattered user multi-source behavior data is organically fused to form a coherent behavior data set, which can accurately reflect the user's current demand and behavior characteristics; This set not only integrates various input information, but also becomes the core basic input for further analysis and mining of user potential demand.

[0023] Preferably, the specific steps of multi-modal data analysis in S2 are as follows:

[0024] S21, multi-modal input type recognition and content analysis: for user multi-source demand information, first translate the voice input into semantics to clarify the instruction intent; carry out feature extraction on image input to identify item attributes; implement keyword extraction on text input to grasp the search core. Through this classification analysis, the specific content of various modal data is accurately mined, laying the foundation for multi-modal feature integration;

[0025] S22, unified label fusion multi-modal feature: after completing the analysis of each type of input, a unified data label system is established for voice, image and text data. With the help of label, different modal information is associated to realize the organic integration of multi-modal features, so that the scattered demand data forms a whole and improves the accuracy of subsequent demand analysis;

[0026] The multi-modal feature analysis expression formula is:

[0027] F m =f parse (D u )={F t ,F v ,F i}

[0028] In the formula, F m : multi-modal input feature set, f parse : multi-modal data analysis function, F t : text input analysis feature, F v : speech input analysis feature, F i : image input analysis feature.

[0029] The multi-modal input data is structured and analyzed to effectively separate and extract features of speech, text and image, ensuring the accuracy and compatibility of data input.

[0030] The problem of fragmented multi-modal input processing of existing e-commerce platforms is solved, ensuring that data such as voice, image and text can be managed in a unified format, and improving the input quality of subsequent label generation and recommendation algorithms.

[0031] Preferably, the user behavior label generation in S3 refers to relying on the multi-modal data analysis results, the system calls a preset user behavior modeling algorithm, integrates multi-source information such as voice, image and text, and generates user behavior labels covering dimensions such as interest preference, browsing habit and price sensitivity. In this process, the algorithm associates multi-modal features with historical behavior patterns, combines with latent demand reasoning logic, makes the label accurately reflect user characteristics, provides reliable basis for subsequent scene adaptation and recommendation matching, and realizes effective conversion from data analysis to behavior description;

[0032] The user behavior label generation expression formula is:

[0033] L u =f label (F m ,B u ,P u )

[0034] In the formula, L u : user behavior label set (such as interest, preference, price sensitivity), f label : behavior label generation function, F m : multi-modal feature set of S2, B u : user historical browsing set, P u : user historical purchase set.

[0035] The input features of the user are fused with historical behaviors to generate dynamic and accurate user behavior labels, reflecting the current interests, purchase preferences and price sensitivity of the user.

[0036] Through label generation, personalized and dynamic user portraits are realized, avoiding the recommendation bias caused by static labels in the prior art, and improving the accuracy of subsequent scene recognition and recommendation.

[0037] Preferably, the user scene intelligent recognition in S4 means that based on the generated user behavior label, the system calls a preset scene model library for matching, and by calculating the similarity degree of the behavior label and each scene model feature, the current consumption scene is automatically recognized according to the association closeness. This process integrates the similarity comparison logic of scene matching, making the scene judgment more accurate, and finally outputs the scene identification parameter, providing a basis for the adaptation of subsequent display forms and recommended content, realizing efficient conversion from behavior label to scene positioning, and improving the fit of recommendation and user demand;

[0038] The scene matching similarity expression formula is:

[0039]

[0040] In the formula, S i : the current user scene matched, S: the preset scene library set, Sim(L u ,S j ): the similarity function between the user label and the scene label.

[0041] The user label set and the preset scene library can be matched according to the similarity to quickly identify the current shopping scene of the user, and realize scene-based product recommendation.

[0042] Through maximum similarity matching, it is ensured that the scene recognition result is highly consistent with the current demand of the user, breaking through the limitations of random recommendation or simple classification recommendation in the prior art.

[0043] Preferably, the specific steps of the immersive interactive display in S5 are as follows:

[0044] S51, page layout dynamically adapts to the scene: based on the scene identification parameter, the platform dynamically adjusts the page layout structure, information level and display mode, integrates the scene-based display weight distribution logic, plans element arrangement according to scene demand priority, makes the page highly consistent with the current scene, and improves the scene immersion and information acquisition efficiency of the user during browsing.

[0045] S52, multi-form content priority loading: according to the scene identification parameter and the weight distribution logic, priority is given to loading adaptive content such as graphic text, video or 3D commodity model, an immersive, multi-dimensional interactive experience is created by reasonably distributing the display weight of each form of content, scene-adapted information presentation is provided for users, and precise recommendation is landed;

[0046] The scene-based display weight distribution expression formula is:

[0047] W d =f disp (S i )={w img ,w vid ,w 3D}

[0048] In the formula, W d : page display weight set, f disp : scene-based display weight distribution function, S i : S4-identified user scene, w img : picture display weight, w vid : video display weight, w 3D : 3D model display weight.

[0049] The display proportion of page images, videos and 3D models is dynamically adjusted to ensure that the commodity display mode meets the best interactive experience under the current user scene.

[0050] The weight is controlled to arrange the page layout, effectively solving the problem of fixed page display and fragmented user experience in existing e-commerce pages, and realizing intelligent dynamic allocation of content.

[0051] Preferably, the real-time personalized recommendation in S6 means that based on the click path and stay node of the user in the immersive interactive page, the system combines these real-time behavior data with the generated user behavior tags, calls the real-time recommendation algorithm to dynamically generate personalized commodity recommendations, the algorithm analyzes the association between the current interactive preference and historical behavior characteristics of the user, accurately matches and adapts commodities, and immediately updates the push on the current page, so that the recommendation result dynamically follows the user behavior, improves the timeliness and relevance of the recommendation, and enhances the user shopping experience.

[0052] Preferably, the active interaction feedback in S7 means that when the user stays at a certain node in the real-time personalized recommendation page for more than a preset time or frequently switches the browsing path, the system will actively start the interaction feedback mechanism, combine the current user scene and the recommended content, accurately push the pop-up prompt, related preferential information, or show the customer service inquiry entrance. This timely response based on user behavior characteristics can not only resolve the user's possible decision hesitation, but also enhance the service stickiness through active interaction, so that the recommendation experience is more in line with the real-time needs of users.

[0053] Preferably, the intelligent customer service linkage in S8 means that when a user requests customer service through the active interactive feedback interface, the system will synchronize the real-time personalized recommendation link and user interactive feedback information to the intelligent customer service system. The customer service system combines the current recommendation link node with the response priority matching logic, automatically calls the corresponding knowledge base content, and gives priority to the needs that are highly relevant to the user's current scenario, so as to achieve a fast and accurate response, so that the customer service is closely connected with the recommendation process, and improve the user's problem-solving efficiency and interactive experience;

[0054] The customer service response priority matching expression formula is:

[0055] P c =f match (R s ,F t )

[0056] Where, P c :Customer service knowledge base call priority, f match :Customer service matching function, R s : Recommended product score set, F t : Active feedback trigger results.

[0057] The customer service system can prioritize the relevant knowledge base content based on recommended links and user feedback, quickly respond to user inquiries, and improve customer service response efficiency.

[0058] Linking recommended paths with customer service requests in real time avoids the disconnect between existing customer service systems, where they are unable to quickly understand users' current needs, and enables intelligent linkage of human-computer interaction.

[0059] Preferably, the specific steps of the closed-loop optimization of interactive data in S9 are as follows:

[0060] S91, Dynamic Adjustment of Algorithms and Display Priorities: The platform collects interactive data from S1 to S8 in real time, combines interaction effects and user feedback, and incorporates recommendation algorithm weight adjustment logic to dynamically optimize algorithm weights. It also adjusts page display priorities based on data feedback, ensuring that recommended content and page presentation are more aligned with user needs, laying the foundation for subsequent interactive experience optimization.

[0061] The formula for weight adjustment of the recommendation algorithm is:

[0062]

[0063] Where, Optimized recommendation algorithm weight Current recommendation algorithm weight, α: weight adjustment coefficient, U pos : User positive feedback ratio (such as clicks, purchases), U negUser negative feedback ratio (such as exit, close);

[0064] According to the user positive and negative feedback dynamic adjustment recommendation algorithm weight, realize the continuous self-adaptive optimization of recommendation strategy, improve the recommendation accuracy and user satisfaction.

[0065] The closed-loop optimization mechanism can automatically adjust the recommendation parameters based on user feedback, break through the problem of fixed recommendation model parameters and long adjustment period of existing e-commerce platforms, and build a self-learning recommendation system.

[0066] S92, feedback trigger mechanism precision optimization: based on the whole process interaction data, the system analyzes the user feedback situation, dynamically adjusts the feedback trigger threshold, combines the recommendation algorithm weight adjustment idea, makes the feedback trigger such as pop-up prompt and customer inquiry more accurate, reduces invalid interference, and continuously improves the smoothness of human-computer interaction and user satisfaction.

[0067] The beneficial effects of the present application are as follows:

[0068] 1、The present application combines multi-modal interactive input and scene intelligent recognition, so that users can search and purchase goods in a more natural way, supports voice input, image recognition, text retrieval and other interactive modes, breaks the dependence of existing e-commerce platforms on single page clicks and keyword input, significantly improves the freedom and convenience of human-computer interaction of e-commerce platforms, meets the interactive needs of different users in different scenarios, and has higher applicability and flexibility.

[0069] 2、The present application dynamically links immersive interactive display and real-time personalized recommendation, can actively push product information related to the current interest during user browsing, realizes scene-based and personalized product display and accurate recommendation, significantly improves user shopping decision efficiency and decision experience, further enhances user stickiness of e-commerce platforms through continuously optimized recommendation path and visual presentation, and breaks through the bottleneck of single e-commerce human-computer interaction experience and inaccurate recommendation.

[0070] 3、The present application establishes a high-response and high-adaptation human-computer interaction service system through active interactive feedback, intelligent customer service linkage and interactive data closed-loop optimization, effectively solves the technical problems of slow response, fragmented customer service system and user feedback lag of traditional e-commerce platforms, significantly improves the human-computer interaction efficiency, intelligent level and service quality of e-commerce platforms through real-time interaction and continuous optimization, and promotes the in-depth development of e-commerce platforms towards intelligence and personalization. BRIEF DESCRIPTION OF DRAWINGS

[0071] Figure 1 The present application is based on the flow chart of the human-computer interaction method of the e-commerce platform. DETAILED DESCRIPTION

[0072] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work belong to the scope of protection of the present application.

[0073] As shown in the figure, the embodiments of the present application provide a human-computer interaction method based on an e-commerce platform, and the specific steps of the human-computer interaction method based on the e-commerce platform are as follows: Figure 1

[0074] S1, user multi-source demand information collection: through collecting multi-source input data of the user on the e-commerce platform, including text search words, voice instructions, image uploading, browsing track, historical purchase record, a user current behavior data set is constructed as a basis input for subsequent data processing;

[0075] S2, multi-modal data analysis: modal recognition and data analysis are performed on the multi-source input data of the user multi-source demand information collection, the specific types and contents of voice input, image input and text input are distinguished, and a unified data label of multi-modal input is established;

[0076] S3, user behavior label generation: based on the input data of the multi-modal data analysis, a preset user behavior modeling algorithm is called to generate multi-dimensional user behavior labels including interest preference, browsing habit, price sensitivity, etc., which are used for accurate matching of subsequent scenes and recommendations;

[0077] S4, user scene intelligent recognition: through the user behavior label generated by the user behavior label, a preset scene model library is matched, the current consumption scene (such as home, office, baby, etc.) of the user is automatically recognized, and the scene identification parameter is outputted, which is used as the basis for subsequent display and recommendation;

[0078] S5, immersive interaction display: based on the scene identification parameter of the user scene intelligent recognition, the layout structure, information level and display mode of the current platform page are dynamically adjusted, the text, video or 3D product model is loaded according to the scene priority, and the immersive, multi-dimensional interaction experience is realized;

[0079] S6, real-time personalized recommendation: according to the user click path and stay node of the immersive interaction display page content, combined with the user behavior label generation, the real-time recommendation algorithm is used to dynamically generate personalized product recommendation results, and the current page is pushed and updated in real time;

[0080] ​S7, active interactive feedback: when the user stays at a certain node of the real-time personalized recommendation page for more than a preset time or frequently switches the browsing path, a pop-up prompt, preferential information or customer service inquiry portal based on the user scenario and the recommended content is actively triggered to realize timely interactive feedback;

[0081] S8, intelligent customer service linkage: when the user requests customer service through the interface of active interactive feedback, the system synchronizes the real-time personalized recommendation link and the active interactive feedback information of the user to the intelligent customer service system in real time, and the intelligent customer service automatically calls the corresponding knowledge base content based on the current recommendation link node to respond quickly;

[0082] S9, interactive data closed loop optimization: the platform collects the whole process interactive data of S1 to S8 in real time, dynamically adjusts the recommendation algorithm weight, page display priority and feedback trigger threshold based on the interactive effect and user feedback, and continuously optimizes the human-computer interaction experience of subsequent users.

[0083] Through the fusion of multi-modal interaction and scene intelligent recognition, an interactive system supporting multi-element input such as voice, image and text is constructed, breaking through the dependence of traditional e-commerce on single search mode, enabling users to initiate commodity inquiry and purchase request in a more natural way, significantly improving the interaction freedom and input convenience, and adapting to the use demand in different scenes. At the same time, with the dynamic cooperation of immersive interactive display and real-time personalized recommendation, scene-based commodity information is actively pushed according to the user browsing behavior, the shopping decision efficiency is optimized through accurate recommendation, the user stickiness is enhanced through dynamic visual presentation, and the problem of inaccurate traditional recommendation is solved. In addition, through the active interactive feedback mechanism, intelligent customer service linkage system and whole process data closed loop optimization, a high-response service link is constructed, the user demand is responded in real time and the interactive strategy is continuously iterated, effectively improving the problems of traditional e-commerce such as response lag and customer service fragmentation, and promoting the deep evolution of the platform to intelligent and personalized service.

[0084] Among them, the user demand collection in S1 refers to the construction of a three-dimensional capture system, active input level, through Nginx log analysis to grab text search words, combined with semantic analysis engine to predict intention; the voice instruction is converted by the edge node ASR engine, and the voiceprint recognition is anchored to the user; after the image data is cached by CDN, the CNN model is used to extract color, texture and other feature vectors.

[0085] For passive data collection, browsing trajectories are reported in real time with the help of front-end embedded SDK, combined with browser fingerprint identification devices; historical purchase records are synchronized incrementally from distributed databases through Binlog, and order status restoration processes are associated. After all data are converged through Kafka data stream, standardized processing is performed through the ETL framework according to user ID and timestamp, forming a raw behavior set. In the data fusion stage, isolated forest algorithm is used to clean noise, knowledge graph technology is used to align cross-source user identification, and time sequence network is used to map multi-source data to a unified time axis according to event semantics, and finally output a fusion data set containing structured interaction parameters, semi-structured word vectors and unstructured multimedia features.

[0086] Among them, the multi-modal data analysis in S2 means that the input type recognition and content analysis need to build a hierarchical processing architecture: the voice input is transcribed in real time through the streaming ASR engine of the edge node (such as the DeepSpeech model based on Transformer), and the semantic understanding network with attention mechanism (such as BERT-whitening) is used to analyze the instruction intent and extract the voiceprint features for identity verification; the image input is cached through CDN, and then the multi-scale feature extraction network (such as ResNet+FPN architecture) is used to identify the bottom features such as item outline and color texture, and then the YOLOv8 target detection model is used to locate the semantic entity; the text input uses the syntax analyzer (such as spaCy) to extract noun phrases, and combines the TF-IDF weight algorithm and the domain dictionary to generate keyword vectors.

[0087] In the unified label fusion stage, first, the knowledge graph embedding model (such as TransR) is used to map the features of each modality to a shared semantic space, and a label system containing entities-attributes-relations is constructed; then, the graph database (such as Neo4j) is used to establish cross-modal association index, for example, the "red dress" label recognized by image is aligned with the text search word "wine red dress" through the synonym dictionary, and the "buy" in the voice instruction corresponds to the "purchase intent" node in the label library. Finally, the multi-modal feature encoder (such as the CLIP model) is used to generate a fusion vector, realizing the spatio-temporal alignment and semantic association of text semantics, image vision and voice rhythm features;

[0088] Among them, the user behavior label generation in S3 means that the system calls a hybrid modeling algorithm to derive labels based on the fusion features obtained by multi-modal analysis (such as cross-modal vectors generated by the CLIP model). Specifically, the interest preference label is modeled by a Transformer encoder for text keywords and image semantic features in time sequence, and the hierarchical interest label is generated by combining the association relationship of commodity categories in the knowledge graph; the browsing habit label relies on the LSTM neural network to analyze the behavior sequence such as page stay time and jump path, and identifies the periodic browsing mode through the hidden Markov model; the price sensitivity label extracts features such as single price and promotion response from historical order data, and maps them to gradientized price interval labels after XGBoost model training.

[0089] When the algorithm is executed, first, the multi-modal features are filtered for redundant dimensions through a feature selection layer (such as LightGBM's GOSS sampling), and then cosine similarity matching is performed with the historical behavior patterns (based on Faiss vector retrieval construction) in the distributed cache. The latent demand reasoning module combines the user portrait entity relationship (stored in the Neo4j graph database) to derive implicit labels through a rule engine, for example, associating the "outdoor hiking shoes" search term with historical camping equipment purchase records to generate a "outdoor sports enthusiast" composite label. The final output label system includes three categories of basic attribute labels, behavior sequence labels, and context association labels, and the label confidence score mechanism ensures the accuracy of the description;

[0090] Among them, the user scene intelligent recognition in S4 means that the system relies on a distributed vector database (such as Milvus) to build a scene model library, each scene model contains a set of feature vectors trained from historical data (such as the "commuting shopping" scene integrates commuting period browsing records, portable goods click features, and workday delivery address preferences). When matching, first map the behavior label to a high-dimensional semantic vector through a feature encoder (such as BERT-whitening), and then use a double-tower neural network to calculate the cosine similarity with each scene model, where the tower structure corresponds to the nonlinear transformation of the user label features and the scene template features, respectively.

[0091] Similarity comparison integrates spatio-temporal context calibration logic: the time dimension analyzes the time entropy of the behavior sequence through a sliding window algorithm (such as Flink-based real-time window), identifying characteristics of early morning peak, late night, etc.; the spatial dimension combines IP address resolution geofencing data (such as shopping mall, residential area labels) with device positioning information for scene anchoring. When the similarity calculation result exceeds the dynamic threshold (adjusted dynamically by the historical scene transition matrix), the scene association reasoning of the graph database (Neo4j) is triggered, for example, after associating "gymnasium surrounding browsing" with "sports clothing purchase label", the "gymnasium after shopping" scene identifier is completed through the rule engine. The final output scene parameters include the main scene category, sub-scene dimension (such as emergency level, social attribute) and confidence score, packaged in JSON format for downstream recommendation system call;

[0092] Among them, the immersive interaction display in S5 refers to the dynamic adaptation of page layout relying on scene identification parameters to drive the front-end rendering engine (such as React+CSSGrid), which dynamically adjusts the DOM structure through scene weight configuration file (JSON format storage element priority). For example, the "commuting shopping" scene will increase the Z-axis level of portable product cards, and use Flexbox layout to compress non-core information area; the "holiday promotion" scene activates the parallax scrolling effect, anchoring the discount label to the golden position of the viewport. The scene display weight logic is parsed by the rule engine, combined with device screen size (real-time detected by window.matchMedia) and interaction history (such as touch hot area heat map) to generate a dynamic layout scheme.

[0093] Multi-form content loading link, through scene-content mapping table (stored in Redis cache) to determine resource priority: 3D exercise equipment model (Three.js rendering) is preferentially preloaded in fitness scene, with WebGL material for real-time preview; video browsing loading (IntersectionObserver monitors the visible area) is enabled in makeup scene, with color matching text set pre-fetching. The weight allocation algorithm combines content volume (such as WebP format picture compression ratio) and user device performance (GPU rendering capability detection) to schedule resource loading order through priority queue (PriorityQueue data structure), realizing parallel pre-fetching of WebVTT subtitles and 3D models, and finally optimizing the secondary access loading link through ServiceWorker cache strategy;

[0094] In S6, real-time personalized recommendation refers to the system capturing click path and stay node data in real time through front-end SDK (such as WebRTC-based interactive tracking component), pushing to message middleware (Kafka) through WebSocket long connection, and analyzing behavior sequence by Flink stream processing engine with microsecond-level granularity. At the same time, user behavior label vector is retrieved from distributed cache (Redis), and real-time features are combined through double-tower neural network (Query tower + Item tower). The Query tower integrates historical behavior timing features extracted by LSTM and attention weights of current interaction, and the Item tower generates structured semantic representation based on product knowledge graph.

[0095] The recommendation algorithm layer adopts a hybrid strategy: short-term preference is calculated through SlidingWindow to obtain click heat map weight of current interaction, combined with SimRank algorithm to mine product association of similar behavior paths; long-term features are matched with historical label clustering centers through Faiss vector retrieval. The dynamic ranking module uses XGBoost model to fuse spatio-temporal features (such as scene preference of current period, rendering performance of device side), generates real-time recommendation list, and pushes to front-end through ServiceWorker, triggering React component's incremental rendering (Hydration), updating the position and display priority of product cards while maintaining page context.

[0096] In S7, active interaction feedback refers to the system monitoring element visible area stay duration in real time through front-end SDK (such as IntersectionObserver API), combined with browser event listening (MutationObserver) to capture browsing path switching frequency. When detecting stay timeout (such as exceeding 3 seconds through requestAnimationFrame timing) or high-frequency jumping (based on sliding window algorithm to identify more than 5 times of page block switching within 10 seconds), behavior data is pushed to Flink stream processing engine through WebSocket, combined with scene identification and recommendation content vector in Redis cache for joint analysis.

[0097] The feedback strategy generation module links rule engine and natural language generation model: the rule engine analyzes preset strategies (such as "triggering color test video popup in makeup scene when stay duration exceeds threshold"), and the NLG model dynamically generates promotional copy (combined with promotion labels in product knowledge graph). The interactive component renders the popup (including CSS transition animation) through WebComponents technology, the customer service entry calls WebSocket long connection to connect to the IM system, and communicates with the recommendation engine through the PostMessage mechanism to dynamically adjust the display weight of subsequent pushed content, realizing real-time response link based on behavior features.

[0098] In S8, intelligent customer service linkage triggers the linkage process through front-end event monitoring (such as capturing customer service portal click events). This process pushes real-time recommendation link data (including the current product list and interaction heat map) and user feedback information (stop nodes and click paths) to the customer service system message queue (Kafka) via a persistent WebSocket connection. After receiving the data, the customer service system first uses a rule engine to parse the recommended link nodes (such as "Commuting Scenario - Portable Water Cup Recommendation") and generates priority tags based on the user's current interaction characteristics.

[0099] The response priority matching module uses the BERT-whitening semantic model to calculate the semantic similarity between user questions and knowledge base entries, and at the same time retrieves the frequently asked questions and answers associated with the recommended nodes in the Neo4j graph database (such as the product parameter knowledge base corresponding to "how long does a water cup last"). When searching through Elasticsearch, the scene label (such as "commuting") and the recommended product ID are used as filtering conditions, and the top 5 answer templates with the highest correlation are returned first. The customer service response content is dynamically rendered by the natural language generation model (NLG), and the recommended product card (including real-time price tags) is embedded. The WebComponents technology is used to achieve a two-way linkage between rich text messages and recommendation pages to ensure the scene semantic consistency between customer service replies and the current recommendation process;

[0100] The closed-loop optimization of interactive data in S9 refers to the adjustment of algorithms and display priorities. This involves collecting full-link interactive data (such as click timing and dwell heat maps analyzed by the Flink stream processing engine) through a distributed log system, and dynamically optimizing the recommendation algorithm weights through reinforcement learning models (such as DDPG). This involves calculating the cosine similarity between user click feedback and recommendation results in real time, and synchronously updating the interaction layer weights of the dual-tower model through a parameter server (such as Horovod). The display priority is determined by the front-end rendering engine (React+CSSGrid) dynamically adjusting the DOM structure based on A / B test tracking data. For example, the WebVitals indicator is used to monitor user attention distribution and automatically increase the Z-axis level and viewport share of high-conversion modules.

[0101] At the level of feedback trigger mechanism optimization, the system analyzes user historical feedback data through sliding window algorithm (rolling window based on Flink) and dynamically calibrates the trigger condition by using adaptive threshold model (such as Gaussian mixture model) - when it is detected that the pop-up click rate is lower than the average for 3 consecutive periods, the stay time threshold is automatically adjusted (the requestAnimationFrame redesign time parameter is used). The rule engine synchronously associates with the recommended algorithm weight change, for example, after the XGBoost model is updated, the scene-threshold mapping table in Redis cache is real-time called, the gradient threshold adjustment is implemented for high-frequency trigger items such as "promotion pop-up", and the co-evolution of feedback mechanism and recommendation strategy is realized.

[0102] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one entity or action from another, without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0103] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, alternatives, and variations can be made in the embodiments without departing from the spirit and scope of the present application as defined by the appended claims and their equivalents.

Claims

1. An e-commerce platform based human-machine interaction method, characterized in that: The specific steps of the e-commerce platform-based human-computer interaction method are as follows: S1, user multi-source demand information collection: through collecting the multi-source input data of the user on the e-commerce platform, including text search words, voice instructions, image uploading, browsing track, and historical purchase record, a user current behavior data set is constructed as the basis input for subsequent data processing; S2, multi-modal data analysis: modal recognition and data analysis are performed on the multi-source input data of the user multi-source demand information collection, the specific types and contents of voice input, image input, and text input are distinguished, and a unified data label of multi-modal input is established; S3, user behavior label generation: based on the input data of multi-modal data analysis, a preset user behavior modeling algorithm is called to generate multi-dimensional user behavior labels including interest preference, browsing habit, and price sensitivity, which are used for accurate matching of subsequent scenes and recommendations; S4, user scene intelligent identification: through the user behavior label generated by the user behavior label, the preset scene model library is matched, the current consumption scene of the user is automatically identified, and the scene identification parameter is outputted as the basis for subsequent display and recommendation; S5, immersive interaction display: based on the scene identification parameter of the user scene intelligent identification, the layout structure, information level, and display mode of the current platform page are dynamically adjusted, the text, video or 3D product model is loaded according to the scene priority, and the immersive, multi-dimensional interaction experience is realized; S6, real-time personalized recommendation: according to the user click path and stay node of the immersive interaction display page content, combined with the user behavior label generation, the real-time recommendation algorithm is used to dynamically generate personalized product recommendation results, and the current page is pushed and updated in real time; S7, active interaction feedback: when the user stays at a certain node of the real-time personalized recommendation page for more than a preset time or frequently switches the browsing path, the pop-up prompt, preferential information or customer service inquiry entrance based on the user scene and the recommended content is actively triggered to realize timely interaction feedback; S8, intelligent customer service linkage: when the user requests customer service through the interface of active interaction feedback, the system synchronizes the real-time personalized recommendation link and the user active interaction feedback information to the intelligent customer service system in real time, and the intelligent customer service automatically calls the corresponding knowledge base content based on the current recommendation link node for quick response; S9, interactive data closed loop optimization: the platform collects the whole process interaction data of S1 to S8 in real time, dynamically adjusts the recommendation algorithm weight, page display priority, and feedback trigger threshold based on the interaction effect and user feedback, and continuously optimizes the subsequent user human-computer interaction experience. 2.The method of human-computer interaction based on an e-commerce platform according to claim 1, characterized in that: The specific steps of the user demand collection in S1 are as follows: S11, multi-source data collection and behavior set construction: through collecting the active input of the user on the e-commerce platform, such as text search words, voice instructions, and uploaded images, combined with passive behavior data such as browsing track and historical purchase record, the system can integrate a comprehensive user current behavior data set; The user behavior data set construction expression formula is: D u = {T u , V u , I u , B u , P u} In the formula, D u : User full behavior data set, T u : User text input set (such as keywords), V u : User voice input set, I u : User image input set, B u : User historical browsing behavior set, P u : User historical purchase record set; S12, data fusion lays the foundation for processing: the organic fusion of dispersed user multi-source behavior data, the construction of a coherent behavior data set can accurately reflect the user's current demand and behavior characteristics. 3.The method of human-computer interaction based on an e-commerce platform according to claim 1, characterized in that: The specific steps of the S2 multi-modal data analysis are as follows: S21, multi-modal input type recognition and content analysis: for user multi-source demand information, first translate the voice input into semantics, and clarify the instruction intent; feature extraction is carried out for image input, and the attribute of the object is identified; Carry out keyword extraction on text input and grasp the search core; S22, unified label fusion multi-modal feature: after completing the analysis of each type of input, a unified data label system is established for voice, image and text data, different modal information is associated through labels, and multi-modal features are organically fused; The multi-modal feature analysis expression formula is: F m = f parse (D u ) = {F t , F v , F i} In the formula, F m : a multi-modal input feature set, f parse : a multi-modal data analysis function, F t : a text input analysis feature, F v : a speech input analysis feature, F i : an image input analysis feature.

4. The human-computer interaction method based on an e-commerce platform according to claim 1, characterized in that: The S3 user behavior label generation refers to relying on the multi-modal data analysis results, the system calls the preset user behavior modeling algorithm, integrates multi-source information such as voice, image and text, and generates user behavior labels covering interest preferences, browsing habits and price sensitivity. In this process, the algorithm associates multi-modal features with historical behavior patterns, combines with the reasoning logic of potential demand, so that the label accurately reflects the user characteristics, provides a reliable basis for subsequent scene adaptation and recommendation matching, and realizes effective conversion from data analysis to behavior description; The user behavior label generation expression formula is: L u = f label (F m , B u , P u ) In the formula, L u : a set of user behavior labels (such as interests, preferences, price sensitivity), f label : a behavior label generation function, F m : a multi-modal feature set of S2, B u : a user historical browsing set, P u : a user historical purchase set.

5. The human-computer interaction method based on an e-commerce platform according to claim 1, characterized in that: The S4 user scene intelligent recognition refers to generating user behavior labels based on the generated user behavior labels, the system calls the preset scene model library for matching, and calculates the similarity degree of the behavior label and each scene model feature, and automatically identifies the current consumption scene according to the association closeness. This process integrates the similarity comparison logic of scene matching, so that the scene judgment is more accurate, and finally outputs the scene identification parameter, which provides the basis for the adaptation of subsequent display form and recommended content, realizes the efficient conversion from behavior label to scene positioning, and improves the fitting degree of recommendation and user demand; The scene matching similarity expression formula is: In the formula, S i : the matched current user scenario, S: the preset scenario library set, Sim(L u ,S j ): the similarity function between the user label and the scenario label.

6. The human-computer interaction method based on an e-commerce platform according to claim 1, characterized in that: The specific steps of the S5 immersive interactive display are as follows: S51, page layout dynamically adapts to scene: based on the scene identification parameter, the platform dynamically adjusts the page layout structure, information level and display method, integrates the scene display weight distribution logic, arranges elements according to the scene demand priority, and makes the page highly consistent with the current scene, improves the scene identification and information acquisition efficiency of the user during browsing; S52, multi-form content is preferentially loaded: according to the scene identification parameter and the weight distribution logic, preferentially load the adaptive content such as text, video or 3D product model, and create an immersive, multi-dimensional interactive experience by reasonably distributing the display weight of each form of content; The scene display weight distribution expression formula is: W d = f disp (S i ) = {w img , w vid , w 3D} In the formula, W d : a set of page presentation weights, f disp : a scene-based presentation weight distribution function, S i : a user scenario identified by S4, w img : a picture presentation weight, w vid : a video presentation weight, w 3D : a 3D model presentation weight.

7. The human-computer interaction method based on an e-commerce platform according to claim 1, characterized in that: The real-time personalized recommendation in S6 refers to that based on the click path and stay nodes of the user in the immersive interaction page, the system combines these real-time behavior data with the generated user behavior tags, calls the real-time recommendation algorithm to dynamically generate personalized product recommendations, and the algorithm accurately matches the products by analyzing the association between the current interaction preferences and historical behavior characteristics, and immediately updates the push on the current page. 8.The method of human-computer interaction based on an e-commerce platform according to claim 1, characterized in that: The active interaction feedback in S7 refers to that when the user stays at a certain node of the real-time personalized recommendation page for more than the preset time or frequently switches the browsing path, the system will actively start the interaction feedback mechanism, accurately push the pop-up window prompts, related preferential information, or show the customer inquiry entrance combined with the current user scenario and recommended content. 9.The human-computer interaction method based on an e-commerce platform of claim 1, wherein: The intelligent customer service linkage in S8 refers to that when the user requests customer service through the active interaction feedback interface, the system synchronizes the real-time personalized recommendation link and user interaction feedback information to the intelligent customer service system, the customer service system combines the current recommendation link node, integrates the response priority matching logic, and automatically calls the corresponding knowledge base content. The customer response priority matching expression formula is: P c = f match (R s , F t ) In the formula, P c : customer service knowledge base call priority, f match : customer service matching function, R s : recommended product score set, F t : active feedback trigger result.

10. The human-computer interaction method based on an e-commerce platform according to claim 1, characterized in that: The specific steps of the interaction data closed-loop optimization in S9 are as follows: S91, algorithm and display priority dynamic adjustment: the platform collects the whole-process interaction data of S1 to S8 in real time, combines the interaction effect and user feedback, integrates the recommendation algorithm weight adjustment logic, dynamically optimizes the algorithm weight, and adjusts the page display priority according to the data feedback; The recommendation algorithm weight adjustment expression formula is: In the formula, Optimized recommendation algorithm weight Current recommendation algorithm weight, α: weight adjustment coefficient, U pos : User positive feedback ratio (such as clicks, purchases), U neg : User negative feedback ratio (such as exit, close) S92, feedback trigger mechanism precision optimization: based on the whole-process interaction data, the system analyzes the user feedback, dynamically adjusts the feedback trigger threshold, and makes the feedback trigger such as pop-up window prompt and customer inquiry more accurate combined with the recommendation algorithm weight adjustment idea.

Citation Information

Cited By

  • Product detail page generation method based on AI large model

    CN121599743A