Intelligent question and answer control method based on natural language processing and related equipment
By acquiring scenes and user images, and adjusting the answering strategy using multi-model recognition technology, the problem that the intelligent question-and-answer system cannot meet the needs of different professional users is solved, and personalized and intelligent question-and-answer services are realized.
Patent Information
- Application Number
- CN202510453525.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
The existing intelligent question and answer system is difficult to dynamically adjust the answer strategy based on the user's career background and scenarios, resulting in the answer being unable to meet the personalized needs of users of different professionals.
The use scene images and user images are obtained through the camera module, and the scene interaction pattern recognition model and career judgment model are used for personalized processing. The answer content and speed are adjusted in combination with the information database and facial expression feedback to achieve personalized questions and answers.
It improves the accuracy and pertinence of Q&A, enhances the user experience, and meets the personalized needs of different scenarios and professional users.
Smart Images

Figure CN120336485A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular, to an intelligent question-answering control method and related devices based on natural language processing. Background Art
[0002] In the current era of rapid technological development, intelligent question-answering systems, as an important application branch of artificial intelligence, have gradually become a key tool for people to obtain information and communicate with each other. They have great potential in improving the efficiency of information dissemination, enhancing service quality, etc., and are widely used in many fields such as commercial services, education, and medical care to assist the efficient operation and innovative development of various fields.
[0003] Existing intelligent question-answering systems mainly use a keyword matching-based method to process user questions. The system will match the questions input by users with a preset question-answer library to find corresponding answers. At the same time, some systems will also adjust the answering method according to pre-set fixed rules, such as adjusting the answering speed according to the urgency of the question.
[0004] However, in actual application scenarios, due to the different professional backgrounds of users, even the same question may require different answering methods. Existing intelligent question-answering systems adopt a unified interaction mode and are difficult to dynamically adjust the answering strategy according to user characteristics in specific scenarios, resulting in the system's answers may be difficult to meet the personalized needs of different professional users in different scenarios. Summary of the Invention
[0005] This application provides an intelligent question-answering control method and related devices based on natural language processing to meet the personalized needs of users in different scenarios for obtaining information in various different occupations.
[0006] In a first aspect, the present application provides an intelligent question-and-answer control method based on natural language processing, which is applied to an intelligent question-and-answer device. The method includes: after receiving the question information of the questioner, obtaining the usage scenario image and the overall user image information within a set range through a camera module; inputting the usage scenario image into a preset scenario interaction mode recognition model to obtain the interaction mode in the usage scenario, and the scenario interaction mode recognition model is obtained by training in advance according to multiple different scenario images and corresponding standard interaction modes; inputting the overall user image information into a preset occupation judgment model to determine the occupation information of the user, and the occupation judgment model is established by training in advance according to the dressing characteristics, behavior characteristics of multiple different occupation groups and corresponding occupation labels; performing natural language processing on the question information, and determining the answer information according to the interaction mode and the occupation information in combination with the existing information library, and the answer information at least includes the answer content and the answer speed, and the information library includes the information of the common professional terms and question-and-answer habits of users of different occupations in different interaction modes; after outputting the answer information through the information output module, obtaining the facial image feedback by the questioner in the usage scenario, and performing feature recognition on the facial image to determine the facial expression of the questioner; determining the emotional feedback information of the user on the answer information in combination with the facial expression to adjust the interaction mode; outputting a new answer information for the question information of the next questioner matching the occupation information according to the final interaction mode.
[0007] By adopting the above technical solution, first, the camera module is used to obtain the usage scenario image and the overall user image information. The scenario interaction mode recognition model can accurately identify the interaction mode in the scenario, and the occupation judgment model can determine the user's occupation. Then, through the scenario interaction mode recognition model and the occupation judgment model, combined with the information library, natural language processing is performed on the question information to determine the answer information, realizing the personalization and intelligence of the answers in different scenarios. Combined with some embodiments of the first aspect, in some embodiments, the step of performing natural language processing on the question information and determining the answer information according to the interaction mode and the occupation information in combination with the existing information library specifically includes: performing semantic analysis on the question information to extract keywords and semantic features; retrieving the corresponding answer content from the information library according to the keywords and semantic features; performing personalized adjustment on the retrieved answer content according to the interaction mode and the occupation information, and the personalized adjustment includes adjusting the usage degree of professional terms in the answer content according to the occupation information, adjusting the detail degree and expression mode of the answer content according to the interaction mode, and adjusting the answer speed of the answer content to match the requirements of the usage scenario.
[0008] By adopting the above technical solution, after performing semantic analysis on the question information to extract keywords and features, the answer content is retrieved based on the information library, and personalized adjustment is made according to the interaction mode and professional information. This method fully considers user differences, makes the answer content more in line with the user's knowledge background and scenario needs, avoids the answer being too professional or simple, effectively improves the quality and effectiveness of the answer, and enables users to obtain the required information more efficiently. Combined with some embodiments of the first aspect, in some embodiments, after the step of performing natural language processing on the question information and determining the answer information according to the interaction mode and the professional information in combination with the existing information library, it further includes: obtaining the historical answer information set of the questioner, where the historical answer information set includes the question content, answer content, and corresponding facial expression data of the questioner during the historical question-and-answer process; inputting the historical answer information set into a preset user emotion analysis model to obtain the user emotion type of the questioner, where the user emotion analysis model is used to judge the user emotion type according to the historical question-and-answer behavior characteristics of the questioner; after determining the current emotion feedback information of the questioner, according to the user emotion type and the current emotion feedback information, correct the subsequent answer information.
[0009] By adopting the above technical solution, obtaining the historical answer information set of the questioner and judging the emotion type through the user emotion analysis model, and then correcting the subsequent answer in combination with the current emotion feedback information. This helps the intelligent question-and-answer device keenly perceive the user's emotion changes, timely adjust the answer strategy, enhance the emotional resonance with the user, make the question-and-answer process more user-friendly, further improve the user satisfaction, and make the question-and-answer communication smoother and more effective. Combined with some embodiments of the first aspect, in some embodiments, after the step of obtaining the usage scenario image and the overall user image information within a set range through the camera module after receiving the question information of the questioner, it further includes: if it is detected that the usage scenario image matches the preset product recommendation scenario, then identify whether the appearance characteristics of the questioner meet the requirements of the product recommendation object according to the overall user image information; if they meet, determine the interest information of the questioner during the question-and-answer process, and in combination with the product information in the product library, actively recommend products that meet the interest information.
[0010] By adopting the above technical solution, when it is detected that the usage scenario matches the product recommendation scenario and the user's appearance meets the requirements, recommend products by combining the interest information determined during the question-and-answer process. This realizes precise marketing, improves the accuracy and timeliness of product recommendations, provides valuable product information for users, and at the same time increases sales opportunities for merchants, enhancing the commercial benefits and the user's shopping experience. In some embodiments in combination with some embodiments of the first aspect, if it meets the requirements, after the step of determining the interest information of the questioner according to the question-and-answer process and actively recommending products that meet the interest information in combination with the product information in the product library, it further includes: collecting the voice information of the questioner during the product recommendation process; performing voice feature analysis on the voice information to obtain voice feature parameters; inputting the voice feature parameters into a preset product interest recognition model to obtain the interest degree score of the questioner for the currently recommended product, and the product interest recognition model is established by training with the voice feature parameters of historical users during the product recommendation process and the corresponding actual purchase behavior data; when the interest degree score is higher than a preset degree threshold, triggering a deep recommendation mechanism, and the deep recommendation mechanism includes: screening out associated products related to the currently recommended product from the product library; making a scenario description of the associated products based on the professional information of the questioner; combining the product advantages of the associated products with the professional information of the questioner to determine personalized explanation information, and sending the personalized explanation information to the information output module for output.
[0011] By adopting the above technical solution, collecting voice information during product recommendation, analyzing features to obtain the interest degree score, and triggering the deep recommendation mechanism when the score is high. This further explores the potential needs of users, provides product recommendations that better meet user expectations, enhances the pertinence and attractiveness of recommendations, improves the user's purchase willingness, enhances the product recommendation effect, and promotes sales conversion. In some embodiments in combination with some embodiments of the first aspect, after the step of determining the emotional feedback information of the user for the answer information based on the facial expression to adjust the interaction mode, it further includes: determining the degree of understanding of the questioner for the answer information according to the emotional feedback information; adjusting the professional term situation in the according to the understanding situation.
[0012] By adopting the above technical solution, determining the degree of understanding and adjusting professional terms by analyzing the user's emotional feedback can make the answer more in line with the user's knowledge level, enhance the degree of understanding and interactive experience, and improve the communication efficiency and system adaptability. In some embodiments in combination with some embodiments of the first aspect, after the step of outputting new answer information for the question information of the next questioner matching the professional information according to the final interaction mode, it further includes: after receiving the dissatisfaction information feedback by the questioner for the new answer information, updating and training the scenario interaction mode recognition model according to the dissatisfaction information; using the updated scenario interaction mode recognition model for subsequent scenario interaction mode recognition.
[0013] By adopting the above technical solution, when receiving dissatisfaction information about the new answer information, the training scenario interaction mode recognition model is updated. If the user is dissatisfied with the answer in a certain scenario multiple times, the model can be optimized to improve the interaction mode in that scenario. This enables the model to continuously learn and improve, adapting to different user needs and scenario changes. In a second aspect, the present application provides an intelligent robot, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the intelligent robot to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0014] In a third aspect, the present application provides a computer-readable storage medium, including instructions, which when running on an intelligent robot, enable the intelligent robot to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0015] In a fourth aspect, the present application provides a computer program product, which when running on an intelligent robot, enables the intelligent robot to execute the method described in the first aspect and any possible implementation manner in the first aspect. One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Due to the technical means of acquiring image information based on the camera module and combining multiple models to recognize user characteristics and scenario interaction modes for personalized processing of questions and answers, effectively solving the technical problem in the prior art that the intelligent question-and-answer system lacks scenario awareness and user personalized understanding, resulting in inaccurate and non-conforming answers to requirements, and thus achieving the technical effect of improving the accuracy and pertinence of question and answer and enhancing the user experience in different scenarios and occupations.
[0016] 2. Due to the technical means of performing semantic analysis on the question information and combining user characteristics and scenario patterns to personalized adjust the answer content, effectively solving the technical problem in the prior art that the answer content is single and does not fully consider user differences, and thus achieving the technical effect of improving the quality and effectiveness of the answer and assisting users in obtaining information efficiently.
[0017] 3. Due to the technical means of collecting voice information, analyzing the degree of interest to trigger in-depth recommendations and combining the user's occupation to describe the product in a scenario-based manner, effectively solving the technical problem in the prior art that product recommendations lack depth and pertinence and do not effectively combine the user's occupation scenario, and thus achieving the technical effect of excavating the user's potential needs, improving the pertinence and attractiveness of recommendations, and promoting sales conversion. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic flowchart of an intelligent question - answering control method based on natural language processing in an embodiment of the present application; Figure 2 It is another schematic flowchart of an intelligent question - answering control method based on natural language processing in an embodiment of the present application; Figure 3 It is a schematic structural diagram of an entity device of an intelligent question - answering device in an embodiment of the present application. Detailed implementation manners
[0019] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms "a", "an", "the", "above - mentioned", "said", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items.
[0020] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0021] For ease of understanding, the method provided in this embodiment is described in terms of a process below. Please refer to Figure 1 , which is a schematic flowchart of an intelligent question - answering control method based on natural language processing in an embodiment of the present application.
[0022] S101. After receiving the question information from the questioner, obtain the usage scenario image and the overall user image information within a set range through the camera module; After receiving the question information from the questioner, the intelligent Q&A device immediately activates the camera module. The camera module includes a high-definition camera and related image processing chips. The camera is responsible for capturing light information within a set range and converting it into a digital image signal. The intelligent Q&A device determines the set range of images to be collected through a preset algorithm, and this range can be flexibly adjusted according to the application scenario and functional requirements of the device. For example, in the shopping guide scenario, the set range may cover the questioner and a certain area around him / her to obtain sufficient scene information; while in the office consultation scenario, the set range may be more focused on the questioner himself / herself and relevant items on the desk. To ensure the integrity and accuracy of image collection, the camera module will collect multiple frames of images from different angles and distances. By using the method of multiple cameras working together, different cameras are responsible for collecting images in different areas, and then these images are stitched and fused to form a complete usage scenario image and user overall image information. This method of multiple cameras working together can effectively avoid the lack of image information caused by occlusion or perspective limitations and improve the reliability of image collection.
[0023] S102. Input the usage scenario image into a preset scenario interaction mode recognition model to obtain the interaction mode in the usage scenario. This scenario interaction mode recognition model is obtained by pre-training according to multiple different scenario images and corresponding standard interaction modes; The specific training process of the scene interaction pattern recognition model in the intelligent Q&A device is as follows: Collect a large amount of image data from different scenarios (such as hospitals, schools, shopping malls, offices, etc.) to ensure that various possible usage scenarios are covered. For each scene image, the corresponding standard interaction pattern is labeled. For example, the interaction pattern in the hospital scene is labeled as quiet, orderly, and professional consultation; the school scene is labeled as active, interactive, and knowledge transfer. Then, based on deep learning technology, a convolutional neural network (CNN) model is selected as the basic architecture of the scene interaction pattern recognition model. Preprocess the collected image data to improve the model training effect, including operations such as rotating, flipping, scaling, and cropping the original images to increase the diversity of training data, so that the model can adapt to scene images under various different angles and lighting conditions. For example, for images in the hospital scene, different-angle images of wards, corridors, consulting rooms, etc. can be generated through data augmentation. Adopt the backpropagation algorithm, using the labeled standard interaction pattern as the supervision information, and continuously adjust the model parameters so that the model can accurately predict the corresponding interaction pattern according to the input scene image. Select an appropriate loss function to measure the difference between the model prediction result and the true label. Common loss functions such as the cross-entropy loss function, etc. During the training process, the model will optimize the parameters according to the value of the loss function to gradually reduce the loss value. During the training process, monitor the performance metrics of the model on the validation set, such as accuracy, recall rate, etc., to evaluate the training effect of the model. Adjust the training parameters, such as the learning rate, number of iterations, etc., according to the monitoring results to optimize the model performance. For example, if it is found that the accuracy of the model on the validation set no longer improves, the learning rate may need to be reduced or training may be stopped early to prevent overfitting. Use an independent test set to evaluate multiple trained model versions and select the model with the best performance as the final scene interaction pattern recognition model. Deploy the trained final model to the intelligent Q&A device for actual scene interaction pattern recognition tasks. During the deployment process, it is necessary to ensure that the model can be seamlessly integrated with other components of the device (such as the camera module, information processing module, etc.) to achieve an efficient Q&A control process. When the scene image is input into the trained scene interaction pattern recognition model, the model will perform a series of feature extraction and analysis on the image. First, the convolutional layer performs convolution operations on the image to extract feature maps at different levels. These feature maps contain various element information in the scene, such as object shapes, color distributions, spatial layouts, etc. Then, the pooling layer downsamples the feature maps to further abstract the feature information. Finally, the fully connected layer makes classification predictions based on these abstracted feature information and outputs the interaction pattern in the usage scenario.
[0024] S103. Input the overall user image information into a preset occupation judgment model to determine the user's occupation information. This occupation judgment model is established in advance by training based on the dressing characteristics, behavior characteristics, and corresponding occupation labels of multiple different occupation groups. The occupation judgment model in the intelligent question-answering device is a multi-modal fusion model that integrates the advantages of the convolutional neural network (CNN) and the recurrent neural network (RNN) in deep learning. The CNN part is used to extract the dressing characteristics in the overall user image, such as clothing styles, colors, textures, etc. Through multiple convolutional layers and pooling layers, the model can automatically learn the key feature patterns of different occupation dressings. For example, doctors' white coats usually have specific colors, styles, and pocket layouts, and the model can learn these features and use them for identification. The RNN part is used to analyze the user's behavior characteristics, such as standing postures, walking manners, hand movements, etc. RNN can process time series data and capture the dynamic information of the user's behavior.
[0025] In the training stage, a large amount of image data from different occupation groups was collected. These data cover people of various work scenarios, ages, genders, and ethnicities. For each image sample, the corresponding occupation label was manually annotated, and the description information of the dressing characteristics and behavior characteristics was recorded in detail. During the training process, the model first extracts the dressing characteristics in the image through the CNN to obtain feature vectors, and then the RNN processes the extracted feature vectors and the corresponding description information of the behavior characteristics to learn the correlation between the dressing characteristics, behavior characteristics, and occupation labels.
[0026] After the overall user image information is input into the trained occupation judgment model, the model first preprocesses the image, including operations such as cropping, scaling, and normalization, to ensure that the image size and pixel values meet the input requirements of the model. Then, the CNN part extracts features from the image to generate dressing feature vectors. The RNN part analyzes the user's behavior based on the dressing feature vectors and the preset behavior feature templates. The model will match and compare the extracted features with the feature patterns of different occupations learned during the training process, and determine the most likely occupation information by calculating the similarity scores between the features.
[0027] It should be noted that in today's society, the diversity of occupations has brought different dress codes, which to a certain extent reflect the occupational characteristics. The occupational judgment model in this application is designed based on this realistic situation, mainly targeting those people with clear occupational dress codes in daily life, such as flight attendants, doctors, nurses, food delivery workers, etc., and other occupations associated with dress are not limited here. By analyzing and learning the dress characteristics of those people with occupational dress codes in daily life, the occupational judgment model in this application can effectively identify different occupations and provide important user occupational information basis for subsequent intelligent question-answering services.
[0028] S104. Perform natural language processing on the question information, and determine the answer information according to the interaction mode and the occupational information in combination with the existing information database. The answer information includes at least the answer content and the answer speed. The information database includes information on the common professional terms and question-and-answer habits of users in different occupations under different interaction modes. The intelligent question-answering device will first perform lexical analysis and syntactic analysis on the received question information, and extract linguistic features such as the keywords and question patterns of the question. Then, based on the syntactic structure, analyze the semantics of the question, determine the intention of the question and the semantic classification of the required answer. In this process, a semantic network knowledge base will be used to judge the semantic type corresponding to the question information. According to the interaction mode determined in the previous steps, the intelligent question-answering device will query the interaction mode knowledge base to obtain information such as the corresponding answer templates and speech styles under this interaction mode. For example, the formal interaction mode will correspond to more rigorous answer templates, and the chatting interaction mode will correspond to more casual speech styles. According to the occupational information of the questioner determined previously, the intelligent question-answering device will query the occupational feature knowledge base to obtain information such as the required amount of professional vocabulary and the degree of detail requirements for explanations for this occupation. For example, when facing a scholar's question, more professional terms and detailed explanations are needed. The intelligent question-answering device will combine the semantic content of the question, the interaction mode information, and the occupational information with the corpus knowledge in the information database, and generate an answer text that conforms to the language expression logic and norms according to grammar and pragmatics rules. Finally, based on the generated answer content and answer speed, determine the personalized answer information that meets the requirements of the interaction mode and occupational characteristics, and finally send the answer information to the speech synthesis module and the display module, and output it in two ways: voice and screen display.
[0029] In some embodiments, after this step, the intelligent Q&A device continuously records and stores the historical Q&A data of the questioner during multiple interactions with the questioner. These data include the specific question content at each question, the answer content given by the device in response to the question, and the facial expression data of the questioner after receiving the answer (captured by the camera module, and the specific expression feature recognition method is as described above). By accumulating these historical answer information sets, the device can establish a database of the questioner's Q&A behavior patterns and emotional response patterns, providing a rich data basis for subsequent in-depth analysis. It is constructed based on the Q&A behavior data of a large number of historical users. During the training phase, the question content, answer content, and corresponding facial expression data of many users in different scenarios and for different questions are collected, and the corresponding emotion types are manually labeled to construct a preset user emotion analysis model. When the historical answer information set of the questioner is input, the model first performs semantic analysis on the question content to understand the key points of interest and the nature of the question of the questioner. For the answer content, the model evaluates its relevance, accuracy, and detail to the question. At the same time, combining the corresponding facial expression data, analyzing the category and intensity of the expression, and integrating this multi-dimensional information, the model uses machine learning algorithms for complex calculations and pattern matching, and finally outputs a result representing the common emotion type of the questioner. After obtaining the emotion feedback information determined by the facial expression of the current questioner for a certain answer, it is combined with the user emotion type of this questioner obtained through the user emotion analysis model before (for example, this user usually shows a confused emotion when facing more professional terms), and a comprehensive judgment is made. According to the comprehensive judgment result, the intelligent Q&A device adjusts the generation strategy of subsequent answer information. If it is judged that the user has difficulty in understanding, the use of professional terms may be reduced, the sentence structure may be simplified, and more explanations or examples may be added. In this way, the answer is continuously optimized according to the user's real-time emotion feedback and historical emotion pattern, making the Q&A interaction more in line with the user's needs and improving user satisfaction.
[0030] S105. After outputting the answer information through the information output module, obtain the facial image feedback by the questioner in this usage scenario, and perform feature recognition on the facial image to determine the facial expression of the questioner; After outputting the response information, the intelligent Q&A device will restart the camera module to take a close-up photo of the questioner's face. The face detection algorithm will first identify the questioner's face area in the obtained face image and locate the key facial feature points. Then, the facial expression recognition algorithm will extract the visual features of the facial expression, including the shape of the mouth, the shape of the eyebrows, wrinkle features, etc. These features reflect the movement of facial muscles. By matching the extracted visual features of the facial expression with the facial expression feature dataset, the category to which the expression belongs is determined, that is, the current facial expression of the questioner, such as happy, surprised, angry, etc., is recognized. The facial expression intensity analysis algorithm can also analyze the subtle differences in the facial expression features to judge the specific facial expression intensity, so as to obtain richer facial emotion data. Finally, the intelligent Q&A device can comprehensively consider the facial expression category and intensity to judge the questioner's emotional feedback on the response information. For example, a strong angry expression can be judged as being very dissatisfied with the answer. This facial expression recognition process can use deep learning algorithms to train an accurate facial expression analysis model through a large-scale facial expression dataset, realizing powerful facial emotion recognition and understanding capabilities.
[0031] S106. Determine the emotional feedback information of the user on the response information in combination with the facial expression, and adjust the interaction mode accordingly; After the intelligent Q&A device recognizes the questioner's facial expression, it will determine the emotional feedback information according to a set of pre-set rules for emotion and interaction mode adjustment, and adjust the interaction mode accordingly. If the questioner shows a confused expression (such as frowning tightly, looking confused, etc.), it may mean that the response content is too complex or professional for the user to understand. At this time, the emotional feedback information is that it is difficult to understand the response content, and the interaction mode needs to be adjusted to a more easy-to-understand way, such as reducing the use of professional terms, simplifying the logical structure of the response, etc. If the questioner shows a satisfied expression (such as smiling, nodding, etc.), it indicates that the current interaction mode and response content are more appropriate, and the existing interaction mode can be maintained or fine-tuned according to other factors to further optimize the Q&A experience. In this way, the intelligent Q&A device can dynamically adjust the interaction mode in real time according to the user's emotional feedback, making the Q&A process more in line with the user's needs and expectations and improving the efficiency and quality of communication.
[0032] S107. Output new response information for the question information of the next questioner matching the occupational information according to the final interaction mode.
[0033] After determining the final interaction mode, when receiving a user question that matches the occupational information of the previous questioner again, the intelligent question-and-answer device will generate and output answer information according to the new interaction mode. Suppose that through the analysis of the questioner's facial expression, it was found that the user was confused by the answers with more professional terms, so the interaction mode was adjusted to a more popular expression. Then the next time a questioner with the same occupation (for example, both are doctors) asks a similar question, the intelligent question-and-answer device will reduce the use of professional terms in the answer content and use more easy-to-understand vocabulary to explain related concepts. At the same time, according to the new interaction mode, other aspects such as the speed of the answer and the language style may also be adjusted.
[0034] If the new interaction mode requires faster answers, the device will speed up the pace of information processing and answer output; if a more vivid expression is required, the answer content will add some metaphors, examples and other elements to make the answer easier to understand and accept. This ensures that users of different professions can get answers that better meet their needs and understanding abilities in different scenarios, realize personalized and intelligent question-and-answer services, and improve user experience and satisfaction.
[0035] In some embodiments, when the intelligent question-and-answer device receives information about dissatisfaction from the questioner regarding new answer information, it conducts an in-depth analysis of the dissatisfaction information, including identifying the specific reasons for the dissatisfaction, determining the direction of adjusting the scene interaction pattern recognition model based on the analysis results of the dissatisfaction information, collecting new data related to the current dissatisfaction situation for updating and training the model, and then using the prepared updated training data to retrain the scene interaction pattern recognition model using a suitable machine learning algorithm.
[0036] In the embodiments of the present application, due to the use of multi-dimensional information collection combined with multi-model precise identification, personalized question and answer based on natural language processing and information database, and the technical means of adjusting the interaction mode by using facial expression feedback, it is possible to fully perceive user characteristics and scene information, provide personalized answers that are highly in line with user needs, and provide accurate, intelligent, and personalized question and answer services for users of different professions in different scenarios. After combining the above content, the following is a more detailed description of the process of the method provided by this implementation. Figure 2 , which is another flow chart of the intelligent question-answering control method based on natural language processing in an embodiment of the present application.
[0037] S201, if it is detected that the usage scenario image matches the preset product recommendation scenario, identifying whether the appearance features of the questioner meet the product recommendation object requirements based on the overall image information of the user; After obtaining the usage scenario image, the intelligent Q&A device will compare and analyze it with various preset product recommendation scenario images. For example, the preset product recommendation scenarios may include the scenario of a cosmetics counter in a mall, the scenario of an electronics display area, the scenario of a bookstore, etc. When the usage scenario image matches one of the preset product recommendation scenarios, the appearance feature recognition process based on the user's overall image information will be started. Using image recognition technology, appearance features such as age, gender, facial features (characteristics of facial features, skin condition, etc.), hairstyle, clothing style, etc. are extracted from the user's overall image. Then, according to the characteristics of the target customer group of the current product recommendation scenario, it is judged whether the appearance features of the questioner meet the requirements of the product recommendation object. Taking the cosmetics recommendation scenario as an example, if the target customer group is mainly young women, then when the questioner is recognized as having the appearance features of a young woman, it is considered to meet the requirements; if it is recognized as having the appearance features of an elderly man, it may not meet the requirements of the cosmetics recommendation object in this scenario.
[0038] S202. If it meets the requirements, then according to the interest information of the questioner determined during the Q&A process and combined with the product information in the product library, actively recommend products that meet this interest information; After determining that the appearance features of the questioner meet the requirements of the product recommendation object, the intelligent Q&A device will analyze the language content of the questioner during the Q&A process, extract information such as keywords and topics related to interests, so as to determine the questioner's interest points. For example, in the cosmetics recommendation scenario, if the questioner asks about cosmetics with moisturizing effects or mentions problems such as dry skin, then it can be determined that they are interested in moisturizing cosmetics. At the same time, the device will access the product library, which stores rich product information, including detailed data such as the names, functions, features, applicable populations, prices, etc. of various products. Combining the determined interest information of the questioner, products that match it are screened out from the product library, and this product information is sorted into recommended content and actively pushed to the questioner by means of voice or screen display, etc., providing targeted product recommendation services for users, increasing the chance for users to discover their favorite products, and also helping to improve the success rate of product sales.
[0039] S203. Collect the voice information of the questioner during the product recommendation process; During the process of recommending products to the questioner, the intelligent Q&A device will start the voice collection function and use the built-in microphone or other audio collection devices to record the voice of the questioner in real time. The collected voice information contains the questioner's feedback on the recommended products, questions, further needs expressions, etc.
[0040] S204. Conduct voice feature analysis on this voice information to obtain voice feature parameters; After collecting the voice information, the intelligent Q&A device will analyze it using voice signal processing technology. First, preprocess the voice, including operations such as noise reduction, removal of silent segments, and audio format conversion, to improve the quality and analyzability of the voice signal. Then, extract various voice feature parameters, such as formant frequency, voice duration, speech rate, volume, intonation changes, etc.
[0041] S205. Input the voice feature parameters into a preset product interest recognition model to obtain the interest degree score of the questioner for the currently recommended product. The product interest recognition model is established by training with the voice feature parameters of historical users and their corresponding actual purchase behavior data during the product recommendation process; The intelligent Q&A device inputs the voice feature parameters (such as formant frequency, voice duration, speech rate, volume, intonation changes, etc.) extracted in S204 into a pre-trained product interest recognition model. This product interest recognition model is trained and constructed based on the voice feature parameters of a large number of historical users during the product recommendation process and their subsequent actual purchase behavior data. The model learns the association patterns between the voice feature parameters and purchase behaviors in these historical data, so as to be able to predict the interest degree of the questioner for the currently recommended product according to the newly input voice feature parameters. If the historical data shows that when users ask questions about a product with a fast speech rate, a high intonation, and mention more questions about the product's advantages, the probability of purchasing the product subsequently is relatively high. Then when the model encounters a new user with similar voice feature parameters, it will give a higher interest degree score. The model will perform a series of complex calculations and analyses based on the input voice feature parameters, and finally output a score representing the interest degree of the questioner for the currently recommended product. This score is usually a numerical value, such as a decimal between 0 and 1, where 0 means completely uninterested and 1 means very interested.
[0042] S206. When the interest degree score is higher than a preset degree threshold, trigger the deep recommendation mechanism, which includes: screening out associated products related to the currently recommended product from the product library; The intelligent Q&A device will preset an interest degree threshold. When the interest degree score of the questioner for the currently recommended product obtained from the product interest recognition model is higher than this threshold, it is considered that the questioner has a high interest in the product. At this time, trigger the deep recommendation mechanism. In the deep recommendation mechanism, first, screen out associated products related to the currently recommended product from the product library. The product library stores rich product information, including the association relationships between products. The device will find other related products from the product library according to these preset association relationships and the information such as the category and characteristics of the currently recommended product, to prepare for subsequent personalized recommendations.
[0043] S207. Describe the associated product in a scenario-based manner based on the occupation information of the questioner; The intelligent question-answering device can pre-establish a knowledge base containing various typical work scenarios, demand characteristics of different occupations, and associations with common products. According to the occupation information of the questioner, relevant scenario features are extracted from the knowledge base. For the selected associated products, the matching points between their functions, characteristics, and the extracted occupation scenario features are analyzed. Based on the matching results, using natural language generation technology, the associated products are integrated into the questioner's occupation scenario for description.
[0044] S208. Combine the product advantages of the associated product with the occupation information of the questioner to determine personalized explanation information, and send the personalized explanation information to the information output module for output.
[0045] The intelligent question-answering device deeply analyzes the various advantages of the associated product, such as performance advantages (high speed, high precision, etc.), function advantages (multiple functions in one, unique function design, etc.), and quality advantages (durability, stability, etc.). Combining with the occupation information of the questioner, the specific value of these advantages in their occupation scenario is determined. According to different occupations and product types, a series of personalized explanation templates are designed, which contain a general structural framework. The specific advantages of the associated product and the value points related to the questioner's occupation are filled into the personalized explanation templates to generate complete personalized explanation information. Then, the generated personalized explanation information is sent to the information output module. The information output module can, according to the device settings and user preferences, choose to present the explanation information to the questioner in ways such as voice playback, text display on the screen, or illustrated display, ensuring that the questioner can clearly and intuitively receive the personalized recommendation content about the associated product.
[0046] In the embodiments of the present application, due to the technical means of combining image recognition based on the usage scenario and user appearance feature analysis to accurately locate the product recommendation object, deeply mining the user's interests in the question-and-answer process, quantifying the user's interest in the recommended product through voice feature analysis, and then triggering the deep personalized recommendation mechanism, it is possible to provide highly demand-fitting product recommendation services for the characteristics of different occupational users in different scenarios, effectively solving the problems of traditional product recommendations lacking in-depth understanding of user needs and not fully combining with the user's occupational scenario. Furthermore, it realizes the technical effects of accurately mining the user's potential needs, significantly improving the recommendation pertinence and attractiveness, and thus effectively promoting sales conversion. The intelligent question-answering device in the embodiments of the present invention application will be described from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of the intelligent question-answering device in the embodiments of the present application.
[0047] It should be noted that, Figure 3The structure of the illustrated intelligent question-answering device is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0048] As Figure 3 shown, the intelligent question-answering device includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage section 308 into a random access memory (RAM) 303, such as executing the method described in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0049] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, a button switch, etc.; an output section 307 including a liquid crystal display (LCD), an audio output device, an indicator light, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as required. A removable medium 311, such as a magnetic disk, an optical disc, a magneto-optical disc, a semiconductor memory, etc., is installed on the drive 310 as required so that a computer program read from it can be installed into the storage section 308 as required.
[0050] Specifically, according to the embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the present invention are executed.
[0051] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus.
[0052] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings.
[0053] Specifically, the intelligent question-answering device of this embodiment includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the intelligent question-answering control method provided in the above embodiment is implemented.
[0054] On the other hand, the present invention also provides a computer-readable storage medium, which may be included in the intelligent question-answering device described in the above embodiment; or it may exist separately and not be assembled into the intelligent question-answering device. The above storage medium carries one or more computer programs. When the one or more computer programs are executed by a processor of the intelligent question-answering device, the intelligent question-answering device implements the intelligent question-answering control method based on natural language processing provided in the above embodiment.
[0055] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.
[0056] As used in the foregoing embodiments, depending on the context, the term "when" may be construed to mean "if", "after", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "upon determining" or "if (the stated condition or event) is detected" may be construed to mean "if determined", "in response to determining", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0057] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the foregoing embodiments can be implemented by a computer program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the foregoing method embodiments. The foregoing storage media include: various media such as ROM or random access memory RAM, magnetic disks, or optical discs that can store program code.
Claims
1. An intelligent question-answering control method based on natural language processing, applied to an intelligent question-answering device, characterized in that, The method includes: After receiving the question information of the questioner, obtaining the usage scenario image and the overall user image information within a set range through the camera module; Inputting the usage scenario image into a preset scenario interaction mode recognition model to obtain the interaction mode in the usage scenario, where the scenario interaction mode recognition model is obtained by training in advance according to multiple different scenario images and corresponding standard interaction modes; Inputting the overall user image information into a preset occupation judgment model to determine the occupation information of the user, where the occupation judgment model is established by training in advance according to the dressing characteristics, behavior characteristics of multiple different occupation groups and corresponding occupation labels; Performing natural language processing on the question information, and determining the answer information according to the interaction mode and the occupation information in combination with the existing information database, where the answer information includes at least the answer content and the answer speed, and the information database includes information on the commonly used professional terms and Q&A habits of users in different occupations under different interaction modes; After outputting the answer information through the information output module, obtaining the facial image feedback by the questioner in the usage scenario, and performing feature recognition on the facial image to determine the facial expression of the questioner; Combining the facial expression to determine the emotional feedback information of the user on the answer information to adjust the interaction mode; Outputting a new answer information according to the final interaction mode for the question information of the next questioner matching the occupation information.
2. The method according to claim 1, characterized in that, In the step of performing natural language processing on the question information and determining the answer information according to the interaction mode and the occupation information in combination with the existing information database, it specifically includes: Performing semantic analysis on the question information to extract keywords and semantic features; Retrieving the corresponding answer content from the information database according to the keywords and semantic features; Performing personalized adjustment on the retrieved answer content according to the interaction mode and the occupation information, where the personalized adjustment includes adjusting the usage degree of professional terms in the answer content according to the occupation information, adjusting the detail level and expression mode of the answer content according to the interaction mode, and adjusting the answer speed of the answer content to match the requirements of the usage scenario.
3. The method according to claim 1, wherein After the step of performing natural language processing on the question information and determining the answer information according to the interaction mode and the occupation information in combination with the existing information database, it further includes: Obtaining the historical answer information set of the questioner, where the historical answer information set includes the question content, answer content and corresponding facial expression data of the questioner in the historical Q&A process; Inputting the historical answer information set into a preset user emotion analysis model to obtain the user emotion type of the questioner, where the user emotion analysis model is used to judge the user emotion type according to the historical Q&A behavior characteristics of the questioner; After determining the current emotional feedback information of the questioner, correcting the subsequent answer information according to the user emotion type and the current emotional feedback information.
4. The method according to claim 1, wherein After the step of obtaining the usage scenario image and the overall user image information within a set range through the camera module after receiving the question information of the questioner, it further includes: If it is detected that the usage scenario image matches the preset product recommendation scenario, then according to the overall user image information, identify whether the appearance features of the questioner meet the requirements of the product recommendation target; If it meets the requirements, then according to the interest information of the questioner determined during the Q&A process, and combined with the product information in the product library, actively recommend products that meet the interest information.
5. The method according to claim 4, wherein If it meets the requirements, after the step of actively recommending products that meet the interest information according to the interest information of the questioner determined during the Q&A process and combined with the product information in the product library, it further includes: Collect the voice information of the questioner during the product recommendation process; Conduct voice feature analysis on the voice information to obtain voice feature parameters; Input the voice feature parameters into a preset product interest recognition model to obtain the interest degree score of the questioner for the currently recommended product, and the product interest recognition model is established by training the voice feature parameters and corresponding actual purchase behavior data of historical users during the product recommendation process; When the interest degree score is higher than the preset degree threshold, trigger a deep recommendation mechanism, and the deep recommendation mechanism includes: screening out associated products related to the current recommended product from the product library; Conduct a scenario description of the associated products based on the professional information of the questioner; Combine the product advantages of the associated products with the professional information of the questioner to determine personalized explanation information, and send the personalized explanation information to the information output module for output.
6. The method according to claim 1, characterized in that, After the step of adjusting the interaction mode by combining the facial expression to determine the emotional feedback information of the user for the answer information, it further includes: Determine the degree of understanding of the questioner for the answer information according to the emotional feedback information; Adjust the professional term situation in the according to the understanding situation.
7. The method according to claim 1, wherein After the step of outputting a new answer information for the question information of the next questioner matching the professional information according to the final interaction mode, it further includes: After receiving the dissatisfaction information feedback by the questioner for the new answer information, update and train the scenario interaction mode recognition model according to the dissatisfaction information; Use the updated scenario interaction mode recognition model for subsequent scenario interaction mode recognition.
8. An intelligent robot, characterized in that, The intelligent robot includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the intelligent robot to execute the method according to any one of claims 1-7.
9. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the intelligent robot, enable the intelligent robot to execute the method according to any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product runs on the intelligent robot, enable the intelligent robot to execute the method according to any one of claims 1-7.