AI agent interaction method and system based on semantic analysis
By analyzing the information complexity and repetitive interaction rate of user voice data, filtering out biased voice data, combining pronouncing value and syntactic analysis to determine the probability of true intentions, the problem of inaccurate intention recognition by AI agents under multiple meanings is solved, and more accurate user intention recognition is achieved.
Patent Information
- Application Number
- CN202510689716.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the multiple meanings of language information have an impact on the semantic analysis results of AI agents, resulting in inaccurate identification of user's real intentions and difficulty in interacting with AI agents.
By collecting user voice data and converting it into text data, analyzing the information complexity and repetitive interaction rate of text data, filtering out biased speech data, combining pronouncing habit characteristic values and syntactic analysis, the true intention probability of biased speech data is determined, and re-screening of speech data for AI agent interaction.
It improves the accuracy of AI agents in user real intention recognition, solves the impact of multiple meanings of language information on semantic analysis results, and improves the accuracy of user intention recognition.
Smart Images

Figure CN120260554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic recognition in human-computer dialogue, and particularly to an AI intelligent agent interaction method and system based on semantic analysis. Background Art
[0002] Semantic analysis is to analyze the meaning of language and understand the specific content expressed by the language. AI intelligent agent interaction refers to the interaction process between artificial intelligence intelligent agents or between artificial intelligence intelligent agents and human users. AI intelligent agent interaction based on semantic analysis means that in the process of AI intelligent agent interaction, semantic analysis technology is used to understand the language meaning in the interaction content, so as to interact more accurately.
[0003] In order to realize AI intelligent agent interaction based on semantic analysis, generally, the voice data of users is collected and semantic analysis is performed on the voice data to achieve interaction. However, in natural language, many words and expressions have multiple meanings, and the specific meaning is affected by the context and context, which often affects the semantic analysis results of AI intelligent agents, resulting in inaccurate recognition of the true intentions of users and making the interaction between users more difficult. Summary of the Invention
[0004] The present invention provides an AI intelligent agent interaction method and system based on semantic analysis to solve the problem that the multiple meanings of language information affect the semantic analysis results of AI intelligent agents, resulting in inaccurate recognition of the true intentions of users. The specific technical solutions adopted are as follows: In the first aspect, an embodiment of the present invention provides an AI intelligent agent interaction method based on semantic analysis, and the method includes the following steps: Collect the voice data of the user and convert it into text data, and obtain the collection time of each character in the text data and the time interval between two adjacent voice data of the user; Identify the entities and non-entities in the text data, determine the information complexity of the user's text data according to the number and proportion of entities in the user's text data, determine the repeated interaction rate of each voice data of the user respectively according to the time interval between two adjacent voice data of the user and the duration of the voice data, perform syntactic analysis on the text data corresponding to two adjacent voice data of the user, and screen the deviated voice data of the user according to the syntactic analysis results corresponding to two adjacent voice data of the user, the difference in the number of entities, the repeated interaction rate of the user's voice data, and the information complexity of the text data corresponding to the voice data; Determine the voice habit characteristic value of the user according to the time interval between different texts in the adjacent voice data of the user and the time interval between two adjacent voice data of the user. Combine the syntactic analysis result corresponding to the deviated voice data of the user and the difference between the number of entities in the text data corresponding to the deviated voice data and the text data corresponding to the previous adjacent voice data to determine the true intention probability of the deviated voice data; Rescreen the voice data from the deviated voice data according to the true intention probability of the deviated voice data. Based on the rescreened voice data and the deviated voice data, realize the interaction of the AI intelligent agent based on semantic analysis.
[0005] Furthermore, the method for determining the information complexity of the user's text data is as follows: Calculate the proportion of the number of entities in the user's text data to the total number of all entities and non-entities. The normalized value of the product of the proportion and the number of entities in the user's text data is denoted as the information complexity of the user's text data.
[0006] Furthermore, the method for determining the repeated interaction rate of each voice data of the user is as follows: Denote the time interval between two adjacent voice data of the user as the adjacent time interval of the latter voice data in the two adjacent voice data; Denote the duration of the previous voice data in two adjacent voice data of the user as the previous voice duration of the latter voice data in the two adjacent voice data; Denote the normalized value of the ratio of the previous voice duration of the user's voice data to the adjacent time interval as the repeated interaction rate of the user's voice data.
[0007] Furthermore, the specific method for screening the deviated voice data of the user according to the syntactic analysis result, entity quantity difference, repeated interaction rate of the user's voice data, and information complexity of the text data corresponding to the voice data includes: The syntactic analysis result corresponding to the voice data is the subject, predicate, object, attributive, adverbial, and complement in the text data corresponding to the voice data; Denote the subject, predicate, and object in the text data as intention keywords, and denote the attributive, adverbial, and complement in the text data as modifying keywords. Denote the cumulative sum of the quantity differences of all modifying keywords of all the same intention keywords in the text data corresponding to two adjacent voice data of the user as the modifying emphasis quantity of the latter voice data in the two adjacent voice data of the user; Denote the absolute value of the difference between the number of entities in the text data corresponding to two adjacent voice data of the user as the entity difference quantity of the latter voice data in the two adjacent voice data of the user; Filter the deviated speech data of the user according to the repeated interaction rate, modification and emphasis amount, number of entity differences of the user's speech data, and information complexity of the text data corresponding to the speech data.
[0008] Further, the specific method for filtering the deviated speech data of the user according to the repeated interaction rate, modification and emphasis amount, number of entity differences of the user's speech data, and information complexity of the text data corresponding to the speech data is as follows: Multiply the repeated interaction rate, modification and emphasis amount of the user's speech data by the information complexity of the text data corresponding to the speech data, and denote it as the first product of the user's speech data. Denote the normalized value of the ratio of the first product of the user's speech data to the number of entity differences as the deviation possibility of the user's speech data; When the deviation possibility of the user's speech data is greater than the preset first determination threshold, denote the user's speech data as deviated speech data.
[0009] Further, the determination method of the user's speech habit characteristic value is as follows: Denote the acquisition time interval between two adjacent characters corresponding to the user's speech data as the time difference of the latter character in the two adjacent characters. Denote the average value of the time differences of all characters included in an entity corresponding to the user's speech data as the average time difference of the entity. Denote the variance of the average time differences of all entities corresponding to all the user's speech data as the time difference of adjacent characters of the user; Denote the average value of the time intervals between all adjacent two speech data of the user as the time difference of adjacent speech of the user; Denote the normalized value of the ratio of the time difference of adjacent speech of the user to the time difference of adjacent characters as the speech habit characteristic value of the user.
[0010] Further, the determination method of the true intention probability of the deviated speech data is as follows: Denote the number of sentence components that do not conform to the Chinese grammar order in the syntactic analysis result corresponding to the user's deviated speech data as the grammar deviation number of the user's deviated speech data. The sentence components are the subject, predicate, object, attributive, adverbial, and complement in the text data; Multiply the speech habit characteristic value of the user by the number of entity differences of the deviated speech data, and denote it as the second product of the deviated speech data. Denote the product of the ratio of the second product of the deviated speech data to the grammar deviation number as the third product of the deviated speech data; Determine the true intention probability of the deviated speech data according to the third product of the deviated speech data. The third product of the deviated speech data and the true intention probability of the deviated speech data have a negative correlation.
[0011] Further, the method of re-screening speech data from the deviation speech data according to the true intention probability of the deviation speech data includes the following specific method: When the true intention probability of the deviation speech data is greater than the second determination threshold, update the deviation speech data to speech data; when the true intention probability of the deviation speech data is less than or equal to the second determination threshold, do not update the deviation speech data.
[0012] Further, the method of implementing AI intelligent agent interaction based on semantic analysis according to the re-screened speech data and the deviation speech data includes the following specific method: The AI intelligent agent interacts with the speech data and issues an interaction prompt of "lack the ability to meet the requirements" for the deviation speech data.
[0013] In a second aspect, an embodiment of the present invention further provides an AI intelligent agent interaction system based on semantic analysis, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the method described in any one of the above.
[0014] The beneficial effects of the present invention are as follows: This application takes into account the characteristics that the greater the amount of information corresponding to the speech data, the higher the information load of the speech data, and the greater the possibility of problems occurring during the interaction of the AI intelligent agent. It evaluates the information complexity of the user's speech data to obtain the information complexity of the user's text data. Since when the AI intelligent agent has a deviation in the user intention recognition process, the user will re-express their needs to the AI intelligent agent. Therefore, it evaluates the time interval and semantic similarity between the user's adjacent two speech data to obtain the repeated interaction rate of the user's speech data. Further, in combination with the characteristic that when the AI intelligent agent has a deviation in the user intention recognition process, the user will specifically modify and emphasize specific components when re-stating the intention, which is manifested as adding more attributives, adverbials, complements and other determiners to the same entity in the text data corresponding to the adjacent two speech data, all the user's deviation speech data are screened out. Then, it analyzes the reasons for the deviation of the AI intelligent agent's recognition of the user intention from the deviation speech data, evaluates the user's speech habit, obtains the speech habit characteristic value of the user, and combines the semantic logic and grammatical structure corresponding to the user's speech data to evaluate the possibility that the AI intelligent agent obtains the true intention of the user according to the deviation speech data, and obtains the true intention probability of the deviation speech data. Finally, based on the true intention probability and speech data of the deviation speech data, it realizes AI intelligent agent interaction based on semantic analysis, solves the problem that the multiple meanings of language information affect the semantic analysis result of the AI intelligent agent, resulting in inaccurate recognition of the user's true intention, and improves the accuracy of the AI intelligent agent interaction based on semantic analysis in recognizing the user's true intention. Brief Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0016] Figure 1 It is a schematic flowchart of an AI intelligent agent interaction method based on semantic analysis provided by an embodiment of the present invention; Figure 2 It is a flowchart for obtaining information complexity provided by an embodiment of the present invention. Detailed Embodiments
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0018] Please refer to Figure 1 , which shows a flowchart of an AI intelligent agent interaction method based on semantic analysis provided by an embodiment of the present invention. The method includes the following steps: Step S001, collect the voice data of the user and convert it into text data, and obtain the collection time of each character in the text data and the time interval between two adjacent voice data of the user.
[0019] Use a microphone or other audio input device to collect the voice data emitted by the user, perform preprocessing such as noise reduction and echo cancellation on the collected voice data to improve the recognition accuracy. Use a speech recognition model to convert the voice data into text data. While collecting the voice data and obtaining the text data, collect the collection time corresponding to each character in the text data, and record the time interval between the collection times of the first characters in the text data corresponding to two adjacent voice data emitted by the user as the time interval between two adjacent voice data emitted by the user.
[0020] Among them, in this embodiment, the Transformer model in the speech recognition model is selected to convert speech data into text data. It is a well-known technology to use the Transformer model to convert speech data into text data, so it will not be elaborated here; preprocessing the collected speech data such as noise reduction and echo cancellation is a well-known technology and will not be elaborated here; as other implementation manners, on the basis of achieving the purpose of converting speech data into text data, implementers can adopt other speech recognition models in the prior art, and this application does not make special restrictions.
[0021] So far, the text data of the user, the acquisition time corresponding to each character in the text data, and the time interval between two adjacent speech data of the user are obtained.
[0022] Step S002: Identify entities and non-entities in the text data. Determine the information complexity of the user's text data according to the number and proportion of entities in the user's text data. Determine the repeated interaction rate of each speech data of the user respectively according to the time interval between two adjacent speech data of the user and the duration of the speech data. Perform syntactic analysis on the text data corresponding to two adjacent speech data of the user. Filter the deviated speech data of the user according to the syntactic analysis results corresponding to two adjacent speech data of the user, the difference in the number of entities, the repeated interaction rate of the user's speech data, and the information complexity of the text data corresponding to the speech data.
[0023] The speech data of the user contains a large amount of information, and the amount of information corresponding to different speech data is different. When the amount of information is too large, it will lead to too high an information load of the speech data. Therefore, the greater the amount of information in the speech data, the greater the possibility of problems occurring when the AI agent interacts. Therefore, it is first necessary to analyze the information complexity of the user's speech data.
[0024] The text data corresponding to the user's speech data consists of several keywords. When the frequency of the keywords appears relatively high, it indicates that there are more subjects in the user's speech data, that is, the amount of information in the speech data is larger, the information complexity of the speech data is larger, it is more difficult for the AI agent to interact and recognize the user's intention, and the possibility of deviation in the recognition result of the user's true intention is greater.
[0025] Use named entity recognition NER to obtain entities and non-entities in the user's text data.
[0026] Determine the information complexity of the user's text data according to the number and proportion of entities in the user's text data.
[0027] Calculate the proportion of the number of entities in the user's text data to the total number of all entities and non-entities, and record the normalized value of the product of the proportion and the number of entities in the user's text data as the information complexity of the user's text data.
[0028] Among them, it is a well-known technology to use named entity recognition (NER) to obtain entities and non-entities in text data, which will not be elaborated here; it should be noted that in this embodiment, the Z-Score standard normalization method is used to calculate the normalization value. In the actual application process, implementers can use other methods of existing technologies, such as the maximum-minimum normalization method, sigmoid function, etc., to calculate the normalization value, which is not limited here.
[0029] When the proportion of the number of entities in the user's text data to the total number of all entities and non-entities is larger and the number of entities in the user's text data is more, the information volume of the voice data is larger and the information complexity of the voice data is greater. It is more difficult for the AI intelligent agent to interactively recognize the user's intention, and the possibility of deviation in the recognition result of the user's true intention is greater. At this time, the information complexity of the user's text data is greater.
[0030] The flowchart for obtaining information complexity is as Figure 2 shown.
[0031] Furthermore, analyze whether there is a deviation in the process of the AI intelligent agent recognizing the user's intention. When there is a deviation in the recognition result of the user's true intention, in order to convey their true intention to the AI intelligent agent in a timely manner, the user will immediately re-express their needs to the AI intelligent agent. Therefore, when there is a deviation in the process of the AI intelligent agent recognizing the user's intention, the time interval between two adjacent voice data of the user is smaller, and the semantic similarity is higher.
[0032] According to the time interval between two adjacent voice data of the user and the duration of the voice data, respectively determine the repeated interaction rate of each voice data of the user.
[0033] For two adjacent voice data of the user, the time interval between the latter voice data and the previous adjacent voice data among the two adjacent voice data of the user is recorded as the adjacent time interval of the latter voice data among the two adjacent voice data of the user; the time interval between the acquisition moments corresponding to the first and last characters in the text data corresponding to the previous adjacent voice data of the user is recorded as the previous voice duration of the latter voice data among the two adjacent voice data of the user; the normalization value of the ratio of the previous voice duration of the user's voice data to the adjacent time interval is recorded as the repeated interaction rate of the user's voice data.
[0034] It can be understood that when there is no previous adjacent voice data for the voice data, the voice data is not analyzed.
[0035] When the time interval between two adjacent voice data of the user is shorter and the duration of the user's previous adjacent voice data is shorter, the greater the likelihood that the user's previous adjacent interaction was interrupted, the greater the likelihood that the user has made a repeated interaction, and the greater the likelihood that the AI agent deviates during the user intention recognition process. At this time, the repeated interaction rate of the user's voice data is greater.
[0036] It can be understood that when the user makes a repeated interaction, the likelihood that the user has made a repeated interaction is relatively high. It may also be that the user has a new instruction for the AI agent. In order to make the recognition of the deviation during the user intention recognition process by the AI agent more accurate, the semantics expressed by the user's voice data is further analyzed to determine whether there are differences in the intentions corresponding to the user's two adjacent voice data.
[0037] When the AI agent deviates during the user intention recognition process, the user will specially modify and emphasize specific components when re - stating the intention, which is manifested as adding more attributives, adverbials, complements and other determiners to the same entity in the text data corresponding to the two adjacent voice data.
[0038] Perform syntactic analysis on the text data corresponding to the user's voice data to obtain the subject, predicate, object, attributive, adverbial and complement in the text data. Denote the subject, predicate and object in the text data as intention keywords, and denote the attributive, adverbial and complement in the text data as modifier keywords. Denote the cumulative sum of the quantity differences of all modifier keywords of all the same intention keywords in the text data corresponding to the user's two adjacent voice data as the modification and emphasis amount of the latter voice data in the user's two adjacent voice data.
[0039] When the modification and emphasis amount of the voice data is greater, the more special modification and emphasis is made on the latter voice data in the user's two adjacent voice data, and the greater the likelihood that the AI agent deviates during the process of recognizing the user intention expressed by the voice data.
[0040] Denote the absolute value of the difference in the number of entities in the text data corresponding to the user's two adjacent voice data as the entity difference quantity of the latter voice data in the user's two adjacent voice data.
[0041] When the entity difference quantity of the voice data is smaller, the smaller the semantic difference between the user's two adjacent voice data, and the greater the likelihood that the AI agent deviates during the process of recognizing the user intention expressed by the voice data, causing the user to re - state the intention.
[0042] Determine the deviation possibility of the user's voice data according to the repeated interaction rate, modification and emphasis amount, entity difference quantity of the user's voice data and the information complexity of the text data corresponding to the voice data.
[0043] The product of the repeated interaction rate of the user's voice data, the amount of modification and emphasis, and the information complexity of the text data corresponding to the voice data is denoted as the first product of the user's voice data. The normalized value of the ratio of the first product of the user's voice data to the number of entity differences is denoted as the deviation possibility of the user's voice data.
[0044] When the deviation possibility of the user's voice data is greater, the possibility of deviation in the process of the AI agent identifying the user intention expressed by the voice data is greater.
[0045] Thus, the deviation possibility of the user's voice data is obtained.
[0046] When the deviation possibility of the user's voice data is greater than the first determination threshold, it is determined that there is a deviation in the process of the AI agent identifying the user intention expressed by the voice data, and the user's voice data is denoted as deviation voice data. Among them, the first determination threshold is a preset threshold, and the value of the first determination threshold in this embodiment is 0.5.
[0047] Thus, all the deviation voice data of the user is obtained.
[0048] Step S003: According to the time interval between different characters in the adjacent voice data of the user and the time interval between two adjacent voice data of the user, determine the voice habit characteristic value of the user. Combine the syntactic analysis result corresponding to the deviation voice data of the user and the difference between the number of entities in the text data corresponding to the deviation voice data and the text data corresponding to the previous adjacent voice data to determine the true intention probability of the deviation voice data.
[0049] For the deviation voice data, it is necessary to analyze the reason for the deviation in the user intention recognition by the AI agent.
[0050] When the user's own speaking habit has problems such as unclear sentence breaks and stuttering, the voice recognition of the AI intelligent body may divide multiple sentences into one sentence, or divide one sentence into multiple sentences, resulting in errors in the process of the AI agent recognizing the user intention.
[0051] The time interval between the acquisition times corresponding to two adjacent characters in the voice data of the user is denoted as the time difference of the latter character in the two adjacent characters. The average value of the time differences of all the characters included in an entity in the voice data of the user is denoted as the average time difference of the entity. The variance of the average time differences of all the entities corresponding to all the voice data of the user is denoted as the time difference of adjacent characters of the user. The average value of the time intervals between all adjacent two voice data of the user is denoted as the adjacent voice time difference of the user. The normalized value of the ratio of the adjacent voice time difference of the user to the time difference of adjacent characters is denoted as the voice habit characteristic value of the user.
[0052] When the voice habit characteristic value of the user is larger, the possibility that the user's own speaking habit has problems such as unclear sentence breaks and stuttering is smaller, and the possibility that the user's speaking habit is relatively good is larger. At this time, the influence of the user's own speaking habit on the user intention recognition by the AI agent is smaller.
[0053] Furthermore, when there are problems such as grammar structures in the user's own speaking habit, the semantic logic corresponding to the user's voice data is relatively low, which will affect the user intention recognition by the AI agent. In order to avoid the influence of the user's grammar structure on the user intention recognition by the AI agent, analyze the grammar structure of the user's deviated voice data.
[0054] Compare the syntactic analysis result corresponding to the user's deviated voice data with the subject-predicate-object grammar order in Chinese, and record the number of sentence components that do not conform to the Chinese grammar order in the syntactic analysis result corresponding to the user's deviated voice data as the grammar deviation number of the user's deviated voice data. Among them, the syntactic analysis result is the subject, predicate, object, attributive, adverbial, and complement in the text data, and the sentence component is the subject, predicate, object, attributive, adverbial, and complement in the text data.
[0055] For example: "I drink water" is the order of the sentence components of the subject, predicate, and object, which conforms to the Chinese grammar order, while "Water drinks me" is the order of the sentence components of the object, predicate, and subject, which does not conform to the Chinese grammar order. The number of sentence components that do not conform to the Chinese grammar order in "Water drinks me" is 2, and the sentence components that do not conform to the Chinese grammar order are the subject and the object. Therefore, when "Water drinks me" is the deviated voice data, the grammar deviation number corresponding to the deviated voice data is 2.
[0056] Record the product of the voice habit characteristic value of the user and the entity difference number of the deviated voice data as the second product of the deviated voice data, and record the product of the ratio of the second product of the deviated voice data to the grammar deviation number as the third product of the deviated voice data.
[0057] Determine the true intention probability of the deviated voice data according to the third product of the deviated voice data, and there is a negative correlation between the third product of the deviated voice data and the true intention probability of the deviated voice data.
[0058] It can be understood that the negative correlation relationship in this application refers to the relationship between the independent variable and the dependent variable. The negative correlation relationship means that the dependent variable decreases (increases) as the independent variable increases (decreases), which can be an inverse ratio relationship, a subtraction relationship, etc.
[0059] Preferably, as an embodiment of this application, record the difference between the number 1 and the normalized value of the third product of the deviated voice data as the true intention probability of the deviated voice data.
[0060] In some other embodiments of the present application, the negative value of the third product of the deviated speech data is used as the exponent of the exponential function, and the exponential value is denoted as the true intention probability of the deviated speech data, where the base of the exponential function is the natural constant.
[0061] The greater the true intention probability of the deviated speech data, the greater the possibility that the grammatical structure of the user's deviated speech data is correct. At this time, the influence of the user's grammatical structure on the AI agent's user intention recognition is smaller, and the possibility for the AI agent to obtain the user's true intention based on the deviated speech data is greater.
[0062] Thus, the true intention probability of the deviated speech data is obtained.
[0063] Step S004: Rescreen the speech data from the deviated speech data according to the true intention probability of the deviated speech data, and implement the interaction of the AI agent based on semantic analysis according to the rescreened speech data and the deviated speech data.
[0064] When the true intention probability of the deviated speech data is greater than the second determination threshold, it is determined that the AI agent can obtain the user's true intention by performing user intention recognition based on the deviated speech data, and the deviated speech data is updated to speech data.
[0065] When the true intention probability of the deviated speech data is less than or equal to the second determination threshold, it is determined that the AI agent cannot obtain the user's true intention by performing user intention recognition based on the deviated speech data, and the deviated speech data is not updated.
[0066] Wherein, the second determination threshold is a preset threshold, and in this embodiment, the value of the second determination threshold is 0.5.
[0067] Thus, the deviated speech data and the speech data are updated to obtain the updated speech data.
[0068] The AI agent interacts with the speech data and issues an interaction prompt of "does not have the ability to meet the requirements" for the deviated speech data.
[0069] Thus, the interaction of the AI agent based on semantic analysis is realized.
[0070] Based on the same inventive concept as the above method, an embodiment of the present invention further provides an AI agent interaction system based on semantic analysis, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above methods for the AI agent interaction method based on semantic analysis are implemented.
[0071] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. AI intelligent agent interaction method based on semantic analysis, characterized in that, The method includes the following steps: Collect the user's voice data and convert it into text data, obtain the collection time of each character in the text data and the time interval between two adjacent voice data of the user; Identify entities and non-entities in the text data, determine the information complexity of the user's text data according to the number and proportion of entities in the user's text data, determine the repeated interaction rate of each voice data of the user respectively according to the time interval between two adjacent voice data of the user and the duration of the voice data, perform syntactic analysis on the text data corresponding to two adjacent voice data of the user, and screen the user's deviated voice data according to the syntactic analysis results corresponding to two adjacent voice data of the user, the difference in the number of entities, the repeated interaction rate of the user's voice data, and the information complexity of the text data corresponding to the voice data; Determine the voice habit eigenvalue of the user according to the time interval between different characters in the adjacent voice data of the user and the time interval between two adjacent voice data of the user, and determine the true intention probability of the deviated voice data in combination with the syntactic analysis result corresponding to the user's deviated voice data and the difference in the number of entities between the text data corresponding to the deviated voice data and the text data corresponding to the previous adjacent voice data; Rescreen the voice data from the deviated voice data according to the true intention probability of the deviated voice data, and realize the interaction of the AI intelligent agent based on semantic analysis according to the rescreened voice data and the deviated voice data.
2. The AI intelligent agent interaction method based on semantic analysis according to claim 1, wherein The method for determining the information complexity of the user's text data is as follows: Calculate the proportion of the number of entities in the user's text data to the total number of all entities and non-entities, and record the normalized value of the product of the proportion and the number of entities in the user's text data as the information complexity of the user's text data.
3. The AI intelligent agent interaction method based on semantic analysis according to claim 1, characterized in that, The method for determining the repeated interaction rate of each voice data of the user is as follows: Record the time interval between two adjacent voice data of the user as the adjacent time interval of the latter voice data in the two adjacent voice data; Record the duration of the previous voice data in the two adjacent voice data of the user as the previous voice duration of the latter voice data in the two adjacent voice data; Record the normalized value of the ratio of the previous voice duration of the user's voice data to the adjacent time interval as the repeated interaction rate of the user's voice data.
4. The AI intelligent agent interaction method based on semantic analysis according to claim 1, wherein, The method for screening the user's deviated voice data according to the syntactic analysis results corresponding to two adjacent voice data of the user, the difference in the number of entities, the repeated interaction rate of the user's voice data, and the information complexity of the text data corresponding to the voice data includes the following specific methods: The syntactic analysis result corresponding to the voice data is the subject, predicate, object, attributive, adverbial, and complement in the text data corresponding to the voice data; Record the subject, predicate, and object in the text data as intention keywords, record the attributive, adverbial, and complement in the text data as modifier keywords, and record the cumulative sum of the quantity differences of all modifier keywords of all the same intention keywords in the text data corresponding to two adjacent voice data of the user as the modifier emphasis amount of the latter voice data in the two adjacent voice data of the user; The absolute value of the difference in the number of entities in the text data corresponding to the user's two adjacent voice data is denoted as the entity difference quantity of the latter voice data among the user's two adjacent voice data. Filter the user's deviated voice data based on the repetition interaction rate, modification emphasis quantity, entity difference quantity of the user's voice data, and the information complexity of the text data corresponding to the voice data.
5. The AI intelligent agent interaction method based on semantic analysis according to claim 4, wherein The specific method included in filtering the user's deviated voice data based on the repetition interaction rate, modification emphasis quantity, entity difference quantity of the user's voice data, and the information complexity of the text data corresponding to the voice data is as follows: The product of the repetition interaction rate, modification emphasis quantity of the user's voice data, and the information complexity of the text data corresponding to the voice data is denoted as the first product of the user's voice data, and the normalized value of the ratio of the first product of the user's voice data to the entity difference quantity is denoted as the deviation possibility of the user's voice data. When the deviation possibility of the user's voice data is greater than a preset first determination threshold, the user's voice data is denoted as deviated voice data.
6. The AI intelligent agent interaction method based on semantic analysis according to claim 1, wherein The determination method of the user's voice habit characteristic value is as follows: The acquisition time interval between two adjacent characters corresponding to the user's voice data is denoted as the time difference of the latter character among the two adjacent characters, the average value of the time differences of all the characters included in an entity corresponding to the user's voice data is denoted as the average time difference of the entity, and the variance of the average time differences of all the entities corresponding to all the user's voice data is denoted as the time difference of adjacent characters of the user. The average value of the time intervals between all the user's two adjacent voice data is denoted as the adjacent voice time difference of the user. The normalized value of the ratio of the user's adjacent voice time difference to the time difference of adjacent characters is denoted as the voice habit characteristic value of the user.
7. The AI intelligent agent interaction method based on semantic analysis according to claim 4, characterized in that, The determination method of the true intention probability of the deviated voice data is as follows: The number of sentence components that do not conform to the Chinese grammar order in the syntactic analysis result corresponding to the user's deviated voice data is denoted as the grammar deviation quantity of the user's deviated voice data, and the sentence components are the subject, predicate, object, attributive, adverbial, and complement in the text data. The product of the user's voice habit characteristic value and the entity difference quantity of the deviated voice data is denoted as the second product of the deviated voice data, and the product of the ratio of the second product of the deviated voice data to the grammar deviation quantity is denoted as the third product of the deviated voice data. Determine the true intention probability of the deviated voice data according to the third product of the deviated voice data, and there is a negative correlation between the third product of the deviated voice data and the true intention probability of the deviated voice data.
8. The AI intelligent agent interaction method based on semantic analysis according to claim 1, characterized in that, The specific method included in re-filtering the voice data from the deviated voice data according to the true intention probability of the deviated voice data is as follows: When the true intention probability of the deviated voice data is greater than the second determination threshold, update the deviated voice data to voice data; when the true intention probability of the deviated voice data is less than or equal to the second determination threshold, do not update the deviated voice data.
9. The AI intelligent agent interaction method based on semantic analysis according to claim 1, wherein The specific method included in realizing the AI intelligent agent interaction based on semantic analysis according to the re-filtered voice data and the deviated voice data is as follows: The AI agent interacts with the voice data and issues an interaction prompt of "not having the ability to meet the requirements" for the deviated voice data.
10. An AI intelligent agent interaction system based on semantic analysis, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-9.
Citation Information
Cited By
Semantic analysis-based AI agent question and answer text generation method
CN120596641A