Off-line and on-line mixed use method of AI voice

By evaluating the network connection status and voice intention complexity, collaborative decision-making on AI voice processing mode, combining offline and online modes, the problems of waste of resources and poor interaction effects in the existing technology are solved, and efficient and accurate voice interaction services are achieved.

CN120299453APending Publication Date: 2025-07-11SHENZHEN MAICHIRUI SOFTWARE CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510541139.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When existing AI voice technology decides to adopt offline or online mode, it fails to fully consider the complexity of the user's voice intention, resulting in waste of resources or poor interaction effects, and the rational allocation and efficient utilization of resources cannot be achieved.

Method used

By evaluating the user's network connection status and voice intent complexity, collaboratively making voice processing modes, combining the advantages of offline and online modes, the local lightweight intention classification model is used for preliminary classification and the cloud intention knowledge base for fine classification, and dynamically adjusting the processing mode to meet user needs.

Benefits of technology

Provide efficient, accurate and flexible voice interaction services under different network conditions to improve user experience, ensure immediate response and accuracy, and optimize resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299453A_ABST
    Figure CN120299453A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI voice, and particularly discloses an off-line and on-line mixed use method of AI voice, which comprises the following six steps of: evaluating a network connection state to obtain a network reliability index; analyzing voice data input by the user and preamble information, determining the complexity of voice intention and generating a demand priority; and then cooperatively selecting a voice processing mode according to the network reliability index, the demand priority and the decision rule. If the mode is an off-line mode, extracting a voice feature vector, classifying responses by using a local model, obtaining a processing mode confidence coefficient according to user feedback, and if the processing mode confidence coefficient is lower than a threshold value, triggering an on-line supplementary response and updating a rule; and if the mode is an online mode, feature vectors are extracted to be matched with a cloud knowledge base, a fine classification result is obtained, and accurate response is performed by means of a cloud large model. The whole process is combined with network conditions and user requirements, and a processing mode is flexibly decided, so that high-quality AI voice interaction service is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI voice technology, and in particular to an offline and online mixed use method of AI voice. Background Art

[0002] In the era of rapid development of artificial intelligence, AI voice technology has been widely used in many fields such as smart speakers, smart phones, and car systems, which has greatly changed the way people interact with devices and improved the convenience of life and work. AI voice processing usually has two modes: offline and online. The offline mode does not rely on the network, can respond quickly in poor or no network environment, and ensure basic functions; the online mode uses the powerful computing power and rich data of the cloud to achieve complex intention understanding and accurate interaction, but has high network requirements. With the rapid development of industries such as the Internet of Things and smart homes, higher requirements are placed on the stability, accuracy, and efficiency of AI voice interaction in different scenarios. A method that can intelligently and flexibly mix offline and online modes is crucial. It can not only optimize resource utilization and improve user experience, but also promote the widespread application of AI voice technology in more complex and changeable scenarios, and has broad development prospects.

[0003] At present, in conventional AI voice technology applications, when deciding whether to use offline or online mode, most of them rely on simple network status judgments, and fail to fully consider the complexity of the user's voice intentions. Some simple voice interaction tasks, even if the network is good, occupy online resources and cause waste; and complex tasks cannot be met by offline mode alone when the network is unstable. It is impossible to achieve reasonable allocation and efficient use of resources. And once the voice processing mode is determined, it will not be dynamically adjusted according to actual conditions such as user feedback during the processing process. If the initial mode is not selected properly, it may lead to poor effect of the entire voice interaction and cannot be optimized in time to meet user needs.

[0004] Therefore, the present invention proposes an offline and online hybrid method for AI voice. Summary of the invention

[0005] The present invention provides an offline and online mixed use method of AI voice, which effectively combines the advantages of offline and online, and can provide users with efficient, accurate and flexible voice interaction services under different network conditions, thereby improving user experience.

[0006] The present invention provides an offline and online mixed use method of AI voice, comprising:

[0007] S1: Evaluate the current network reliability index based on the user's network connection status;

[0008] S2: Analyze the complexity of the user's voice intention based on the voice data input by the user and the previous context information, and generate demand priority based on the complexity of the user's voice intention;

[0009] S3: Collaboratively decide the voice processing mode based on the current network reliability index, demand priority, and mode decision rules;

[0010] S4: Classify and respond to the intent based on the voice processing mode, including:

[0011] When the voice processing mode is the offline processing mode, extract the feature vector of the voice data input by the user and perform preliminary intent classification and response based on the local lightweight intent classification model, and obtain the user feedback after the response;

[0012] Obtain the processing mode confidence based on the user feedback after the response. When the processing mode confidence is less than the confidence threshold, trigger the online processing mode for supplementary response and update the current mode decision rules;

[0013] When the voice processing mode is the online processing mode, extract the feature vector of the voice data input by the user and match it with the cloud intent knowledge base to obtain the refined intent classification result, and perform an accurate response based on the refined intent classification result and the cloud large model.

[0014] Preferably, for the offline and online mixed use method of AI voice, S1: Evaluate the current network reliability index based on the user's network connection status, including:

[0015] S101: Collect the network signal strength, network latency, network bandwidth, and network packet loss rate of the user access device in real time as the user's network connection status;

[0016] S102: Evaluate the current network reliability index based on the user's network connection status.

[0017] Preferably, for the offline and online mixed use method of AI voice, S102: Evaluate the current network reliability index based on the user's network connection status, including:

[0018]

[0019] In the formula, E is the current network reliability index, exp is the natural exponential function, α is the weight of the network bandwidth, W is the network bandwidth of the user access device, W0 is the standard bandwidth, β is the weight of the network signal strength, SI is the network signal strength of the user access device, SI0 is the standard network signal strength, γ is the weight of the network latency, t delay is the network latency of the user access device, t 0delay is the standard network latency, δ is the weight of the network packet loss rate, R PL is the network packet loss rate of the user access device, R 0PL is the standard packet loss rate.

[0020] Preferably, for the offline and online mixed use method of AI voice, S2: Analyze the complexity of the user's voice intention based on the voice data input by the user and the previous information, and generate a requirement priority based on the complexity of the user's voice intention, including:

[0021] S201: Collect the voice data input by the user and the previous information;

[0022] S202: Determine whether there are real-time knowledge retrieval requirements and third-party service call requirements in the voice data input by the user. If so, directly output the requirement priority as the first level. Otherwise, summarize the literal translation text of the voice data and the previous information to obtain the input text, analyze the topic correlation degree between the literal translation texts of adjacent voice data inputs in the input text, and perform topic differentiation on the input text based on the topic correlation degree to obtain the latest topic text input by the user;

[0023] S203: Input the latest topic text input by the user and the historical topic text into the local sentiment analysis model to obtain the current sentiment tendency of the user;

[0024] S204: Analyze the complexity of the user's voice intention based on the number of rounds of progress and the current sentiment tendency of the user's latest topic text, and generate a requirement priority based on the complexity of the user's voice intention.

[0025] Preferably, for the offline and online mixed use method of AI voice, analyze the complexity of the user's voice intention based on the number of rounds of progress and the current sentiment tendency of the user's latest topic text, including:

[0026] Determine whether the literal translation text of the voice data newly input by the user is included in the preset literal translation requirement list. If so, set the complexity of the user's latest voice intention as the lowest complexity. Otherwise, input the literal translation text of the voice data newly input by the user into the local voice intention complexity recognition model to obtain the complexity of the user's latest voice intention;

[0027] Calculate the complexity of the user's voice intention based on the number of rounds of progress of the latest topic text, the average number of rounds of progress of the online processed progress text of the user's historical topic, the average number of rounds of progress of the offline processed progress text of the historical topic, the current sentiment tendency, and the complexity of the user's latest voice intention:

[0028]

[0029] In the formula, C is the complexity of the user's voice intention. α is the weight of the complexity of the latest voice intention, C new is the complexity of the user's latest voice intention, β is the weight of the difference in the number of rounds of progress, is the average number of rounds of progress of the online processed progress text of the user's historical topic, The average number of progress rounds of the offline processed text for the historical topic, n new The number of progress rounds of the latest topic text Denote taking and the maximum value in, τ is 10 -6 , γ is the weight of the current sentiment tendency, and F is the value determined after retrieving the sentiment tendency assignment list based on the current sentiment tendency.

[0030] Preferably, for the offline and online mixed use method of AI voice, generating a demand priority based on the complexity of the user's voice intention, including:

[0031] Determining the initial demand priority based on the complexity of the user's voice intention and the initial demand priorities corresponding to different intervals of the pre-set complexity of the user's voice intention;

[0032] Judging whether there is a time limit mark in the user's latest topic text. If so, inputting the user's latest input voice data into the local time urgency analysis model to obtain the response urgency of the latest voice data. Otherwise, setting the response urgency of the latest voice data as the lowest urgency threshold;

[0033] Judging whether there is a response importance emphasis mark in the user's latest topic text. If so, inputting the user's latest input voice data into the local response importance analysis model to obtain the response importance of the latest voice data. Otherwise, setting the response importance of the latest voice data as the lowest importance threshold;

[0034] Determining the demand priority based on the initial demand priority, the response urgency and response importance of the latest voice data, and the user's response weight.

[0035] Preferably, for the offline and online mixed use method of AI voice, obtaining a processing mode confidence based on the user feedback after response, including:

[0036] Continuously tracking the voice data input by the user after the current offline processing response, and evaluating the topic correlation degree between the literal translation text of the latest input voice data and the literal translation text of the corresponding previous input voice data, until the topic correlation degree drops suddenly, determining that a topic progress has been completed;

[0037] Performing content analysis on the literal translation text and response results of all voice data of the latest completed progress topic and constructing a topic progress content tree;

[0038] Evaluating the topic dispersion degree, the lower-level progress degree of the same branch, and the lower-level progress degree of different branches of the latest completed progress topic based on the topic progress content tree;

[0039] Based on the topic dispersion degree, the lower-level progress degree within the same branch, and the lower-level progress degree between different branches of the most recently completed progress topic, the confidence level of the processing mode is evaluated.

[0040] Preferably, for the method of mixing offline and online of AI voice, based on the topic progress content tree, the topic dispersion degree, the lower-level progress degree within the same branch, and the lower-level progress degree between different branches of the most recently completed progress topic are evaluated, including:

[0041] Determine the general concept of each branch in the topic progress content tree, and count the similarity degree between the general topic concepts of every two branches in the topic progress content tree;

[0042] Based on the similarity degree between the general topic concepts of every two branches in the topic progress content tree, analyze the topic dispersion degree of the topic progress content tree;

[0043] Take the ratio of the difference degree between the literal translation texts of the voice data of the adjacent nodes of each branch in the topic progress content tree to the difference degree between the literal translation texts of the voice data of the two end nodes of the corresponding branch as the relative lower-level refinement degree of the adjacent nodes of the corresponding branch;

[0044] Based on the relative lower-level refinement degree of the adjacent nodes of each branch in the topic progress content tree, determine the lower-level progress degree within the same branch of the most recently completed progress topic;

[0045] Regard two nodes belonging to different branches and adjacent layers in the topic progress content tree as adjacent-layer nodes between different branches;

[0046] Based on all the adjacent-layer nodes between different branches in the topic progress content tree, analyze the lower-level progress degree between different branches of the most recently completed progress topic.

[0047] Preferably, for the method of mixing offline and online of AI voice, based on all the adjacent-layer nodes between different branches in the topic progress content tree, analyze the lower-level progress degree between different branches of the most recently completed progress topic, including:

[0048] Based on the difference degree between the literal translation texts of the voice data of all the adjacent-layer nodes between different branches in the topic progress content tree and the difference degree between the literal translation texts of the voice data of the two end nodes of the corresponding two branches, analyze the lower-level progress degree between different branches of the most recently completed progress topic.

[0049] Preferably, for the method of mixing offline and online of AI voice, based on the topic dispersion degree, the lower-level progress degree within the same branch, and the lower-level progress degree between different branches of the most recently completed progress topic, evaluate the confidence level of the processing mode, including:

[0050] When the topic dispersion degree of the latest completed progress topic exceeds the preset dispersion degree threshold, the confidence degree of the processing mode is evaluated based on the lower-level progress of different branches of the topic progress content tree and the topic dispersion degree;

[0051] When the topic dispersion degree of the latest completed progress topic does not exceed the preset dispersion degree threshold, the confidence degree of the processing mode is evaluated based on the lower-level progress of the same branch of the topic progress content tree and the topic dispersion degree.

[0052] The beneficial effects of the present invention compared with the prior art are as follows: By evaluating the user's network connection status, the current network reliability index can be obtained, and the network condition can be grasped in real time, providing a network-level basis for subsequent mode decisions. Analyze the complexity of the speech intention based on the user input speech data and the previous information and generate the demand priority level, fully considering the characteristics of the user's needs, making the speech processing more in line with the actual needs of the user. Collaboratively decide the speech processing mode according to the network reliability index, the demand priority level and the mode decision rule, realizing the organic combination of the network condition and the user's needs, and improving the scientificity of the mode decision. In the intention classification and response link, in the offline processing mode, the local lightweight intention classification model is used for quick preliminary classification and response, improving the instant response speed. After obtaining the user's feedback, it is decided whether to trigger the online processing mode to supplement the response according to the confidence degree of the processing mode, ensuring both quick response and accuracy, and at the same time updating the mode decision rule to optimize the subsequent decision. The online processing mode uses the cloud intention knowledge base and the large model to achieve fine intention classification and accurate response, providing high-quality services. The overall method effectively combines the advantages of offline and online, and can provide efficient, accurate and flexible speech interaction services for users under different network conditions, improving the user experience.

[0053] Other features and advantages of the present invention will be described in the following description, and, in part, will be obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structure specifically pointed out in the present application document.

[0054] The technical solutions of the present invention will be further described in detail below through the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings are used to provide a further understanding of the present invention, and constitute a part of the description, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0056] Figure 1 is the flowchart of the offline and online mixed use method of AI speech in the embodiment of the present invention;

[0057] Figure 2 is the specific execution method flowchart of step S1 in the embodiment of the present invention;

[0058] Figure 3 This is the flowchart of the specific implementation method for step S2 in the embodiments of the present invention. Detailed implementation manner

[0059] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not used to limit the present invention.

[0060] Embodiment 1:

[0061] The present invention provides a method for mixed use of offline and online AI voice. Refer to Figure 1 , including:

[0062] S1: Evaluate the current network reliability index based on the user's network connection status;

[0063] S2: Analyze the complexity of the user's voice intention based on the voice data input by the user and the previous information, and generate a demand priority based on the complexity of the user's voice intention;

[0064] S3: Collaboratively decide the voice processing mode based on the current network reliability index, demand priority, and mode decision rules;

[0065] S4: Perform intention classification and response based on the voice processing mode, including:

[0066] When the voice processing mode is the offline processing mode, extract the feature vector of the voice data input by the user and perform preliminary intention classification and response based on the local lightweight intention classification model, and obtain the user feedback after the response;

[0067] Obtain the processing mode confidence based on the user feedback after the response. When the processing mode confidence is less than the confidence threshold, trigger the online processing mode for supplementary response and update the current mode decision rules;

[0068] When the voice processing mode is the online processing mode, extract the feature vector of the voice data input by the user and match it with the cloud intention knowledge base to obtain a fine intention classification result, and perform an accurate response based on the fine intention classification result and the cloud large model.

[0069] In this embodiment, the current network reliability index: is calculated by comprehensively considering indicators such as the network signal strength, network latency, network bandwidth, and network packet loss rate of the user's access device, reflecting the reliability of the current network, and providing a network status reference for voice processing mode decision-making. For example, the higher the index, the better the network status.

[0070] In this embodiment, the previous information refers to the relevant voice or text information during the interaction with the AI voice system before the user inputs the voice data this time. This information helps the system to more comprehensively and accurately understand the user's current voice intention, such as the previously discussed topics, the requirements put forward, and other contents.

[0071] In this embodiment, the complexity of the user's voice intention is obtained by comprehensively considering various factors, such as the number of rounds of progress of the user's latest topic text, the current emotional tendency, whether the latest input voice is in the preset literal translation requirement list, and the difference in the average number of rounds of progress of the online / offline processing of the historical topic. It is used to measure the complexity of the user's voice intention. The higher the complexity, the more complex the user's needs may be.

[0072] In this embodiment, the requirement priority is determined based on the complexity of the user's voice intention, combined with whether there is a time limit mark and a response importance emphasis mark in the latest topic text. It represents the priority order of the user's voice requirements during the processing. The requirements with higher priority will be processed first.

[0073] In this embodiment, the mode decision rule is a set of rules based on the current network reliability index and the requirement priority to determine whether to adopt the offline or online voice processing mode. For example, it is stipulated that the online mode is adopted when the network reliability is high and the requirement priority is high, and the offline mode is adopted vice versa and other specific rules.

[0074] In this embodiment, based on the current network reliability index, the requirement priority, and the mode decision rule, jointly determine the voice processing mode: that is, according to the network condition reflected by the current network reliability index, the urgency and complexity of the user's requirements reflected by the requirement priority, and in accordance with the established mode decision rule, jointly determine whether to adopt the offline processing mode or the online processing mode to achieve efficient and accurate voice interaction.

[0075] In this embodiment, the offline processing mode: in this mode, the system extracts the feature vector of the user's input voice data, and uses the local lightweight intention classification model for preliminary intention classification and response, and can give a response quickly. Then, according to the user feedback, evaluate the confidence of the processing mode. If it is lower than the threshold, trigger the online processing mode to supplement the response and update the mode decision rule. It is applicable to scenarios with poor network or simple requirements.

[0076] In this embodiment, the feature vector of the voice data is the vector representation formed after feature extraction of the user's input voice data, which contains feature information such as the frequency, duration, and intonation of the voice, and is used for subsequent analysis by the intention classification model to identify the user's intention.

[0077] In this embodiment, the local lightweight intent classification model is a model deployed on a local device for performing preliminary intent classification on the feature vectors of voice data. It is characterized by small computational complexity and fast response speed. It can quickly give preliminary intent classification and responses in the offline processing mode, but its ability to handle complex intents is relatively weak.

[0078] In this embodiment, the user feedback after response refers to the feedback information given by the user to the system through subsequent input voices or other means after the system gives preliminary intent classification and responses in the offline processing mode. It can help the system evaluate the effect of the current processing mode, such as whether the user is satisfied, whether the topic progresses as expected, etc.

[0079] In this embodiment, the confidence level of the processing mode is obtained by constructing a topic progression content tree based on the user feedback after response, tracking the topic correlation degree, and evaluating the topic dispersion degree, the sub - progression degree within the same branch, the sub - progression degree between different branches, etc. It is used to measure the degree of trust in the current offline processing mode and reflects the degree to which this mode meets the user's needs.

[0080] In this embodiment, the confidence threshold is a preset standard value used to compare with the confidence level of the processing mode. When the confidence level of the processing mode is less than the confidence threshold, it indicates that the current offline processing mode may not well meet the user's needs, and it is necessary to trigger the online processing mode for supplementary response.

[0081] In this embodiment, triggering the online processing mode for supplementary response means that when the confidence level of the processing mode is less than the confidence threshold, the online processing mode is started, and the cloud resources are used to supplement and optimize the results of the previous offline processing to improve the accuracy and quality of voice interaction and better meet the user's needs.

[0082] In this embodiment, updating the current mode decision rule means adjusting and optimizing the original mode decision rule according to the situation of triggering the online processing mode for supplementary response and related data. For example, if it is found that under a certain combination of network conditions and demand priorities, the online supplementary response is frequently triggered, then the decision rule corresponding to this combination is appropriately adjusted to make the subsequent mode decision more reasonable.

[0083] In this embodiment, the online processing mode is that the system extracts the feature vectors of the user - input voice data, matches them with the cloud - based intent knowledge base to obtain a refined intent classification result, and then combines with the cloud - based large model for accurate response. This mode relies on the powerful computing power and rich data in the cloud and can handle complex intents, but it depends on a good network environment.

[0084] In this embodiment, the feature vectors of the voice data input by the user are extracted and matched with the cloud intention knowledge base to obtain a refined intention classification result: in the online processing mode, the feature vectors of the voice data are first extracted and then compared with the rich intention knowledge stored in the cloud, so as to accurately identify the category to which the user's intention belongs and provide a basis for subsequent accurate responses.

[0085] In this embodiment, the cloud large model: a model with powerful computing capabilities and rich parameters, deployed on the cloud server. In the online processing mode, based on the refined intention classification result, the cloud large model is used to generate high-quality and accurate response content to meet the complex voice interaction needs of users.

[0086] In this embodiment, an accurate response is made based on the refined intention classification result and the cloud large model: according to the refined intention classification result obtained by matching with the cloud intention knowledge base, the powerful capabilities of the cloud large model are used to generate an accurate, detailed and context-compliant response to the user's voice demand, realizing a high-quality voice interaction service.

[0087] Embodiment 2:

[0088] On the basis of Embodiment 1, an off-line and on-line mixing method for AI voice, S1: evaluate the current network reliability index based on the user's network connection status, referring to Figure 2 , including:

[0089] S101: Collect the network signal strength, network latency, network bandwidth, and network packet loss rate of the user's access device in real time as the user's network connection status;

[0090] S102: Evaluate the current network reliability index based on the user's network connection status.

[0091] In this embodiment, the user access device: refers to the device used by the user to interact with the AI voice system, such as a smart speaker, a smart phone, a vehicle-mounted system, etc. These devices have a network connection function, can collect the user's voice data and transmit it to the system, and at the same time receive the system feedback, and are the hardware carriers for realizing voice interaction.

[0092] In this embodiment, the network signal strength: characterizes the strength of the network signal received by the user access device.

[0093] In this embodiment, the network latency: refers to the time elapsed from when the user access device sends a data request to when it receives the response data.

[0094] In this embodiment, the network bandwidth: represents the amount of data that can be transmitted by the network per unit time.

[0095] In this embodiment, the network packet loss rate: refers to the proportion of data packets lost during network transmission.

[0096] Example 2 refines the process of evaluating the current network reliability index. First, it collects in real time the network signal strength, network latency, network bandwidth, and network packet loss rate of the user access device, comprehensively covering the key factors affecting network quality. The network signal strength reflects the strength of signal reception and directly affects the stability of data transmission; the network latency reflects the response speed of data transmission and is crucial for voice interaction with high real-time requirements; the network bandwidth determines the rate of data transmission and is related to whether cloud resources can be obtained quickly; the network packet loss rate indicates the integrity during data transmission. By integrating these indicators, the user's network connection status can be characterized more accurately. Based on these comprehensively collected network connection status data, the current network reliability index is evaluated, making the index more scientific and accurate. It is no longer a rough judgment based on a single factor but a comprehensive consideration of the network condition. This provides a solid foundation for subsequent collaborative decision-making on the voice processing mode based on the network reliability index, enabling a more reasonable decision on whether to use the offline or online processing mode. Furthermore, under different network conditions, the efficiency and stability of AI voice processing can be ensured, the user's voice interaction experience can be optimized, and the voice interaction service can be ensured to adapt to the complex and changing network environment.

[0097] Example 3:

[0098] Based on Example 2, for the method of mixing offline and online use of AI voice, S102: Evaluate the current network reliability index based on the user's network connection status, including:

[0099]

[0100] In the formula, E is the current network reliability index, exp is the natural exponential function, α is the weight of the network bandwidth, W is the network bandwidth of the user access device, W0 is the standard bandwidth, β is the weight of the network signal strength, SI is the network signal strength of the user access device, SI0 is the standard network signal strength, γ is the weight of the network latency, t delay is the network latency of the user access device, t 0delay is the standard network latency, δ is the weight of the network packet loss rate, R PL is the network packet loss rate of the user access device, R 0PL is the standard packet loss rate.

[0101] In this embodiment, the weight of network bandwidth: In the formula for calculating the current network reliability index, it is a value used to measure the degree of influence of network bandwidth on network reliability. It reflects the proportion of network bandwidth in evaluating the overall reliability of the network. For example, in an application scenario where data transmission speed is crucial, the weight of network bandwidth can be appropriately increased to make it have a greater impact on the calculation result of the network reliability index. Thus, when the network bandwidth is good, an online processing mode with higher bandwidth requirements is more likely to be selected.

[0102] In this embodiment, the standard bandwidth: It serves as a reference benchmark value for measuring the network bandwidth of the user access device. When calculating the network reliability index, the actual network bandwidth of the user access device is compared with the standard bandwidth, and through a certain calculation method, the influence of the network bandwidth factor on network reliability is reflected. The standard bandwidth can be set according to different network environments, application scenarios, or industry standards. For example, in the scenario of a general home broadband network, a common stable bandwidth value can be set as the standard bandwidth.

[0103] In this embodiment, the weight of network signal strength: In the formula for calculating the network reliability index, it is used to determine the importance of network signal strength in the process of evaluating network reliability. It reflects the contribution ratio of network signal strength to the network reliability index. For example, in some scenarios with extremely high requirements for network connection stability, the weight of network signal strength will be relatively high. In this way, the change in network signal strength has a more significant impact on the network reliability index. If the signal strength is weak, it may reduce the network reliability index and affect the selection of the voice processing mode.

[0104] In this embodiment, the standard network signal strength: It is a standard value used to compare the current network signal strength of the user access device. By comparing the actual signal strength with the standard network signal strength and combining its weight, the role of network signal strength in network reliability is comprehensively evaluated. The standard network signal strength can be determined according to factors such as device type and network coverage environment. For example, in an ideal indoor network coverage environment, the value corresponding to the full signal strength that the device can receive is set as the standard network signal strength.

[0105] In this embodiment, the weight of network latency: In the formula for calculating the current network reliability index, it is a parameter that measures the magnitude of the influence of network latency on network reliability. It indicates the relative importance of network latency in evaluating the overall reliability of the network. In the AI voice interaction scenario with high real-time requirements, the weight of network latency will be set relatively high because excessive latency will seriously affect the fluency of voice interaction. At this time, the change in network latency has a more prominent impact on the network reliability index. Higher latency may lead to a decrease in the network reliability index and affect the decision-making of the voice processing mode.

[0106] In this embodiment, the standard network latency: serves as a reference standard for evaluating the network latency of the user access device. When calculating the network reliability index, the actual network latency of the user access device is compared with the standard network latency for calculation, so as to reflect the effect of the network latency factor on the network reliability. The standard network latency can be set according to different network application scenarios and the user's acceptable degree of latency. For example, for the real-time voice call scenario, the value at which the human ear can hardly perceive the latency can be set as the standard network latency.

[0107] In this embodiment, the weight of the network packet loss rate: in the formula for calculating the network reliability index, it is used to measure the influence degree of the network packet loss rate on the network reliability. It reflects the component that the network packet loss rate occupies in evaluating the overall network reliability. If the integrity of data transmission is very critical in a specific application scenario, then the weight of the network packet loss rate will be relatively high, and the change of the network packet loss rate will have a greater impact on the network reliability index. A higher packet loss rate may significantly reduce the network reliability index, thereby affecting the selection of the voice processing mode.

[0108] In this embodiment, the standard packet loss rate: is a reference value for measuring the network packet loss situation of the user access device. When calculating the network reliability index, the actual network packet loss rate of the user access device is compared and calculated with the standard packet loss rate to determine the influence of the network packet loss rate on the network reliability. The standard packet loss rate can be determined according to factors such as different network types and service quality requirements. For example, for the network service scenario with high quality requirements, a very low packet loss rate value can be set as the standard packet loss rate.

[0109] The above solution evaluates the current network reliability index through a specific calculation formula, showing significant advantages in many aspects. Weights are set for network bandwidth, network signal strength, network latency, and network packet loss rate in the formula. By comprehensively considering these key network indicators, the weights of each indicator can be flexibly adjusted according to the actual application scenario to accurately quantify their influence on the network reliability. For example, in the scenario with high real-time requirements, the weight of the network latency is emphasized to highlight its importance. The natural exponential function is used to integrate various indicators for calculation, giving scientificity and rationality to the finally obtained network reliability index. The natural exponential function can effectively simulate the complex interaction relationship between various factors in the actual network environment, rather than a simple linear superposition. Using the standard bandwidth, standard network signal strength, standard network latency, and standard packet loss rate as reference benchmarks makes the evaluation results comparable and universal. No matter what network environment the user is in, a relatively objective network reliability index can be obtained based on a unified standard.

[0110] Embodiment 4:

[0111] Based on Embodiment 1, for the method of mixing offline and online use of AI voice, S2: Analyze the complexity of the user's voice intention based on the voice data input by the user and the previous information, and generate a requirement priority based on the complexity of the user's voice intention, including:

[0112] S201: Collect the voice data input by the user and the previous information;

[0113] S202: Determine whether there are real-time knowledge retrieval requirements and third-party service invocation requirements in the voice data input by the user. If so, directly output the requirement priority as level one. Otherwise, summarize the literal translation text of the voice data and the previous information to obtain the input text, analyze the topic correlation degree between the literal translation texts of the adjacent voice data inputs in the input text, and perform topic differentiation on the input text based on the topic correlation degree to obtain the latest topic text input by the user;

[0114] S203: Input the latest topic text and the historical topic text input by the user into the local sentiment analysis model to obtain the current sentiment of the user;

[0115] S204: Analyze the complexity of the user's voice intention based on the number of rounds the latest topic text of the user has advanced and the current sentiment, and generate a requirement priority based on the complexity of the user's voice intention.

[0116] In this embodiment, the real-time knowledge retrieval requirement: refers to the requirement that the user expresses through voice data to obtain specific knowledge information immediately. For example, when the user asks "What is the rise and fall situation of the stock market today", such requirements need the system to query relevant knowledge in real time to respond.

[0117] In this embodiment, the third-party service invocation requirement: means that the user reflects in the voice expression the requirement to achieve a certain purpose by relying on the services provided by a third party. For example, when the user says "Help me book a train ticket to Beijing on a certain platform", it involves invoking the third-party service of train ticket reservation.

[0118] In this embodiment, determine whether there are real-time knowledge retrieval requirements and third-party service invocation requirements in the voice data input by the user: The system analyzes the voice data input by the user to identify whether it contains the intention to obtain knowledge in real time or invoke a third-party service. If such requirements exist, the requirement priority will be directly set to level one and processed preferentially to meet the user's requirements for timeliness or specific services.

[0119] In this embodiment, the topic relevance degree between the literal translation texts of the voice data input successively in the input text is analyzed: After converting the user input voice data into literal translation texts, the texts input twice successively are analyzed to judge their relevance degree in topic content. Existing methods based on keyword matching or similarity calculation methods based on word vectors can be used, or methods based on topic models can also be used. For example, if the previous sentence says "I want to travel" and the next sentence says "Which city is suitable for playing in summer", the topic relevance degree of these two sentences is relatively high; if the next sentence says "What should I do if my computer crashes", then the topic relevance degree is relatively low.

[0120] In this embodiment, the input text is topic-distinguished based on the topic relevance degree to obtain the latest topic text input by the user: According to the calculated topic relevance degree, the input text is divided according to different topics. When the topic relevance degree changes greatly, it can be regarded as the start of a new topic, so as to determine the latest topic text that the user is currently discussing.

[0121] In this embodiment, the local sentiment tendency analysis model: is a model deployed locally. It takes the latest topic text and historical topic text input by the user as inputs, and by analyzing features such as vocabulary, tone, and semantics in the text, judges the user's current emotional state and obtains the user's current sentiment tendency.

[0122] In this embodiment, the latest topic text and historical topic text input by the user: The latest topic text is the topic content that the user is currently discussing determined through topic relevance degree analysis; the historical topic text is the topic text involved by the user in the previous interaction process with the system.

[0123] In this embodiment, the user's current sentiment tendency: After analyzing the latest topic text and historical topic text input by the user through the local sentiment tendency analysis model, the user's current emotional state is obtained. It may be manifested as positive, negative or neutral. For example, if the user says "This product is really great", it reflects a positive sentiment tendency; if the user says "This product is too difficult to use", it is a negative sentiment tendency.

[0124] The above solution analyzes the complexity of the user's speech intention from multiple dimensions and generates the requirement priority, bringing many beneficial effects. By collecting the user's input speech data and the previous information, it provides rich materials for comprehensively understanding the user's intention. It judges whether there are real-time knowledge retrieval requirements and third-party service call requirements in the speech data. If so, it directly sets the requirement priority to the first level. This mechanism ensures that critical and urgent requirements can be processed first, improving the timeliness and pertinence of service response. When the above specific requirements do not exist, the input text is obtained by summarizing the literal translation text of the speech data and the previous information, and then the topic relevance between adjacent input texts is analyzed to distinguish topics and obtain the latest topic text, which helps to accurately grasp the current focus of the user's attention. Inputting the latest topic text and the historical topic text into the local sentiment analysis model can understand the user's current sentiment tendency and provide a basis for judging the intention complexity from the sentiment dimension. Based on the number of rounds of progress and the current sentiment tendency of the latest topic text, the intention complexity is analyzed and the requirement priority is generated. Considering the dialogue process and sentiment factors comprehensively, the requirement priority is more in line with the actual requirement degree of the user. For example, if the user has had multiple rounds of conversations on a certain topic and has strong emotions, it means that the requirement is more urgent and complex, and a higher priority should be given.

[0125] Example 5:

[0126] Based on the method of mixing offline and online use of AI voice on the basis of Example 4, the complexity of the user's speech intention is analyzed based on the number of rounds of progress of the latest topic text of the user and the current sentiment tendency, including:

[0127] Judge whether the literal translation text of the user's latest input speech data is included in the preset literal translation requirement list. If so, set the complexity of the user's latest speech intention to the lowest complexity (not 0). Otherwise, input the literal translation text of the user's latest input speech data into the local speech intention complexity recognition model to obtain the complexity of the user's latest speech intention;

[0128] Calculate the complexity of the user's speech intention based on the number of rounds of progress of the latest topic text, the average number of rounds of progress of the online processed progress text of the user's historical topics, the average number of rounds of progress of the offline processed progress text of the historical topics, the current sentiment tendency, and the complexity of the user's latest speech intention:

[0129]

[0130] In the formula, C is the complexity of the user's speech intention. α is the weight of the complexity of the latest speech intention, C new is the complexity of the user's latest speech intention, β is the weight of the difference in the number of rounds of progress, is the average number of rounds of progress of the online processed progress text of the user's historical topics, is the average number of rounds of progress of the offline processed progress text of the historical topics, nnew is the number of rounds advanced for the latest topic text, Table

[0131] indicating taking and the maximum value in, τ is 10 -6 , γ is the weight of the current sentiment tendency, and F is the value determined after retrieving the sentiment tendency assignment list based on the current sentiment tendency.

[0132] In this embodiment, the preset literal translation requirement list: is a list set in advance, which contains literal translation texts corresponding to common and simple voice requirements. The system quickly determines the complexity of the user's intention by judging whether the literal translation text of the user's latest input voice data is in this list.

[0133] In this embodiment, the literal translation text of the voice data: is the text content after directly converting the voice data input by the user into text form, retaining the original expression of the voice without performing in-depth semantic understanding.

[0134] In this embodiment, the complexity of the user's latest voice intention: refers to the complexity of the intention determined according to the user's latest input voice data.

[0135] In this embodiment, the local voice intention complexity recognition model: is a model deployed locally. When the literal translation text of the user's latest input voice data is not in the preset literal translation requirement list, this model analyzes it. By understanding and processing aspects such as the semantics, syntactic structure, and vocabulary of the text, it gives an evaluation of the complexity of the user's latest voice intention, assisting the system to accurately grasp the user's complex intention.

[0136] In this embodiment, the number of rounds advanced for the latest topic text: refers to the number of times of interacting with the system around this topic from the moment the user starts discussing the current latest topic to the current moment.

[0137] In this embodiment, the average number of rounds advanced for the user's historical topic online processing progress text: counts the number of rounds advanced for each historical topic that the user has interacted with in the past through the online processing mode, and then calculates the average value of these rounds. This average value can be used as a reference and compared with the number of rounds advanced for the latest topic text to assist in judging the complexity of the current topic. If the number of rounds advanced for the current topic is much higher than this average value, it may imply that the current user's intention is more complex.

[0138] In this embodiment, the average number of progress rounds of the offline processed progress text of historical topics: Similar to the average number of progress rounds of the online processed progress text of the user's historical topics, it is the average value calculated by counting the number of progress rounds of each historical topic that the user interacted with in the past through the offline processing mode. It is also used to compare with the number of progress rounds of the latest topic text, and assist in evaluating the complexity of the current user's speech intention from the perspective of offline processing.

[0139] The above solution further refines the analysis process of the complexity of the user's speech intention and has positive effects in many aspects. First, by determining whether the literal translation requirement list contains the literal translation text of the user's latest input speech data, if it contains, directly set the complexity of the latest speech intention to 0. This fast judgment mechanism can identify common and simple requirements, reduce unnecessary complex calculations, and improve processing efficiency. For the text not in the list, use the local speech intention complexity recognition model to obtain the complexity of the latest speech intention, and with the help of the intelligent analysis ability of the model, more accurately grasp complex intentions. When comprehensively calculating the complexity of the user's speech intention, the formula considers multiple key factors. Incorporate the complexity of the latest speech intention, the difference in the number of progress rounds, and the current emotional tendency into the calculation, and set weights for each factor to comprehensively and accurately measure the intention complexity. For example, the difference in the number of progress rounds reflects the comparison of the number of rounds in this conversation with the historical online and offline processed conversation rounds. If the difference from the average number of rounds of historical online processing is large, it may mean that the intention is more complex, and adjust its impact on the overall complexity through weights. The current emotional tendency participates in the calculation after determining the value by retrieving the assignment list, so that the emotional factor can be quantitatively incorporated. If the user's emotion is strong, its corresponding value will increase the intention complexity in the calculation, reflecting that the user may have higher requirements for the service due to strong emotion.

[0140] Embodiment 6:

[0141] Based on the method of mixing offline and online use of AI voice in Embodiment 4, generating demand priorities based on the complexity of the user's speech intention, including:

[0142] Determine the initial demand priority based on the complexity of the user's speech intention and the initial demand priorities corresponding to different intervals of the complexity of the user's speech intention set in advance;

[0143] Judge whether there is a time limit mark in the user's latest topic text. If so, input the user's latest input speech data into the local time urgency analysis model to obtain the response urgency of the latest speech data. Otherwise, set the response urgency of the latest speech data to the lowest urgency threshold (not 0);

[0144] Determine whether there is a response importance emphasis flag in the user's latest topic text. If so, input the user's latest input voice data into the local response importance analysis model to obtain the response importance of the latest voice data. Otherwise, set the response importance of the latest voice data to the lowest importance threshold (not zero).

[0145] Based on the initial requirement priority, the response urgency and response importance of the latest voice data, and the user's response weight, determine the requirement priority.

[0146] In this embodiment, the initial requirement priorities corresponding to different intervals of the complexity of the user's voice intention, which is a mechanism for initially dividing the importance of requirements according to the complexity of the voice intention.

[0147] In this embodiment, the initial requirement priority: the requirement priorities corresponding to different complexity intervals preset according to the complexity of the user's voice intention. It is the basis for determining the final requirement priority and initially reflects the priority degree of the user's requirements based on complexity. For example, for a voice intention with high complexity, its initial requirement priority may be set to a higher level.

[0148] In this embodiment, the time limit mark: refers to an identifier in the user's latest topic text that can reflect the time urgency of the user's response to the requirement, such as words or expressions like "as soon as possible", "right away", "immediately", etc. It is used to prompt the system about the urgency of the requirement in terms of time.

[0149] In this embodiment, determine whether there is a time limit mark in the user's latest topic text: The system performs text analysis on the user's input latest topic text to find out whether it contains words or expressions that can indicate time urgency. If there is a time limit mark, it means that the user's requirement has time urgency and the system needs to further evaluate and adjust the requirement priority.

[0150] In this embodiment, the local time urgency analysis model: a model deployed locally. When the system determines that there is a time limit mark in the user's latest topic text, the user's latest input voice data is input into this model. The model analyzes the information related to time urgency in the text to evaluate the response urgency of the latest voice data, so as to quantify the urgency of the user's requirement in the time dimension.

[0151] In this embodiment, the response urgency of the latest voice data: a quantified value obtained by analyzing the latest voice data with a time limit mark through the local time urgency analysis model, representing the time urgency of the user's response to this requirement. The higher the value, the more urgent the requirement, and when determining the final requirement priority, the priority of this requirement will be correspondingly increased.

[0152] In this embodiment, the response importance emphasis flag: In the user's latest topic text, it is an identifier used to highlight the importance of demand response, such as words or expressions like "very important", "must", "key", etc., to prompt the system that the response to this demand has a relatively high importance.

[0153] In this embodiment, the local response importance analysis model: A model deployed locally. When the system detects that there is a response importance emphasis flag in the user's latest topic text, the system inputs the latest voice data input by the user into this model. The model determines the response importance of the latest voice data by analyzing the text content to quantify the degree of importance of this demand.

[0154] In this embodiment, the response importance of the latest voice data: A quantified value obtained after the local response importance analysis model analyzes the latest voice data with a response importance emphasis flag, which reflects the importance degree of the user's response to this demand. The higher the value, the more important the demand is, and it promotes the improvement of the demand priority when determining the final demand priority.

[0155] In this embodiment, based on the initial demand priority, the response urgency and response importance of the latest voice data, and the user's response weight, the demand priority is determined: The system comprehensively considers the initial demand priority (pre-set based on the complexity of voice intent), the response urgency of the latest voice data (reflecting the time urgency), the response importance (reflecting the importance degree), and the user's response weight (reflecting the user's specific weight preference), and integrates these factors through certain calculations or rules (such as rounding up the product of these parameters to obtain the demand priority), and finally determines the priority of the user's voice demand.

[0156] The above solution generates requirement priorities through multiple steps. By presetting corresponding initial requirement priorities for different intervals of the complexity of the user's voice intention, a basic priority framework based on the complexity of the intention is provided, enabling the system to initially distinguish the processing priorities required for different intentions and laying a foundation for subsequent fine-tuning. Judgments are made on the time limit mark and response importance emphasis mark for the user's latest topic text, which fully considers the factors of time urgency and importance in the user's requirements. If there is a time limit mark, the response urgency is obtained using the local time urgency analysis model to accurately evaluate the urgency of the requirement in the time dimension; if not, it is set to the lowest urgency threshold to ensure a unified standard. Similarly, for the judgment and processing of the response importance emphasis mark, the special requirements of the user for the importance of the response can be identified, and the response importance is obtained through the local response importance analysis model. When there is no mark, it is set to the lowest importance threshold. Finally, the final requirement priority is determined by integrating the initial requirement priority, the response urgency of the latest voice data, the response importance, and the user's response weight, comprehensively integrating multiple factors such as intention complexity, time urgency, importance, and the user's own weight preference.

[0157] Embodiment 7:

[0158] Based on Embodiment 1, the offline and online mixed use method of AI voice obtains the processing mode confidence based on the user feedback after the response, including:

[0159] Continuously track the voice data input by the user after the current offline processing response, and evaluate the topic relevance between the literal translation text of the latest input voice data and the literal translation text of the corresponding previous input voice data until the topic relevance drops sharply, and determine that a topic progression has been completed;

[0160] Perform content analysis on the literal translation text and response results of all voice data of the latest completed topic progression to construct a topic progression content tree;

[0161] Based on the topic progression content tree, evaluate the topic dispersion degree, the same-branch lower-level progression degree, and the different-branch lower-level progression degree of the latest completed topic progression;

[0162] Based on the topic dispersion degree, the same-branch lower-level progression degree, and the different-branch lower-level progression degree of the latest completed topic progression, evaluate the processing mode confidence.

[0163] In this embodiment, the current offline processing response refers to the response given by the system to the user after initially classifying the intention of the user input voice data based on the local lightweight intention classification model in the offline processing mode. For example, when the user asks "What restaurants are near here", in the offline processing mode, the system may quickly reply "We have found several popular restaurants nearby, such as XX Restaurant and XX Restaurant".

[0164] In this embodiment, the topic relevance between the literal translation text of the latest input voice data and the literal translation text of the corresponding previous input voice data is evaluated: The system compares and analyzes the literal translation text converted from the user's latest input voice and the literal translation text converted from the previous input voice, and judges the degree of relevance between the two in terms of topic content through technical means such as semantic understanding and keyword matching.

[0165] In this embodiment, there is a sudden drop in topic relevance: When the system continuously evaluates the topic relevance between the literal translation texts of the latest and previous input voice data, it is found that the relevance value suddenly drops significantly, which indicates that the user's topic has changed greatly, meaning the completion of a topic progression. For example, when the user was originally discussing travel plans and suddenly started asking about work-related matters, the topic relevance would drop suddenly.

[0166] In this embodiment, the literal translation texts and response results of all voice data of the latest completed topic progression: From the start to the sudden drop in topic relevance, during the process of completing a topic conversion, the literal translation texts converted from all the user's input voice data, as well as the offline processing response results given by the system for these inputs. For example, in a discussion about the travel topic, the user asks "I want to travel to the seaside. What places are recommended?", and the system replies "Sanya, Qingdao and other places are very suitable for seaside travel" and a series of conversation contents.

[0167] In this embodiment, the literal translation texts and response results of all voice data of the latest completed topic progression are content-analyzed and a topic progression content tree is constructed: The system deeply analyzes all the above-mentioned texts and results of completing a topic, extracts key information, and constructs a tree structure according to the logical relationship and hierarchical structure of the topic. For example, taking the topic theme as the root node and different sub-topics or related contents as branch nodes, a tree that can intuitively show the development context of the topic is constructed.

[0168] In this embodiment, the topic dispersion degree, the lower-level progression degree within the same branch, and the lower-level progression degree between different branches of the latest completed topic progression:

[0169] Topic dispersion degree: It is analyzed by determining the general concepts of each branch in the topic progression content tree and counting the similarity between the general topic concepts of every two branches. It reflects the degree of dispersion of the topic during the unfolding process. If the topic dispersion degree is high, it means that the topic unfolds in multiple different directions during the discussion.

[0170] Same-branch lower-level progress: By comparing the difference between the literal translations of the speech data of adjacent nodes in each branch of the topic progression content tree with the difference between the literal translations of the speech data of the two end nodes of the corresponding branch, the relative lower-level refinement degree is obtained, and then the same-branch lower-level progress is determined. It reflects the refinement and progression of the topic within the same branch. The higher the progress, the better the in-depth development of the topic within the branch.

[0171] Different-branch lower-level progress: Nodes in different branches and adjacent layers in the topic progression content tree are regarded as different-branch adjacent-layer nodes, and it is analyzed based on the difference between the literal translations of the speech data of these nodes and the difference between the literal translations of the speech data of the two end nodes of the corresponding two branches. It reflects the horizontal correlation and progression between different topic branches. A high progress means better coordinated development between different topic branches. These metrics are used to evaluate the degree to which the offline processing mode meets user needs.

[0172] The above solution obtains the confidence level of the processing mode in a unique way based on the user feedback after the response. Continuously track the speech data input by the user after the current offline processing response, and determine the topic progression by evaluating the topic relevance. This method can closely analyze around the user's conversation flow, accurately capture the nodes of the user's topic conversion, and provide a clear timeline and topic context for subsequent in-depth analysis. Parse the content of a completed progressive topic and construct a topic progression content tree, integrating various information in the topic in a structured way. In this way, the system can intuitively understand the internal logical structure of the topic and the relationship between the contents of each part, providing strong data support for further evaluating topic characteristics. Evaluate the topic dispersion, same-branch lower-level progress, and different-branch lower-level progress based on the topic progression content tree, and quantitatively analyze the topic from multiple dimensions. Topic dispersion reflects the degree of dispersion of the topic during the unfolding process, helping to judge whether the user's thinking is centered around the main line; the same-branch lower-level progress and different-branch lower-level progress measure the progression of the topic in terms of depth and breadth from the perspectives of the same logical branch and different logical branches respectively. These metrics comprehensively characterize the characteristics of topic progression and can effectively reflect the degree to which the offline processing mode meets user needs. Evaluating the confidence level of the processing mode by integrating these three dimensions makes the determination of the confidence level more scientific and reasonable.

[0173] Example 8:

[0174] Based on Example 7, the offline and online mixed use method of AI speech evaluates the topic dispersion, same-branch lower-level progress, and different-branch lower-level progress of the latest completed progressive topic based on the topic progression content tree, including:

[0175] Determine the generalized concept of each branch in the topic progression content tree, and calculate the similarity between the generalized topic concepts of every two branches in the topic progression content tree;

[0176] Based on the similarity between the generalized topic concepts of each two branches in the topic progression content tree, the topic dispersion of the topic progression content tree is analyzed;

[0177] The ratio of the difference between the voice data literal translation texts of the adjacent nodes of each branch of the topic progression content tree and the difference between the voice data literal translation texts of the two end nodes of the corresponding branch is regarded as the relative lower refinement degree of the adjacent nodes of the corresponding branch;

[0178] Determine the lower advancement degree of the same branch of the most recently completed progress topic based on the relative lower refinement degree of the adjacent nodes of each branch of the topic progress content tree;

[0179] Treat two nodes belonging to different branches and adjacent layers in the topic progression content tree as different-branch adjacent layer nodes;

[0180] Based on all the different-branch adjacent layer nodes in the topic progress content tree, the different-branch lower-level advancement degree of the current most recently completed progress topic is analyzed.

[0181] In this embodiment, the general concept of each branch: In the topic progression content tree, each branch represents the core and general concept of the topic. It is at the beginning of the branch, commanding the specific content under the branch, and can reflect the main direction of the branch topic. For example, in the topic progression content tree discussing tourism, a branch revolves around "domestic tourist attractions", then "domestic tourist attractions" is the general concept of this branch.

[0182] In this embodiment, the similarity between the generalized topic concepts of every two branches in the topic progression content tree is statistically calculated: the system compares and analyzes the generalized concepts of any two branches in the topic progression content tree, and obtains the similarity between them through methods such as semantic similarity calculation and keyword matching. For example, if one branch generalized concept is "natural scenic spots" and the other is "historical and cultural attractions", the similarity between the two concepts in the scope of tourism topics can be obtained through analysis.

[0183] In this embodiment, based on the proximity between the general topic concepts of every two branches in the topic progression content tree, the topic dispersion degree of the topic progression content tree is analyzed: If the proximity between the general concepts of each branch is generally high, it indicates that the topics are relatively concentrated in similar fields and the topic dispersion degree is low; on the contrary, if the proximity is low, it shows that the topics are developed in different directions and the topic dispersion degree is high. For example, if multiple branches respectively revolve around different core concepts such as "seaside tourism", "mountain hiking", and "city sightseeing", their proximity is low and the topic dispersion degree is high.

[0184] In this embodiment, the relative lower-level refinement degree of adjacent nodes of a branch: In each branch of the topic progression content tree, the ratio of the difference degree between the literal translation texts of adjacent node voice data to the difference degree between the literal translation texts of the voice data of the two end nodes of the corresponding branch is calculated, and the obtained value is the relative lower-level refinement degree. It reflects the degree of content refinement when the topic progresses between adjacent nodes. For example, if the starting node of a branch is "selection of tourist destinations" and the adjacent node becomes "popular seaside tourist destinations", the relative lower-level refinement degree can be calculated by comparing with the difference degree between the two end nodes of the branch.

[0185] In this embodiment, based on the relative lower-level refinement degree of adjacent nodes of each branch of the topic progression content tree, the same-branch lower-level advancement degree of the currently latest completed topic progression is determined: The relative lower-level refinement degrees of adjacent nodes of each branch are comprehensively considered, such as calculating the uniformity or using other statistical methods, so as to determine the advancement degree of the entire topic within the same branch from the start to the end. If the uniformity of the relative lower-level refinement degrees of adjacent nodes of each branch is high, it indicates that the topic continuously and evenly deepens and refines within the same branch, and the same-branch lower-level advancement degree is high, reflecting that the offline processing mode guides the topic to have good depth expansion in the direction of this branch.

[0186] The above solution further refines the process of evaluating topic features based on the topic progression content tree, bringing many beneficial effects. Determine the general concepts of each branch of the topic progression content tree, and count the similarity between every two general topic concepts of the branches, providing a basis for measuring topic dispersion. Similarity analysis can intuitively show the degree of tight connection between the topics of each branch. If the similarity is high, it means that the topics are relatively concentrated around the core; otherwise, it indicates that the topics are relatively dispersed. In this way, the topic dispersion is accurately quantified, enabling the system to clearly understand the discrete situation of the topics during the conversation of the user, and determining whether the offline processing mode can guide the user to stay within a relatively focused discussion scope. Determine the relative lower-level refinement degree by the ratio of the literal translation text difference degree between the adjacent nodes of the branch to the difference degree between the two end nodes, and then obtain the lower-level progression degree of the same branch, deeply analyzing the development situation from within the topic branch. This method considers the refinement process of the topic within the same branch. If the relative lower-level refinement degree is high, it means that the topic progression changes greatly between adjacent nodes within the branch, reflecting that the topic has good depth expansion within this branch, which helps the system evaluate the effect of the offline processing mode in guiding the topic to go deeper. By defining the adjacent layer nodes of different branches and analyzing the lower-level progression degree of different branches based on these nodes, the topic development is investigated from the relationship between different branches. This can reflect the horizontal association and progression between different topic branches. If the lower-level progression degree of different branches is high, it indicates that there is good coordinated development between different topic branches, showing that the offline processing mode can effectively guide the user to naturally transition between multiple related topics. Through the above multi-dimensional evaluation methods, the system can comprehensively and meticulously analyze the topic dispersion, the lower-level progression degree of the same branch, and the lower-level progression degree of different branches based on the topic progression content tree, thereby more accurately judging the degree to which the offline processing mode guides the user's topics and meets the user's needs.

[0187] Embodiment 9:

[0188] Based on Embodiment 8, for the offline and online mixed use method of AI voice, analyze the lower-level progression degree of different branches of the currently latest completed progression topic based on all adjacent layer nodes of different branches in the topic progression content tree, including:

[0189] Based on the difference degree between the literal translation texts of the speech data of all non-branch adjacent layer nodes in the topic progression content tree and the difference degree between the literal translation texts of the speech data of the two end nodes corresponding to the two branches, analyze the non-branch lower-level progression degree of the currently latest completed progression topic, including: comparing and analyzing the difference degree between the literal translation texts of the speech data of the non-branch adjacent layer nodes with the difference degree between the literal translation texts of the speech data of the two end nodes corresponding to the two branches. If the difference degree between the non-branch adjacent layer nodes is relatively small and has a certain correlation with the difference degree between the end nodes of the branches, it indicates that the topic conversion between different branches is relatively natural and the progression is orderly at the adjacent levels, and the non-branch lower-level progression degree is relatively high, meaning that the offline processing mode effectively guides the user to switch and progress between different topic branches; on the contrary, if the difference degree is too large or lacks correlation, it may indicate that there are problems in the progression between branches and the non-branch lower-level progression degree is relatively low. For example, from "mountains in natural landscapes" to "rivers in natural landscapes", if the difference between the texts of these two non-branch adjacent layer nodes is within a reasonable range and the relationship with the difference degrees of the starting and ending nodes of their respective branches is coordinated, it indicates that the progression between different branches is good and the non-branch lower-level progression degree is high.

[0190] In this embodiment, the difference degree between the literal translation texts of the speech data of all non-branch adjacent layer nodes in the topic progression content tree: In the topic progression content tree, for the literal translation texts of the speech data corresponding to the nodes at adjacent levels in different branches, it is the degree of difference obtained through comparative analysis in terms of semantics, vocabulary, etc. For example, in one branch, a certain layer of nodes discusses "mountains in natural landscapes", and in another branch at the adjacent layer, the nodes discuss "rivers in natural landscapes". The system analyzes the literal translation texts of the speech data corresponding to these two nodes to determine the difference degree between them, which is used to measure the difference in topic content between different branches at the adjacent levels.

[0191] The above solution analyzes the progress of the lower-level nodes of different branches in the content tree based on the topic in a specific way, bringing the following beneficial effects. From the perspective of the integrity of the topic development assessment, the progress of the lower-level nodes of different branches is analyzed by using the difference degree of the literal translations of the voice data of the adjacent-layer nodes of different branches and the difference degree of the literal translations of the voice data of the two end nodes of the corresponding two branches, providing a quantitative basis for comprehensively understanding the horizontal expansion and progress of the topic between different branches. This method not only considers the direct relationship between different branches at the same level (through adjacent-layer nodes), but also takes into account the differences between the two end nodes of the entire branch, making the analysis of the relationship between different branches more comprehensive and three-dimensional. In evaluating the effect of the voice processing mode, this method can more accurately reflect the ability of the offline processing mode to guide users to switch and progress between different topic branches. If the relationship between the difference degree between the adjacent-layer nodes of different branches and the difference degree between the two end nodes of the branch shows a reasonable change, it indicates that the offline processing mode effectively promotes the natural transition between different related topics and helps users explore in depth at multiple topic levels. This helps the system determine whether the offline processing mode successfully guides the conversation, and thus provides strong support for determining the confidence level of the processing mode. For optimizing the AI voice interaction experience, by accurately analyzing the progress of the lower-level nodes of different branches, the system can better understand the user's behavior pattern and demand changes during the topic switching process, and thus adjust the subsequent voice processing strategy accordingly. If it is found that there are problems with the progress of different branches, a more suitable processing mode (such as the online processing mode) can be triggered in time for optimization.

[0192] Embodiment 10:

[0193] Based on the topic dispersion degree, the progress of the lower-level nodes of the same branch, and the progress of the lower-level nodes of different branches of the currently latest completed progressive topic, on the basis of Embodiment 8, an AI voice off-line and on-line mixed use method evaluates the confidence level of the processing mode, including:

[0194] If the topic dispersion degree of the currently latest completed progressive topic exceeds the preset dispersion degree threshold, then the confidence level of the processing mode is evaluated based on the progress of the lower-level nodes of different branches and the topic dispersion degree of the topic progress content tree;

[0195] If the topic dispersion degree of the currently latest completed progressive topic does not exceed the preset dispersion degree threshold, then the confidence level of the processing mode is evaluated based on the progress of the lower-level nodes of the same branch and the topic dispersion degree of the topic progress content tree.

[0196] In this embodiment, the preset dispersion degree threshold: a standard value preset for measuring the topic dispersion degree. When evaluating the confidence level of the processing mode, it is used as a boundary to judge the high or low of the topic dispersion degree. The specific value can be set according to the actual application scenario and experience.

[0197] In this embodiment, the confidence level of the processing mode is evaluated based on the lower-level progress of different branches and the topic dispersion degree of the topic progress content tree: when the topic dispersion degree exceeds the preset dispersion threshold, it means that the topic development is relatively discrete. At this time, the lower-level progress of different branches can reflect the ability of the offline processing mode to establish connections between different discrete topics and promote the development of the conversation. If the lower-level progress of different branches is high, it indicates that the offline processing mode can better coordinate the conversion and progress between different topic branches, and can effectively guide the conversation even if the topics are dispersed, so the confidence level of the processing mode is relatively high; on the contrary, if the lower-level progress of different branches is low, it means that the offline processing mode has poor effect in processing dispersed topics, and the confidence level of the processing mode will decrease.

[0198] In this embodiment, the confidence level of the processing mode is evaluated based on the lower-level progress of the same branch and the topic dispersion degree of the topic progress content tree: when the topic dispersion degree does not exceed the preset dispersion threshold, that is, the topic is relatively concentrated. The lower-level progress of the same branch reflects the refinement and in-depth development of the topic within the same branch. If the lower-level progress of the same branch is high, it indicates that the offline processing mode can guide the user to discuss deeply around the core topic, and the confidence level of the processing mode is relatively high; if the lower-level progress of the same branch is low, it means that the offline processing mode has poor effect in guiding the topic to develop deeply, and the confidence level of the processing mode will decrease.

[0199] The above solution has significant advantages in many aspects by evaluating the confidence level of the processing mode according to different situations of the topic dispersion degree, combined with the lower-level progress of the same branch or different branches. When the topic dispersion degree exceeds the preset dispersion threshold, it indicates that the topic development is relatively discrete. At this time, evaluating the confidence level of the processing mode based on the lower-level progress of different branches and the topic dispersion degree can focus on the association and progress between different topic branches. Because topic dispersion may mean that the user is discussing in multiple different directions, and the lower-level progress of different branches can reflect the ability of the offline processing mode to establish connections between these dispersed topics and promote the development of the conversation. Through this evaluation method, it can accurately judge whether the offline processing mode can still effectively control the conversation direction in the case of relatively dispersed topics, and then determine the confidence level of this processing mode. If the topic dispersion degree does not exceed the preset threshold, it means that the topic is relatively concentrated. Evaluating the confidence level of the processing mode based on the lower-level progress of the same branch and the topic dispersion degree focuses on the in-depth development of the topic within the same branch. The lower-level progress of the same branch reflects the refinement and promotion effect of the topic under a single context. Evaluating the confidence level of the processing mode in this way can judge the effectiveness of the offline processing mode in guiding the user to discuss deeply around the core topic. This way of flexibly adjusting the evaluation basis according to the topic dispersion degree makes the evaluation of the confidence level of the processing mode more scientific and accurate.

[0200] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. An offline and online mixed use method for AI voice, characterized in that, include: S1: Evaluate the current network reliability index based on the user's network connection status; S2: Analyze the complexity of the user's voice intention based on the voice data input by the user and the previous context information, and generate demand priority based on the complexity of the user's voice intention; S3: Collaboratively decide on the voice processing mode based on the current network reliability index, demand priority, and mode decision rules; S4: Intent classification and response based on speech processing mode, including: When the voice processing mode is the offline processing mode, the feature vector of the voice data input by the user is extracted and preliminary intent classification and response are performed based on the local lightweight intent classification model, and user feedback after the response is obtained; The confidence of the processing mode is obtained based on the user feedback after the response. When the confidence of the processing mode is less than the confidence threshold, the online processing mode is triggered to make a supplementary response and update the current mode decision rule; When the voice processing mode is the online processing mode, the feature vector of the voice data input by the user is extracted and matched with the cloud intent knowledge base to obtain a refined intent classification result, and an accurate response is made based on the refined intent classification result and the cloud large model.

2. The offline and online mixed use method of AI voice according to claim 1, characterized in that, S1: Evaluate the current network reliability index based on the user's network connection status, including: S101: collecting network signal strength, network delay, network bandwidth, and network packet loss rate of the user's access device in real time as the user's network connection status; S102: Evaluate the current network reliability index based on the user's network connection status.

3. The offline and online mixed use method of AI voice according to claim 2, characterized in that S102: Evaluate the current network reliability index based on the user's network connection status, including: Where, E is the current network reliability index, exp is the natural exponential function, α is the weight of network bandwidth, W is the network bandwidth of the user access device, W0 is the standard bandwidth, β is the weight of network signal strength, SI is the network signal strength of the user access device, SI0 is the standard network signal strength, γ is the weight of network delay, t delay is the network delay of the user access device, t 0delay is the standard network delay, δ is the weight of network packet loss rate, R PL is the network packet loss rate of the user access device, R 0PL is the standard packet loss rate.

4. The offline and online mixed use method of AI voice according to claim 1, characterized in that S2: Analyze the complexity of the user's voice intention based on the voice data input by the user and the previous context information, and generate demand priorities based on the complexity of the user's voice intention, including: S201: Collecting voice data and previous text information input by the user; S202: Determine whether the voice data input by the user has a real-time knowledge retrieval demand and a third-party service call demand. If so, directly output the demand priority as level one. Otherwise, aggregate the translated text of the voice data with the previous text information to obtain the input text, analyze the topic relevance between the translated texts of the voice data inputted adjacently in the input text, and distinguish the input text by topic based on the topic relevance to obtain the latest topic text input by the user. S203: Input the latest topic text and the historical topic text input by the user into the local sentiment tendency analysis model to obtain the user's current sentiment tendency; S204: Analyze the complexity of the user's voice intention based on the number of rounds of the user's latest topic text and the current emotional tendency, and generate a demand priority based on the complexity of the user's voice intention.

5. The offline and online mixed use method of AI voice according to claim 4, characterized in that The complexity of the user's voice intent is analyzed based on the number of rounds of the user's latest topic text and the current sentiment tendency, including: Determine whether the literal translation text of the user's latest input voice data is included in the preset literal translation requirement list. If so, set the complexity of the user's latest voice intention to the lowest complexity. Otherwise, input the literal translation text of the user's latest input voice data into the local voice intention complexity recognition model to obtain the complexity of the user's latest voice intention; Calculate the complexity of the user's voice intention based on the number of rounds of progress of the latest topic text, the average number of rounds of progress of the user's historical topic online processed progress text, the average number of rounds of progress of the historical topic offline processed progress text, the current sentiment tendency, and the complexity of the user's latest voice intention: Where C is the complexity of the user's speech intention. α is the weight of the complexity of the latest speech intention, C new is the complexity of the user's latest speech intention, β is the weight of the difference in the number of progress rounds, is the average number of progress rounds of the online processed progress text of the user's historical topic, is the average number of progress rounds of the offline processed progress text of the historical topic, n new is the number of progress rounds of the latest topic text, denotes taking and the maximum value of, τ is 10 -6 , γ is the weight of the current emotional tendency, and F is the value determined after retrieving the emotional tendency assignment list based on the current emotional tendency.

6. The offline and online mixed use method of AI voice according to claim 4, characterized in that Generate a requirement priority based on the complexity of the user's voice intention, including: Determine the initial requirement priority based on the complexity of the user's voice intention and the initial requirement priorities corresponding to different intervals of the preset complexity of the user's voice intention; Determine whether there is a time limit mark in the user's latest topic text. If so, input the user's latest input voice data into the local time urgency analysis model to obtain the response urgency of the latest voice data. Otherwise, set the response urgency of the latest voice data to the lowest urgency threshold; Determine whether there is a response importance emphasis flag in the user's latest topic text. If so, input the user's latest input voice data into the local response importance analysis model to obtain the response importance of the latest voice data. Otherwise, set the response importance of the latest voice data to the lowest importance threshold; Determine the requirement priority based on the initial requirement priority, the response urgency and response importance of the latest voice data, and the user's response weight.

7. The offline and online mixed use method of AI voice according to claim 1, characterized in that Obtain the processing mode confidence based on the user feedback after the response, including: Continuously track the voice data input by the user after the current offline processing response, and evaluate the topic relevance between the literal translation text of the latest input voice data and the literal translation text of the corresponding previous input voice data until the topic relevance drops sharply, and determine that a topic progress has been completed; Parse the content of the literal translation text and response results of all voice data of the latest completed progress topic and construct a topic progress content tree; Evaluate the topic dispersion, the lower-level progress within the same branch, and the lower-level progress between different branches of the latest completed progress topic based on the topic progress content tree; Evaluate the processing mode confidence based on the topic dispersion, the lower-level progress within the same branch, and the lower-level progress between different branches of the latest completed progress topic.

8. The offline and online mixed use method of AI voice according to claim 7, characterized in that, Evaluate the topic dispersion, the lower-level progress within the same branch, and the lower-level progress between different branches of the latest completed progress topic based on the topic progress content tree, including: Determine the general concept of each branch in the topic progress content tree and count the similarity between the general topic concepts of every two branches in the topic progress content tree; Analyze the topic dispersion of the topic progress content tree based on the similarity between the general topic concepts of every two branches in the topic progress content tree; The ratio of the difference degree between the literal translation texts of the voice data of adjacent nodes of each branch of the topic progression content tree to the difference degree between the literal translation texts of the voice data of the two end nodes of the corresponding branch is regarded as the relative lower-level refinement degree of the adjacent nodes of the corresponding branch; Based on the relative lower-level refinement degree of the adjacent nodes of each branch of the topic progression content tree, determine the lower-level advancement degree of the same branch of the currently latest completed progression topic; Regard two nodes that belong to different branches and adjacent layers in the topic progression content tree as adjacent nodes of different branches and adjacent layers; Based on all adjacent nodes of different branches and adjacent layers in the topic progression content tree, analyze the lower-level advancement degree of different branches of the currently latest completed progression topic.

9. The offline and online mixed use method of AI voice according to claim 8, characterized in that, Based on all adjacent nodes of different branches and adjacent layers in the topic progression content tree, analyze the lower-level advancement degree of different branches of the currently latest completed progression topic, including: Based on the difference degree between the literal translation texts of the voice data of all adjacent nodes of different branches and adjacent layers in the topic progression content tree and the difference degree between the literal translation texts of the voice data of the two end nodes of the corresponding two branches, analyze the lower-level advancement degree of different branches of the currently latest completed progression topic.

10. The offline and online mixed use method of AI voice according to claim 8, characterized in that, Based on the topic dispersion degree, the lower-level advancement degree of the same branch, and the lower-level advancement degree of different branches of the currently latest completed progression topic, evaluate the confidence degree of the processing mode, including: If the topic dispersion degree of the currently latest completed progression topic exceeds the preset dispersion threshold, then evaluate the confidence degree of the processing mode based on the lower-level advancement degree of different branches and the topic dispersion degree of the topic progression content tree; If the topic dispersion degree of the currently latest completed progression topic does not exceed the preset dispersion threshold, then evaluate the confidence degree of the processing mode based on the lower-level advancement degree of the same branch and the topic dispersion degree of the topic progression content tree.

Citation Information

Cited By

  • Data display method and system for service suitable for aging, medium and program product

    CN120544572A

  • Pet equipment voice control method and system based on intelligent switching

    CN120808775A

  • Voice equipment response method and device and electronic equipment

    CN121237091A