Real-time conversation assistance method, electronic device, storage medium, and program product

By extracting the rhythmic features of real-time voice dialogue data to generate early warning signals and initiating the script preloading process, the problem of delayed script push in real-time dialogue assistance systems is solved, enabling timely assistance.

CN122454979APending Publication Date: 2026-07-24ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG GEELY HLDG GRP CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

The problem of real-time dialogue assistance systems failing to function due to delays in script delivery.

Method used

By extracting the dialogue rhythm features from real-time voice dialogue data, a dialogue bottleneck warning signal is generated, and in response to this signal, a script preloading process is initiated to generate candidate scripts, thereby achieving dialogue assistance.

Benefits of technology

It effectively avoids response delays in real-time dialogue assistance systems, ensures timely delivery of dialogue scripts, and improves the effectiveness of real-time dialogue assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454979A_ABST
    Figure CN122454979A_ABST
Patent Text Reader

Abstract

The application provides a real-time dialogue assistance method, an electronic device, a storage medium and a program product, relates to the technical field of voice data processing, and acquires real-time voice dialogue data; extracts dialogue rhythm characteristics of the real-time voice dialogue data, and generates a dialogue bottleneck early warning signal based on the dialogue rhythm characteristics; a dialogue technique preloading process is started in response to the dialogue bottleneck early warning signal to generate a candidate dialogue technique; and dialogue assistance is performed based on the candidate dialogue technique. The application can solve the problem that a real-time dialogue assistance system cannot function due to lag in dialogue technique pushing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice data processing technology, and in particular to a real-time dialogue assistance method, electronic device, storage medium, and program product. Background Technology

[0002] In related technologies, real-time dialogue assistance systems typically adopt a trigger-based script generation mode. Only after the user triggers the assistance request (such as manually pressing the trigger button) will the system sequentially carry out operations such as automatic speech recognition (ASR) transcription, semantic parsing, and script generation, and finally push the generated script to the user terminal. Because each functional step in this mode needs to be completed sequentially, the overall time consumption is relatively long, which directly causes the system script push to be delayed, thus failing to play the role of real-time dialogue assistance. Summary of the Invention

[0003] The main objective of this application is to provide a real-time dialogue assistance method, electronic device, storage medium, and program product, which aims to solve the problem that real-time dialogue assistance systems cannot function due to delays in script delivery.

[0004] To achieve the above objectives, a first aspect of this application proposes a real-time dialogue assistance method, the method comprising: Acquire real-time voice dialogue data; Extract the dialogue rhythm features from the real-time voice dialogue data, and generate a dialogue bottleneck warning signal based on the dialogue rhythm features; In response to the dialogue bottleneck warning signal, the script preloading process is initiated to generate candidate scripts; Dialogue assistance is provided based on the candidate dialogues.

[0005] In some embodiments, generating a dialogue bottleneck warning signal based on the dialogue rhythm features includes: The dialogue rhythm features are compared with a preset rhythm warning threshold to obtain a first comparison result; If the first comparison result indicates that the dialogue rhythm feature is greater than or equal to the rhythm warning threshold, a dialogue bottleneck warning signal is generated.

[0006] In some embodiments, the method further includes: The dialogue rhythm features are compared with a preset rhythm trigger threshold to obtain a second comparison result; the rhythm trigger threshold is greater than the rhythm warning threshold. If the second comparison result indicates that the dialogue rhythm feature is greater than or equal to a preset rhythm trigger threshold, a dialogue assistance trigger signal is generated. The dialogue assistance based on the candidate dialogues includes: In response to the dialogue assistance trigger signal, the candidate dialogues are actively pushed to assist in the dialogue.

[0007] In some embodiments, the method includes: Semantic analysis is performed on the dialogue text data corresponding to the real-time voice dialogue data to obtain semantic analysis results; A dialogue bottleneck warning signal is generated based on the semantic analysis results.

[0008] In some embodiments, the semantic analysis result includes a dialogue score, and generating a dialogue bottleneck warning signal based on the semantic analysis result includes: The dialogue score is compared with a preset first score threshold to obtain a third comparison result; If the third comparison result indicates that the dialogue score is less than the first score threshold, a dialogue bottleneck warning signal is generated.

[0009] In some embodiments, the dialogue assistance based on the candidate dialogues includes at least one of the following: If the dialogue score is less than a preset second score threshold, the candidate dialogue words are actively pushed to assist in the dialogue; the second score threshold is less than the first score threshold. If the number of times the dialogue bottleneck warning signal is generated is greater than or equal to a preset threshold, the candidate dialogue is actively pushed to assist in the dialogue.

[0010] In some embodiments, the number of candidate dialogues is greater than one, and after the dialogue preloading process is initiated in response to the dialogue bottleneck warning signal to generate candidate dialogues, the method further includes: Based on the semantic analysis results, the push weights of each of the multiple candidate dialogues are adjusted to obtain multiple candidate dialogues with adjusted push weights. The dialogue assistance based on the candidate dialogues includes: The target candidate dialogue is pushed from among the multiple candidate dialogues to assist in the dialogue; the push weight of the target candidate dialogue is higher than the push weight of other candidate dialogues among the multiple candidate dialogues excluding the target candidate dialogue.

[0011] In some embodiments, after initiating the dialogue preloading process to generate candidate dialogues in response to the dialogue bottleneck warning signal, the method further includes: New candidate statements are generated in response to a topic switching signal; the topic switching signal is generated when a topic switch occurs in the real-time voice dialogue data. The dialogue assistance based on the candidate dialogues includes: The new candidate phrases are pushed to assist in the conversation.

[0012] To achieve the above objectives, a second aspect of this application proposes a real-time dialogue assistance system, the system comprising: The acquisition module is used to acquire real-time voice dialogue data; The prediction module is used to extract the dialogue rhythm features of the real-time voice dialogue data and generate a dialogue bottleneck warning signal based on the dialogue rhythm features. The preloading module is used to initiate a dialogue preloading process to generate candidate dialogues in response to the dialogue bottleneck warning signal; An auxiliary module is used to provide dialogue assistance based on the candidate dialogues.

[0013] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the real-time dialogue assistance method described in the first aspect.

[0014] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the real-time dialogue assistance method described in the first aspect.

[0015] To achieve the above objectives, a fifth aspect of the present application provides a computer program product, which includes a computer program that, when executed by a processor, implements the real-time dialogue assistance method provided in the first aspect above.

[0016] The real-time dialogue assistance method, system, electronic device, computer-readable storage medium, and computer program product proposed in this application acquire real-time voice dialogue data; extract the dialogue rhythm features of the real-time voice dialogue data, and generate a dialogue bottleneck warning signal based on the dialogue rhythm features; initiate a script preloading process to generate candidate scripts in response to the dialogue bottleneck warning signal; and provide dialogue assistance based on the candidate scripts.

[0017] Compared to traditional solutions that only sequentially perform voice ASR transcription, semantic parsing, and script generation after a user triggers an assistance request, ultimately pushing the script to the user's terminal, this application embodiment extracts the dialogue rhythm features of real-time voice dialogue data and generates a dialogue bottleneck warning signal based on these features. In response to this warning signal, a script preloading process is initiated to generate candidate scripts, which are then used for dialogue assistance. Thus, this embodiment eliminates the need to wait for a user to trigger an assistance request before analyzing the voice and generating scripts. Instead, it generates a dialogue bottleneck warning signal based on the dialogue rhythm features and automatically initiates the script preloading process to generate candidate scripts. This effectively avoids delays in script delivery by the real-time dialogue assistance system, reduces the response latency of the real-time voice assistance system, and ultimately solves the problem of the real-time dialogue assistance system failing to function due to delayed script delivery. Attached Figure Description

[0018] Figure 1 A flowchart illustrating the steps of the real-time dialogue assistance method provided in some embodiments of this application; Figure 2 for Figure 1 A detailed flowchart of step S102; Figure 3 A flowchart illustrating the steps of the real-time dialogue assistance method provided in this application in some other embodiments; Figure 4 A flowchart illustrating the steps of the real-time dialogue assistance method provided in some other embodiments of this application; Figure 5 for Figure 4 A detailed flowchart of step S402; Figure 6 for Figure 1 A detailed flowchart of step S104; Figure 7 A flowchart illustrating the steps of the real-time dialogue assistance method provided in some other embodiments of this application; Figure 8 A flowchart illustrating the steps of the real-time dialogue assistance method provided in some other embodiments of this application; Figure 9 The real-time dialogue assistance method provided in the embodiments of this application includes system flow diagrams in some embodiments; Figure 10 This is a schematic diagram of the structure of the real-time dialogue assistance system provided in the embodiments of this application; Figure 11 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0020] It should be noted that although functional modules are divided in the device / system schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device / system or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0022] First, a brief explanation of the overall concept of the real-time dialogue assistance method provided in the embodiments of this application will be given.

[0023] With the deepening of digital transformation in the sales field, real-time dialogue assistance systems have been widely used in online and offline communication scenarios between sales and customers. By providing precise script support and communication guidance for the sales side, they effectively improve sales communication efficiency and business conversion results, becoming a core component of the sales digitalization tool system.

[0024] In related technologies, real-time dialogue assistance systems typically adopt a trigger-based script generation mode. Only after the user presses the trigger button does the system sequentially perform automatic speech recognition (ASR) transcription, semantic parsing, script generation, and other processes, ultimately pushing the generated script to the terminal. Because the functional steps in this mode need to be completed one by one in sequence, the overall time consumption is relatively long, which directly causes the system script push to be delayed and disconnected from the dialogue scenario, thus failing to play the role of real-time dialogue assistance.

[0025] To address the aforementioned issues, this application proposes a real-time dialogue assistance method, system, electronic device, computer-readable storage medium, and computer program product, aiming to solve the problem that real-time dialogue assistance systems cannot function due to delays in script delivery.

[0026] In this embodiment, real-time voice dialogue data is acquired; the dialogue rhythm features of the real-time voice dialogue data are extracted, and a dialogue bottleneck warning signal is generated based on the dialogue rhythm features; in response to the dialogue bottleneck warning signal, a script preloading process is initiated to generate candidate scripts; and dialogue assistance is provided based on the candidate scripts.

[0027] Compared to traditional solutions that only sequentially perform voice ASR transcription, semantic parsing, and script generation after a user triggers an assistance request, ultimately pushing the script to the user's terminal, this application embodiment extracts the dialogue rhythm features of real-time voice dialogue data and generates a dialogue bottleneck warning signal based on these features. In response to this warning signal, a script preloading process is initiated to generate candidate scripts, which are then used for dialogue assistance. Thus, this embodiment eliminates the need to wait for a user to trigger an assistance request before analyzing the voice and generating scripts. Instead, it generates a dialogue bottleneck warning signal based on the dialogue rhythm features and automatically initiates the script preloading process to generate candidate scripts. This effectively avoids delays in script delivery by the real-time dialogue assistance system, reduces the response latency of the real-time voice assistance system, and ultimately solves the problem of the real-time dialogue assistance system failing to function due to delayed script delivery.

[0028] Next, the real-time dialogue assistance method, system, electronic device, computer-readable storage medium, and computer program product provided in this application will be specifically described through the following embodiments, and firstly, the various detailed embodiments of the real-time dialogue assistance method provided in this application will be described in detail.

[0029] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0030] It should be noted that the real-time dialogue assistance method provided in this application relates to the field of voice data processing technology. The real-time dialogue assistance method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be an electronic device such as a smartphone, tablet, laptop, desktop computer, in-vehicle terminal, or a terminal associated with a vehicle that can communicate and interact with the vehicle via a network. The server can be a backend server terminal device of the terminal, which can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The software can be an application implementing the real-time dialogue assistance method, a computer program, and a storage medium carrying the computer program. It should be understood that, based on different design needs of practical applications, the terminal, server, and software of the real-time dialogue assistance method provided in this application may also be other forms not listed here, and the real-time dialogue assistance method provided in this application does not specifically limit these.

[0031] Furthermore, this application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: terminal devices, vehicle terminals, personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, personal computers (PCs), minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via communication networks. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0032] For ease of understanding and explanation, the following text will use the real-time dialogue assistance method provided in the embodiments of this application applied to a terminal device as an example to describe the various specific embodiments of this application in detail. In some descriptions, the terminal device may be simply referred to as a terminal. The implementation of the real-time dialogue assistance method provided in the embodiments of this application for any of the above-described forms of subject matter can refer to the implementation process of the real-time dialogue assistance method described below.

[0033] Please refer to Figure 1 , Figure 1 The flowchart illustrates the steps of the real-time dialogue assistance method provided in some embodiments of this application. It should be understood that, although... Figure 1 The flowcharts illustrating subsequent steps show the execution order of some method steps. However, based on different design needs in practical applications, the real-time dialogue assistance method provided in this application embodiment can, of course, adopt an execution order of method steps different from that shown in the figures. That is, Figure 1 The order of the method steps shown does not constitute a limitation on the execution logic order of the real-time dialogue assistance method provided in the embodiments of this application. Any other order based on... Figure 1 Reasonable changes to the sequence of steps shown should be included within the protection scope of the real-time dialogue assistance method provided in the embodiments of this application.

[0034] like Figure 1 As shown, in some embodiments, the real-time dialogue assistance method provided in this application may include steps S101 to S104 as shown below.

[0035] Step S101: Obtain real-time voice dialogue data.

[0036] It should be noted that real-time voice dialogue data can be audio data generated in offline multi-person (at least two-person) dialogue scenarios. For example, this real-time voice dialogue data is the audio data obtained by collecting the voices of the salesperson and the customer in real time during a face-to-face communication scenario. In addition, real-time voice dialogue data can also be audio data generated in online multi-person dialogue scenarios. For example, this real-time voice dialogue data is the audio data obtained by collecting the voices of the participants / callers in real time during an online conference / telephone scenario.

[0037] When a terminal device performs a real-time dialogue assistance task in a multi-person dialogue scenario, it first acquires the real-time voice dialogue data in that multi-person dialogue scenario.

[0038] In some embodiments, the terminal device can use an embedded or external sound acquisition device (such as a microphone) to collect voice conversations in real time in a multi-person dialogue scenario, thereby obtaining a dialogue recording in that scenario. In this way, the terminal device can use the dialogue recording as real-time voice conversation data.

[0039] Step S102: Extract the dialogue rhythm features of the real-time voice dialogue data, and generate a dialogue bottleneck warning signal based on the dialogue rhythm features.

[0040] It should be noted that the characteristics of dialogue rhythm can include at least the duration of pauses, fluctuations in speech rate, and the number of times speech overlaps.

[0041] After acquiring real-time voice dialogue data, the terminal device further processes the data by extracting rhythm features, thereby obtaining the pause duration, speech rate fluctuations, and number of speech overlaps. For example, the terminal device can use a built-in rhythm feature extraction algorithm to calculate the speech rate (unit: words / second, calculated by dividing the number of text characters by the speech duration), pause duration (calculated by detecting the duration of continuous speech absence, distinguishing between customer pauses and sales pauses), and count the number of overlapping speech per minute (the number of times the salesperson and customer speak simultaneously) and the total duration, thus forming a complete dataset of the dialogue rhythm features of the real-time voice dialogue data.

[0042] After extracting the dialogue rhythm features from real-time voice dialogue data, the terminal device can further analyze the dialogue rhythm based on these features to predict whether a dialogue bottleneck (also known as a communication bottleneck) is about to be encountered in the current multi-person dialogue scenario. If a bottleneck is predicted, a dialogue bottleneck warning signal can be generated. For example, the terminal device can compare the dialogue rhythm features (pause duration, speech rate fluctuations, and the number of speech overlaps) with corresponding thresholds to predict whether a dialogue bottleneck is about to be encountered. That is, if the dialogue rhythm features exceed the corresponding threshold (such as the pause duration exceeding the pause duration warning threshold), the terminal device determines that a dialogue bottleneck is about to be encountered, and thus generates a dialogue bottleneck warning signal.

[0043] In some embodiments, the terminal device may preprocess the real-time voice dialogue data before extracting dialogue rhythm features, and then use a rhythm feature extraction algorithm to extract dialogue rhythm features from the preprocessed real-time voice dialogue data. For example, the terminal device may perform preprocessing operations such as noise reduction and echo removal on the real-time voice dialogue data to remove environmental noise, device noise, and echo interference, thereby ensuring the clarity of the voice data.

[0044] In some embodiments, the terminal device can simultaneously perform preprocessing operations such as noise reduction and echo removal while continuously collecting real-time voice dialogue data, thereby directly obtaining preprocessed real-time voice dialogue data, which can then be directly processed for subsequent extraction of dialogue rhythm features.

[0045] In some embodiments, the terminal device can use a preset data acquisition and preprocessing module to collect real-time voice dialogue data in multi-person dialogue scenarios and simultaneously perform noise reduction and echo removal preprocessing to obtain preprocessed real-time voice dialogue data. Simultaneously, it extracts dialogue rhythm features such as speech rate fluctuations, pause durations, and the number of speech overlaps from this real-time voice dialogue data, forming a complete dataset of dialogue rhythm features. Then, the terminal device can use a dialogue state detection and prediction module to receive the dialogue rhythm features transmitted by the data acquisition and preprocessing module, analyze these features to perform bottleneck prediction, and generate a dialogue bottleneck warning signal.

[0046] Step S103: In response to the dialogue bottleneck warning signal, initiate the dialogue preloading process to generate candidate dialogues.

[0047] It should be noted that candidate dialogues can be guiding or objection-handling dialogues used to assist one or more participants in a multi-person conversation to continue communication with other participants. For example, in a salesperson's conversation with a customer, the candidate dialogue could be, "Please briefly tell me your concerns, whether it's about price or features, and I can answer them in detail to help you find the most suitable solution." Or, the candidate dialogue could be, "Regarding your concern about price, we have a dedicated solution. For the budget needs of small and medium-sized enterprises, we can provide installment payment services. You only need to pay a small monthly fee to enjoy full product features and after-sales support." Furthermore, there can be one or more candidate dialogues.

[0048] After generating a dialogue bottleneck warning signal, the terminal device immediately responds to the warning signal and initiates a script preloading process to generate one or more candidate scripts that are adapted to the current multi-person dialogue scenario and can assist in dialogue communication.

[0049] In some embodiments, when the terminal device initiates the speech preloading process to generate candidate speech, it can combine the current dialogue content (obtained based on speech ASR recognition of real-time voice dialogue data), retrieve relevant knowledge from the knowledge base, and provide the retrieved knowledge to the large model for speech generation, thereby obtaining one or more candidate speech output by the large model.

[0050] In some embodiments, the terminal device can receive a dialogue bottleneck warning signal through a preset dialogue preloading module, and then respond to the dialogue bottleneck warning signal by performing operations to search the dialogue library and generate candidate dialogues through a lightweight large model, thereby obtaining one or more candidate dialogues that are adapted to the current multi-person dialogue scenario and can assist in dialogue communication.

[0051] Step S104: Provide dialogue assistance based on the candidate dialogues.

[0052] After the terminal device initiates the script preloading process to generate candidate scripts, it can proactively push the candidate scripts to users with assistance needs in the current multi-person dialogue scenario to assist them in the dialogue.

[0053] In some embodiments, the terminal device can also respond to instructions / signals triggered by users with assistance needs in the current multi-person dialogue scenario by passively pushing the candidate dialogue to the user to assist them in the dialogue.

[0054] In some embodiments, the terminal device pushes the candidate dialogues generated by the dialogue preloading process to the local cache to build a dialogue cache pool based on the candidate dialogues. This allows the terminal device to actively push candidate dialogues from the dialogue cache pool, or, in the case of passive pushing, to push candidate dialogues from the dialogue cache pool in response to user-triggered instructions / signals.

[0055] In some embodiments, the terminal device can receive candidate dialogues generated by the dialogue preloading module through a preset local cache management module, and construct a local candidate dialogue cache pool to store these preloaded candidate dialogues. Then, the terminal device can use the dialogue push module to extract candidate dialogues from the dialogue cache pool and push them in dual-mode, either actively or passively.

[0056] In some embodiments, when a terminal device pushes candidate dialogues for dialogue assistance, it can stream the pushed dialogues.

[0057] In some embodiments, when a terminal device pushes a script in response to a user-triggered instruction / signal, if there is a lack of candidate scripts in the local cache that are suitable for the current multi-person dialogue scenario and can assist in dialogue communication, an emergency generation process can be triggered to generate scripts in real time and push them to the user to assist in the dialogue.

[0058] In this embodiment, when a terminal device performs a real-time dialogue assistance task in a multi-person dialogue scenario, it first acquires real-time voice dialogue data in the multi-person dialogue scenario. Then, it extracts rhythm features from the real-time voice dialogue data to obtain the dialogue rhythm features of the real-time voice dialogue data. Based on the dialogue rhythm features, it performs dialogue rhythm analysis to predict whether the current multi-person dialogue scenario is about to encounter a dialogue bottleneck. If a bottleneck is predicted, a dialogue bottleneck warning signal is generated. Then, it immediately responds to the dialogue bottleneck warning signal and initiates a script preloading process to generate one or more candidate scripts that are suitable for the current multi-person dialogue scenario and can assist in dialogue communication. Finally, it actively pushes the candidate scripts to users with assistance needs in the current multi-person dialogue scenario to assist them in dialogue through active or passive push methods.

[0059] Compared to traditional solutions that only sequentially perform voice ASR transcription, semantic parsing, and script generation after a user triggers an assistance request, ultimately pushing the script to the user's terminal, this application embodiment does not require analyzing the voice and generating scripts only after the user triggers an assistance request. Instead, it can generate a dialogue bottleneck warning signal based on the dialogue rhythm characteristics, and then automatically start the script preloading process to generate candidate scripts. This effectively avoids the situation where the real-time dialogue assistance system pushes scripts late, reduces the response latency of the real-time voice assistance system, and thus solves the problem that the real-time dialogue assistance system cannot function due to the delay in script push.

[0060] Please refer to Figure 2 , Figure 2 for Figure 1 A detailed flowchart of step S102.

[0061] like Figure 2 As shown, in some embodiments, step S102 above: generating a dialogue bottleneck warning signal based on the dialogue rhythm features may include steps S201 and S202 as shown below.

[0062] Step S201: Compare the dialogue rhythm features with the preset rhythm warning threshold to obtain the first comparison result.

[0063] It should be noted that the preset rhythm warning threshold can be used as a standard to determine whether a multi-person dialogue scenario is about to encounter a dialogue bottleneck. The rhythm warning threshold can include thresholds that correspond one-to-one with the specific types of dialogue rhythm features. That is, when the dialogue rhythm features include pause duration, speech rate fluctuation, and the number of speech overlaps, the rhythm warning threshold can include pause duration warning threshold, speech rate fluctuation warning threshold, and speech overlap number warning threshold.

[0064] When a terminal device performs dialogue rhythm analysis based on dialogue rhythm features to predict whether a dialogue bottleneck is about to occur in a multi-person dialogue scenario, it can compare the dialogue rhythm features with preset rhythm warning thresholds to obtain a first comparison result. For example, the terminal device can compare the pause duration in the dialogue rhythm features with the pause duration warning threshold, compare the speech rate fluctuation in the dialogue rhythm features with the speech rate fluctuation warning threshold, and / or compare the number of speech overlaps in the dialogue rhythm features with the number of speech overlaps warning threshold to obtain a first comparison result.

[0065] Step S202: If the first comparison result shows that the dialogue rhythm feature is greater than or equal to the rhythm warning threshold, a dialogue bottleneck warning signal is generated.

[0066] After receiving the first comparison result, if the first comparison result indicates that the dialogue rhythm feature is greater than or equal to the corresponding rhythm warning threshold, the terminal device determines that the current multi-person dialogue scenario is about to encounter a dialogue bottleneck, and thus generates a dialogue bottleneck warning signal to initiate the script preloading process to generate candidate scripts. For example, if the first comparison result indicates that the pause duration is greater than or equal to the pause duration warning threshold, the speech rate fluctuation is greater than or equal to the speech rate fluctuation warning threshold, and / or the number of speech overlaps is greater than or equal to the speech overlap number warning threshold, the terminal device determines that the current multi-person dialogue scenario is about to encounter a dialogue bottleneck, and thus generates a dialogue bottleneck warning signal to initiate the script preloading process to generate candidate scripts.

[0067] For example, in a scenario where a salesperson and a customer are communicating face-to-face, the dialogue rhythm characteristics of the real-time voice dialogue data collected by the terminal device include: pause duration specifically refers to the customer's silence duration, speech rate fluctuation specifically refers to the salesperson's speech rate fluctuation, and speech overlap frequency refers to the number of times the salesperson and the customer's speech overlaps per minute. Corresponding to these dialogue rhythm characteristics, the pause duration warning threshold is the same as the customer silence duration warning threshold (e.g., 1.4 seconds), the speech rate fluctuation warning threshold is 21% (compared to the salesperson's average speech rate), and the speech overlap frequency warning threshold is 2 times. Thus, if the customer's silence duration is ≥1.4 seconds, the salesperson's speech rate fluctuation is ≥21%, and / or the speech overlap frequency is ≥2 times, the terminal device determines that the salesperson is about to encounter a communication bottleneck in the current scenario, thereby generating a dialogue bottleneck warning signal to initiate the script preloading process and begin generating candidate scripts.

[0068] In some embodiments, after obtaining the first comparison result, if the first comparison result indicates that the dialogue rhythm feature is less than the corresponding rhythm warning threshold, the terminal device will continue to monitor and will not initiate the preloading process, thereby reducing the consumption of system resources.

[0069] In this embodiment, when the terminal device analyzes the dialogue rhythm based on dialogue rhythm features to predict whether a dialogue bottleneck is about to be encountered in a multi-person dialogue scenario, it compares the dialogue rhythm features with a preset rhythm warning threshold to obtain a first comparison result. Then, if the first comparison result indicates that the dialogue rhythm features are greater than or equal to the corresponding rhythm warning threshold, it is determined that the current multi-person dialogue scenario is about to encounter a dialogue bottleneck, thereby generating a dialogue bottleneck warning signal to initiate the script preloading process and begin generating candidate scripts. Thus, compared to traditional solutions, this embodiment can, based on real-time extraction of rhythm features such as the speech rate and pause duration of each speaker, and combined with warning thresholds, detect potential communication bottlenecks between speakers in advance and generate warning signals, providing accurate basis for script preloading. This effectively solves the problem of traditional solutions lacking predictive ability and failing to detect communication difficulties in advance.

[0070] Please refer to Figure 3 , Figure 3 The following are schematic flowcharts of the steps in some other embodiments of the real-time dialogue assistance method provided in the embodiments of this application.

[0071] like Figure 3 As shown, in some embodiments, the real-time dialogue assistance method provided in this application may further include steps S301 and S302 as shown below.

[0072] Step S301: Compare the dialogue rhythm features with a preset rhythm trigger threshold to obtain a second comparison result; the rhythm trigger threshold is greater than the rhythm warning threshold.

[0073] It should be noted that the preset rhythm trigger threshold can be a standard based on the analysis of rhythm features to trigger proactively pushed dialogue for dialogue assistance. This rhythm trigger threshold can also include thresholds that correspond one-to-one with specific types of dialogue rhythm features. That is, when dialogue rhythm features include pause duration, speech rate fluctuations, and the number of speech overlaps, the rhythm trigger threshold can include pause duration trigger thresholds, speech rate fluctuation trigger thresholds, and the number of speech overlaps trigger thresholds. However, the rhythm trigger threshold is larger than the rhythm warning threshold; for example, the rhythm warning threshold can be 70% of the rhythm trigger threshold.

[0074] When a terminal device performs dialogue rhythm analysis based on dialogue rhythm features to predict whether a dialogue bottleneck is about to occur in a multi-person dialogue scenario, it can also compare the dialogue rhythm features with preset rhythm trigger thresholds to obtain a second comparison result. For example, the terminal device can compare the pause duration in the dialogue rhythm features with the pause duration trigger threshold, compare the speech rate fluctuation in the dialogue rhythm features with the speech rate fluctuation trigger threshold, and / or compare the number of speech overlaps in the dialogue rhythm features with the number of speech overlaps trigger threshold, thereby obtaining a second comparison result.

[0075] Step S302: If the second comparison result is that the dialogue rhythm feature is greater than or equal to the preset rhythm trigger threshold, a dialogue assistance trigger signal is generated.

[0076] After receiving the second comparison result, if the second comparison result indicates that the dialogue rhythm feature is greater than or equal to the corresponding rhythm trigger threshold, the terminal device determines that the current multi-person dialogue scenario needs to trigger proactive dialogue to assist in the dialogue, thereby generating a dialogue assistance trigger signal. For example, if the second comparison result indicates that the pause duration is greater than or equal to the pause duration trigger threshold, the speech rate fluctuation is greater than or equal to the speech rate fluctuation trigger threshold, and / or the number of speech overlaps is greater than or equal to the speech overlap number trigger threshold, the terminal device determines that proactive dialogue needs to be triggered to assist in the dialogue, thereby generating a dialogue assistance trigger signal.

[0077] For example, in a scenario where a salesperson and a customer are communicating face-to-face, the dialogue rhythm characteristics of the real-time voice dialogue data collected by the terminal device include: pause duration specifically refers to the customer's silence duration, speech rate fluctuation specifically refers to the salesperson's speech rate fluctuation, and speech overlap frequency refers to the number of times the salesperson and customer's speech overlaps per minute. The corresponding rhythm trigger thresholds for these dialogue rhythm characteristics are: pause duration trigger threshold (e.g., 2 seconds), speech rate fluctuation trigger threshold (30% compared to the salesperson's average speech rate), and speech overlap frequency trigger threshold (3 times). Thus, if the customer's silence duration is ≥2 seconds, the salesperson's speech rate fluctuation is ≥30%, and / or the speech overlap frequency is ≥3 times, the terminal device determines that a script needs to be pushed to the salesperson for dialogue assistance in the current scenario, thereby generating a dialogue assistance trigger signal.

[0078] In some embodiments, the terminal device can receive dialogue rhythm features and perform rhythm analysis through the dialogue state detection and prediction module described above. In the process of completing the bottleneck prediction operation to generate a dialogue bottleneck warning signal, it can simultaneously determine whether the current multi-person dialogue scenario needs to trigger active push of dialogue words for dialogue assistance, and generate a dialogue assistance trigger signal if the determination is yes.

[0079] In some embodiments, step S104 above: providing dialogue assistance based on the candidate dialogues, may include the following steps: In response to the dialogue assistance trigger signal, the candidate dialogues are actively pushed to assist in the dialogue.

[0080] After generating a dialogue assistance trigger signal, the terminal device immediately responds to the dialogue assistance trigger signal and executes an active push mode. It actively pushes the candidate dialogues generated by the dialogue script preloading process that was initiated when responding to the dialogue bottleneck warning signal to users who have assistance needs in the current multi-person dialogue scenario, so as to provide them with dialogue assistance.

[0081] In some embodiments, the terminal device may use the above-mentioned dialogue push module to receive the dialogue assistance trigger signal generated by the dialogue state detection and prediction module, and then respond to the dialogue assistance trigger signal to perform active push, pushing the candidate dialogues in the local candidate dialogue cache pool built by the local cache management module to the user with assistance needs in the current multi-person dialogue scenario in a streaming output mode, so as to provide dialogue assistance.

[0082] In this embodiment, the terminal device further compares the dialogue rhythm features with a preset rhythm trigger threshold to obtain a second comparison result. Then, if the second comparison result shows that the dialogue rhythm features are greater than or equal to the corresponding rhythm trigger threshold, it is determined that the current multi-person dialogue scenario needs to trigger the active push of dialogue scripts for dialogue assistance, thereby generating a dialogue assistance trigger signal. Then, in response to the dialogue assistance trigger signal, the active push mode is executed, and the candidate dialogue scripts generated by the dialogue script preloading process that was initiated when responding to the dialogue bottleneck warning signal are actively pushed to users with assistance needs in the current multi-person dialogue scenario to provide them with dialogue assistance. Thus, compared to traditional solutions, this embodiment can extract rhythmic features such as the speaking speed and pause duration of each speaker in real time, and combine them with multi-level thresholds (early warning thresholds and trigger thresholds) to detect potential communication bottlenecks between speakers in advance, generating early warning signals to provide accurate basis for script preloading, and generating trigger signals to provide accurate basis for proactive script delivery. This enables the real-time dialogue assistance system to have proactive prediction and intervention capabilities, upgrading from passive response to proactive prediction and intervention (proactively predicting the risks of dialogue communication and proactively intervening to assist), thereby fully leveraging the dialogue assistance role, avoiding communication bottlenecks in advance, effectively avoiding dialogue communication failures, and improving communication conversion rates.

[0083] Please refer to Figure 4 , Figure 4 A flowchart illustrating the steps of the real-time dialogue assistance method provided in some other embodiments of this application.

[0084] like Figure 4As shown, in some embodiments, the real-time dialogue assistance method provided in this application may further include steps S401 and S402 as shown below.

[0085] Step S401: Perform semantic analysis on the dialogue text data corresponding to the real-time voice dialogue data to obtain the semantic analysis results.

[0086] After acquiring real-time voice dialogue data, the terminal device can preprocess the real-time voice dialogue data and convert it into dialogue text data in real time, thereby performing in-depth semantic analysis on the dialogue text data and obtaining semantic analysis results.

[0087] In some embodiments, the terminal device can perform real-time semantic analysis of the dialogue text data based on a large model, and at the same time combine the speaker's tags and the current dialogue stage (predicted from historical dialogue data and current text) to classify the speaker's intent and score the dialogue content, thereby obtaining semantic analysis results.

[0088] For example, in a scenario where a salesperson is conversing with a customer, the terminal device can perform real-time semantic analysis of the dialogue text data based on a large model. Simultaneously, it can combine customer tags (such as budget, focus areas, etc.) and the current sales stage (predicted from historical dialogue data and current text, such as needs assessment, objection handling, test drives, etc.) to complete customer intent classification, customer intention scoring, and sales response quality scoring. In other words, the terminal device can provide the following input data to the large model: "Based on the following dialogue text, customer tags and sales stage, complete the customer intent classification, intent rating and sales response quality rating. The rating is based on a 1-10 point scale, with higher scores indicating stronger intent / better response quality." Speaker 1 (Customer): This product is priced much higher than the competitors, which is a bit beyond our budget; Speaker 2 (Sales): Our products are of better quality and have a high cost-performance ratio; Speaker 1 (Customer): The quality is indeed good, but the price is still a bit expensive. I'll think about it some more.

[0089] Customer profile: Limited budget, focused on cost-effectiveness; Current sales stage: Objection handling stage (price objection). The semantic analysis results are as follows after the large model performs deep semantic analysis: Customer intent can be categorized as: price objections, expressing hesitation to buy; Customer intention rating (also known as customer willingness to communicate rating): 5 points (1-10 points, the higher the score, the stronger the intention). Sales response quality rating: 3 points (1-10 points, failed to specifically address the customer's budget issue, response too general). Step S402: Generate a dialogue bottleneck warning signal based on the semantic analysis results.

[0090] After obtaining the semantic analysis results, the terminal device can further predict whether the current multi-person dialogue scenario is about to encounter a dialogue bottleneck based on the semantic analysis results, and generate a dialogue bottleneck warning signal if a bottleneck is predicted to be encountered.

[0091] In some embodiments, the terminal device can combine information such as intent and objection type in the semantic analysis results to predict whether a dialogue bottleneck is about to be encountered.

[0092] Please refer to Figure 5 , Figure 5 for Figure 4 A detailed flowchart of step S402.

[0093] like Figure 5 As shown, in some embodiments, when the semantic analysis result includes a dialogue score, the above step S402: generating a dialogue bottleneck warning signal based on the semantic analysis result may include steps S501 and S502 as shown below.

[0094] Step S501: Compare the dialogue score with a preset first score threshold to obtain a third comparison result.

[0095] It should be noted that dialogue scoring can include intention scoring and response quality scoring. For example, in a scenario where a salesperson is talking to a customer, the intention scoring can be the customer intention score output by the aforementioned large model, while the response quality scoring can be the sales response quality score output by the large model. Furthermore, the first scoring threshold can also serve as a criterion for determining whether a dialogue bottleneck is about to be encountered in a multi-person conversation. This first scoring threshold can include thresholds that correspond one-to-one with the specific type of dialogue scoring; that is, when dialogue scoring includes intention scoring and response quality scoring, the first scoring threshold can include an intention scoring warning threshold and a response quality scoring warning threshold.

[0096] When a terminal device predicts whether it is about to encounter a dialogue bottleneck based on the semantic analysis results, it can compare the dialogue score in the semantic analysis results with the corresponding first scoring threshold to obtain a third comparison result. For example, the terminal device can compare the intention score in the dialogue score with the intention score warning threshold, and / or compare the response quality score in the dialogue score with the response quality score warning threshold to obtain a third comparison result.

[0097] Step S502: If the third comparison result shows that the dialogue score is less than the first score threshold, a dialogue bottleneck warning signal is generated.

[0098] After receiving the third comparison result, if the dialogue score is lower than the corresponding first scoring threshold, the terminal device determines that the current multi-person dialogue scenario is about to encounter a dialogue bottleneck, and generates a dialogue bottleneck warning signal to initiate the script preloading process to generate candidate scripts. For example, if the third comparison result is that the intention score is lower than the intention score warning threshold, and / or the response quality score is lower than the response quality score warning threshold, the terminal device determines that the current multi-person dialogue scenario is about to encounter a dialogue bottleneck, and generates a dialogue bottleneck warning signal to initiate the script preloading process to generate candidate scripts.

[0099] For example, in a scenario where a salesperson and a customer are communicating face-to-face, after the terminal device performs semantic analysis on the dialogue text data corresponding to the real-time voice dialogue data, the dialogue score in the semantic analysis result includes a customer intention score and a sales response quality score. Among the first scoring thresholds corresponding to these dialogue scores, the intention score warning threshold is the customer intention score warning threshold (e.g., 3 points), and the response quality score warning threshold is the sales response quality score warning threshold (e.g., 4 points). Thus, when the sales response quality score is lower than 4 points (e.g., 3 points in the example above), and / or the customer intention score is lower than 3 points, the terminal device determines that the salesperson is about to encounter a communication dilemma in the current scenario, thereby generating a dialogue bottleneck warning signal to initiate the script preloading process and begin generating candidate scripts.

[0100] In some embodiments, after obtaining the third comparison result, if the third comparison result is that the dialogue score is greater than or equal to the first score threshold, the terminal device will continue to monitor and will not start the preloading process, thereby reducing the consumption of system resources.

[0101] In this embodiment, the terminal device compares the dialogue score in the semantic analysis result with the corresponding first scoring threshold to obtain a third comparison result. Then, if the third comparison result shows that the dialogue score is less than the corresponding first scoring threshold, it is determined that the current multi-person dialogue scenario is about to encounter a dialogue bottleneck, thereby generating a dialogue bottleneck warning signal to initiate the script preloading process and generate candidate scripts. Thus, compared to traditional solutions, this embodiment can achieve dual-dimensional prediction of dialogue rhythm and semantic fusion. That is, in addition to extracting rhythmic features such as the salesperson's and customer's speaking speed and pause duration in real time to detect potential bottlenecks in advance, it can also combine information such as customer intentions and objection types obtained from semantic analysis to detect potential communication bottlenecks between the speakers in advance and generate warning signals. This provides accurate basis for script preloading and effectively solves the problem of traditional solutions lacking predictive capabilities and failing to detect communication difficulties in advance.

[0102] Please refer to Figure 6 , Figure 6 for Figure 1 A detailed flowchart of step S104.

[0103] like Figure 6 As shown, in some embodiments, step S104 above: providing dialogue assistance based on the candidate dialogue, includes at least one of steps S601 and S602 as shown below.

[0104] Step S601: If the dialogue score is less than a preset second score threshold, actively push the candidate dialogue to assist in the dialogue; the second score threshold is less than the first score threshold.

[0105] It should be noted that the second scoring threshold can be a standard for triggering proactive push messages for dialogue assistance based on semantic depth analysis. This second scoring threshold can also include thresholds corresponding one-to-one with specific types of dialogue scoring. That is, when dialogue scoring includes intention scoring and response quality scoring, the second scoring threshold can include intention scoring trigger thresholds and response quality scoring trigger thresholds. However, the second scoring threshold is smaller than the first scoring threshold. For example, if the intention scoring warning threshold in the first scoring threshold is 3 points (customer intention scoring warning threshold) and the response quality scoring warning threshold is 4 points (sales response quality scoring warning threshold), then in the second scoring threshold, the customer intention scoring warning threshold can be 2 points and the sales response quality scoring warning threshold can be 3 points.

[0106] When a terminal device predicts whether it is about to encounter a dialogue bottleneck based on the semantic analysis results, it can also compare the dialogue score in the semantic analysis results with the corresponding second score threshold. If the dialogue score is less than the corresponding second score threshold, it determines that the current multi-person dialogue scenario needs to trigger proactive push of dialogue scripts for dialogue assistance. It then immediately executes the proactive push mode, proactively pushing the candidate dialogue scripts generated by the dialogue script preloading process that was initiated in response to the dialogue bottleneck warning signal to the users in the current multi-person dialogue scenario who have assistance needs, so as to provide them with dialogue assistance.

[0107] For example, in a scenario where a salesperson and a customer are communicating face-to-face, if the sales response quality score is below 3 points and / or the customer intention score is below 2 points, the terminal device determines that a script needs to be pushed to the salesperson to assist them in the conversation. Thus, it executes a proactive push mode, which proactively pushes the candidate scripts generated by the script preloading process that was initiated when responding to the conversation bottleneck warning signal to the user with assistance needs in the current multi-person conversation scenario to assist them in the conversation.

[0108] Step S602: If the number of times the dialogue bottleneck warning signal is generated is greater than or equal to a preset threshold, actively push the candidate dialogue to assist in the dialogue.

[0109] The terminal device can continuously perform semantic analysis on the dialogue text data of the constantly changing real-time voice dialogue data based on a large model, and predict whether a dialogue bottleneck is about to be encountered based on the semantic analysis results output by the large model in real time. In this process, it counts the number of times a dialogue score is determined to be less than the corresponding first score threshold in the third comparison result, thus generating a dialogue bottleneck warning signal. If the number is greater than or equal to a preset threshold (e.g., 3 times), the terminal device can also determine that the current multi-person dialogue scenario needs to trigger proactive push of dialogue scripts for dialogue assistance. In this case, it will also execute the proactive push mode, and proactively push the candidate dialogue scripts generated by the dialogue script preloading process that was initiated in response to the dialogue bottleneck warning signal to users in the current multi-person dialogue scenario who have assistance needs, so as to provide them with dialogue assistance.

[0110] In this embodiment, by using terminal devices to make predictions based on a dual dimension of dialogue rhythm analysis and semantic depth analysis, a prediction model can be constructed by integrating dialogue rhythm features and semantic analysis results. This model can extract rhythm features such as the salesperson's and customer's speaking speed and pause duration in real time, preset multi-level thresholds to detect potential communication bottlenecks in advance and trigger proactive script pushes. It can also combine information such as customer intentions and objection types obtained from semantic analysis to detect potential communication bottlenecks in advance and trigger proactive script pushes. This can effectively solve the problems of traditional solutions lacking prediction capabilities, being unable to detect communication difficulties in advance, and being unable to prepare appropriate scripts and proactively push scripts for auxiliary intervention.

[0111] Please refer to Figure 7 , Figure 7 A flowchart illustrating the steps of the real-time dialogue assistance method provided in some other embodiments of this application.

[0112] like Figure 7 As shown, in some embodiments, when the number of candidate dialogues is greater than one, the real-time dialogue assistance method provided in this application embodiment may further include the following step S701 after step S103: in response to the dialogue bottleneck warning signal, starting the dialogue preloading process to generate candidate dialogues.

[0113] Step S701: Adjust the push weight of each of the multiple candidate dialogues based on the semantic analysis results to obtain multiple candidate dialogues after adjusting the push weight.

[0114] It should be noted that the push weight is used to characterize the likelihood that a candidate statement will be pushed for dialogue assistance. The higher the push weight, the greater the likelihood that the candidate statement will be pushed.

[0115] After generating a dialogue bottleneck warning signal based on the semantic analysis results to initiate the dialogue preloading process and generate multiple candidate dialogues, the terminal device can further adjust the push weight of each of the multiple candidate dialogues based on the semantic analysis results to obtain multiple candidate dialogues with adjusted push weights.

[0116] In some embodiments, the terminal device can comprehensively evaluate the degree of fit between each candidate dialogue and the real-time voice dialogue data in the current dialogue scenario based on information such as intention score and objection type in the semantic analysis results, and adjust the push weight of each candidate dialogue accordingly based on the degree of fit. For example, if the degree of fit of candidate dialogue 1 in answering a customer's price objection is higher than that of candidate dialogue 2 in answering a customer's price objection, then the terminal device can adjust the push weight of candidate dialogue 1 to be higher than that of candidate dialogue 2.

[0117] In some embodiments, step S104 above: providing dialogue assistance based on the candidate dialogues, may include the following steps: The target candidate dialogue is pushed from among the multiple candidate dialogues to assist in the dialogue; the push weight of the target candidate dialogue is higher than the push weight of other candidate dialogues among the multiple candidate dialogues excluding the target candidate dialogue.

[0118] After the terminal device adjusts the push weight of each of the multiple candidate dialogues, whether it actively pushes dialogues for dialogue assistance or passively pushes dialogues for dialogue assistance, it can select the target candidate dialogue with the highest push weight from the multiple candidate dialogues, and then push the target candidate dialogue to the user with assistance needs in the current multi-person dialogue scenario to assist them in the dialogue.

[0119] In some embodiments, the terminal device may adjust the push weights of multiple candidate dialogues, then further sort the candidate dialogues according to the adjusted push weights, and then, when a dialogue needs to be pushed, select the target candidate dialogue with the highest push weight from the sorted list for push. For example, the terminal device may sort multiple candidate dialogues in descending order of push weight, and then select the candidate dialogue ranked first as the target candidate dialogue for push.

[0120] In some embodiments, the terminal device can preprocess real-time voice dialogue data through the aforementioned data acquisition and preprocessing module, and then convert it into dialogue text data in real time. Subsequently, the dialogue state detection and prediction module receives the dialogue text data converted by the data acquisition and preprocessing module, performs real-time semantic analysis, and further predicts whether the current multi-person dialogue scenario is about to encounter a dialogue bottleneck based on the semantic analysis results. If a bottleneck is predicted, a dialogue bottleneck warning signal is generated, which is transmitted to the script preloading module to generate candidate scripts. Simultaneously, the dialogue state detection and prediction module transmits information such as customer intention score and objection type from the semantic analysis results to the local cache management module. The local cache management module can then adjust the push weights of the multiple candidate scripts generated by the script preloading module based on this customer intention score and objection type information. Finally, the script push module, either through an active push mode or a passive push mode, pushes the adjusted push weights of the candidate scripts from the local candidate script cache pool built by the local cache management module in a streaming output mode to users with assistance needs in the current multi-person dialogue scenario, providing them with dialogue assistance.

[0121] In this embodiment, the terminal device adjusts the push weights of multiple candidate dialogues based on semantic analysis results, resulting in multiple candidate dialogues with adjusted push weights. Then, whether the dialogue is actively or passively pushed for dialogue assistance, the target candidate dialogue with the highest push weight can be selected from these multiple candidate dialogues and pushed to users with assistance needs in the current multi-person dialogue scenario. This effectively improves dialogue adaptability and enhances the value of real-time dialogue assistance.

[0122] Please refer to Figure 8 , Figure 8 A flowchart illustrating the steps of the real-time dialogue assistance method provided in some other embodiments of this application.

[0123] like Figure 8 As shown, in some embodiments, the real-time dialogue assistance method provided in this application embodiment may further include the following step S801 after step S103: in response to the dialogue bottleneck warning signal, starting the dialogue preloading process to generate candidate dialogues.

[0124] Step S801: Generate new candidate dialogue in response to a topic switching signal; the topic switching signal is generated when a topic switch occurs in the real-time voice dialogue data.

[0125] After acquiring real-time voice dialogue data, the terminal device can also identify changes in the dialogue topic in real time based on this data, and generate a topic switching signal when a change in the dialogue topic is detected. Then, if the terminal device generates the topic switching signal after the script preloading process has already started to generate candidate scripts, it will immediately respond to the topic switching signal and regenerate new candidate scripts.

[0126] In some embodiments, if the terminal device has not yet started the script preloading process to generate candidate scripts when generating a topic switching signal (e.g., it has not determined that a dialogue bottleneck is about to be encountered through rhythm analysis and / or semantic analysis, and therefore has not generated a dialogue bottleneck warning signal), then it will still not start the script preloading process to avoid consuming system resources.

[0127] In some embodiments, the terminal device can use a large model to take the dialogue text data corresponding to the real-time voice dialogue data as input, identify changes in the dialogue topic in real time, and then generate a topic switching signal when the recognition result output by the large model indicates that a topic switch has occurred.

[0128] In some embodiments, when the terminal device identifies whether the conversation topic has changed based on a large model, it can combine a topic library (such as topic tags containing product features, prices, after-sales service, competitor comparisons, etc.) related to the conversation subject (such as products recommended by sales personnel) to determine whether the current conversation topic has switched.

[0129] For example, in a scenario where a salesperson is conversing with a customer, the terminal device can extract a portion of historical dialogue fragments from real-time voice dialogue data. This historical dialogue fragment, along with a product topic library, is input into a large model. The large model then uses this historical dialogue fragment and the product topic library to identify the current dialogue topic tag, determine if a topic switch has occurred, and output a topic switch indicator and the new topic type. In other words, the input data from the terminal device to the large model can be as follows: "Based on historical dialogue fragments and product topic database, identify the current dialogue topic tags, determine whether a topic switch has occurred, and output the topic switch indicator and new topic type; " Excerpt from a historical dialogue: The first 3 minutes mainly discussed product features and labor cost savings; Product Topic Library: Features, Pricing, After-Sales Service, Competitor Comparison, Sales Policy, ... Speaker 2 (Sales): The core function of this product is automated management, which can help you save 50% of your labor costs; Speaker 1 (Customer): The features sound great, but what kind of after-sales service do you provide? How long does it take to resolve any problems? Thus, after identifying whether the topic of the conversation has changed based on the input data, the large model can output the following recognition result: Topic switching indicator: Yes; New topic type: After-sales guarantee. In this way, the terminal device can generate a topic switching signal based on the topic switch indicated by the recognition result. The topic switching signal can indicate that the current new topic type is after-sales guarantee, which makes it easier to regenerate candidate scripts related to the topic of after-sales guarantee.

[0130] In some embodiments, when the terminal device regenerates new candidate dialogue in response to a topic switching signal, it can also clear the candidate dialogue that has already been generated and cached, and then re-cache the newly generated candidate dialogue for use in subsequent push notifications.

[0131] In some embodiments, step S104 above: providing dialogue assistance based on the candidate dialogues, may include the following steps: The new candidate phrases are pushed to assist in the conversation.

[0132] After the terminal device generates new candidate dialogue in response to a topic switching signal, whether it actively or passively pushes the dialogue to assist the conversation, it can push the new candidate dialogue to users in the current multi-person conversation scenario who need assistance, in order to provide them with conversation support. At this time, if there are multiple new candidate dialogues, the target candidate dialogue with the highest push weight can still be selected from these multiple candidate dialogues (it can be selected according to the order of push weight before adjustment or according to the order of push weight after adjustment), and then the target candidate dialogue can be pushed.

[0133] In some embodiments, the terminal device can use the aforementioned dialogue state detection and prediction module to identify changes in the dialogue topic in real time (e.g., based on a large model, combined with product-related topic tags such as product functions, prices, after-sales service, and competitor comparisons in a product-related topic library to determine whether the current dialogue topic has switched). Upon detecting a topic switch, it generates a topic switch signal and synchronously transmits the signal and the new topic type to the script preloading module and the local cache management module. If the script preloading module has already responded to the dialogue bottleneck warning signal and started generating candidate scripts, it immediately stops loading the original topic scripts and starts preloading the new topic scripts to regenerate new candidate scripts suitable for the new topic type. The local cache management module clears the original cached scripts and synchronously caches the new candidate scripts. Finally, the script push module, still following an active or passive push mode, pushes the newly generated candidate scripts from the local candidate script cache pool built by the local cache management module in a streaming output to users with assistance needs in the current multi-person dialogue scenario to provide dialogue assistance.

[0134] In this embodiment, the terminal device identifies changes in the conversation topic in real time based on real-time voice dialogue data. Upon detecting a change in the conversation topic, a topic switching signal is generated, and then new candidate dialogue is regenerated in response to this signal. Subsequently, whether the dialogue is actively or passively pushed for dialogue assistance, this new candidate dialogue can be pushed to users in the current multi-person conversation scenario who require assistance. This effectively avoids the preloading and caching of invalid dialogue unrelated to the current conversation topic, thereby improving the efficiency and adaptability of dialogue preloading.

[0135] Next, a complete embodiment of the real-time dialogue assistance method provided in this application will be presented.

[0136] In some embodiments, the terminal device can execute the real-time dialogue assistance method provided in this application embodiment through a real-time dialogue assistance system. This real-time dialogue assistance system may include the aforementioned data acquisition and preprocessing module, dialogue state detection and prediction module, script preloading module, local cache management module, and script push module. In a scenario where sales personnel and customers are conversing, the data acquisition and preprocessing module can be used to collect real-time voice dialogue data between sales personnel and customers, simultaneously perform noise reduction and echo cancellation preprocessing, and convert it into clear text data; at the same time, it extracts dialogue rhythm features such as speech rate and pause duration to form a complete dataset, which is then transmitted to subsequent modules; the dialogue state detection and prediction module can integrate three sub-modules: rhythm analysis, semantic depth analysis, and topic switching judgment, thereby completing bottleneck prediction, semantic parsing, and topic recognition based on preprocessed data, generating warning signals, triggering warning signals, and topic switching signals; the script preloading module can receive warning signals... The system retrieves the script library based on the account and topic information, generates candidate scripts using a lightweight large model, and pushes them to the local cache. When a topic switch signal is received, the preloaded content is adjusted synchronously. The local cache management module can build a local candidate script cache pool to store preloaded scripts. When a topic switch signal is received, the original cache is cleared and the new topic scripts are refreshed to ensure that the cached content is consistent with the current scenario. The script push module can push scripts in two modes: listening for sales passive trigger signals and instantly retrieving cached scripts, and receiving trigger warning signals and actively pushing scripts. Scripts are output in streaming mode, and an emergency generation process is triggered when the cache is missing.

[0137] Please refer to Figure 9 , Figure 9 The real-time dialogue assistance method provided in the embodiments of this application is illustrated in some system flow diagrams.

[0138] In some embodiments, the terminal device can use the real-time dialogue assistance system described above, according to... Figure 9 The workflow shown is used to execute the real-time dialogue assistance method provided in the embodiments of this application. This workflow can be specifically embodied in steps 1 to 4 as shown below.

[0139] Step 1: Real-time data acquisition and preprocessing.

[0140] After the system starts, the data acquisition and preprocessing module continuously and synchronously collects real-time voice dialogue data between sales personnel and customers. During the acquisition process, noise reduction and echo removal preprocessing operations are performed simultaneously to remove environmental noise, equipment noise, and echo interference, ensuring the clarity of the voice data. The preprocessed voice data is then converted into text data in real time.

[0141] Meanwhile, the module's built-in rhythm feature extraction algorithm calculates in real time the salesperson's and customer's speech rate (unit: words / second, calculated by dividing the number of text characters by the speech duration), pause duration (calculated by detecting the duration of continuous no speech signal, distinguishing between customer pauses and salesperson pauses), counts the number of overlapping speech per minute (the number of times the salesperson and customer speak at the same time) and the total duration, forming a complete dialogue rhythm feature dataset.

[0142] The processed dialogue text data (labeled with speakers and timestamps) and rhythm feature dataset are transmitted in real time to the dialogue state detection and prediction module.

[0143] Step 2, Dialogue State Detection and Prediction.

[0144] The dialogue state detection and prediction module predicts upcoming communication bottlenecks in sales by analyzing the dialogue rhythm, semantic depth, and topic switching in real time, generating early warning signals and topic switching signals to provide a basis for script preloading. It is divided into three sub-steps: 2.1 to 2.3.

[0145] 2.1 Analysis of Dialogue Rhythm.

[0146] The dialogue rhythm analysis submodule receives a real-time transmitted rhythm feature dataset and monitors changes in dialogue rhythm in real time based on preset communication bottleneck trigger thresholds and warning thresholds. Dialogue assistance trigger thresholds: customer silence duration ≥ 2 seconds, sales speech rate fluctuation ≥ 30% (compared to the salesperson's own average speech rate), and overlapping speech ≥ 3 times per minute. Reaching any of these thresholds will be considered a communication bottleneck that requires proactive intervention.

[0147] Dialogue bottleneck warning threshold: 70% of the trigger threshold, i.e., customer silence duration ≥ 1.4 seconds, sales speech rate fluctuation ≥ 21%, and overlapping speech ≥ 2 times per minute. A dialogue bottleneck warning signal will be generated when any of these thresholds are reached.

[0148] When the conversation pace reaches the warning threshold, a conversation bottleneck warning signal is generated, predicting that the salesperson will encounter a communication difficulty. The warning signal is then transmitted to the script preloading module to initiate the script preloading process. When the conversation pace reaches the trigger threshold, a conversation assistance trigger signal is generated and transmitted synchronously to the script push module to provide a trigger basis for the proactive push intervention mode. If the conversation pace does not reach the warning threshold, monitoring continues, and the preloading process is not initiated to reduce system resource consumption.

[0149] 2.2, Semantic Deep Analysis.

[0150] The semantic deep analysis submodule performs real-time semantic parsing of dialogue text data based on a large model. It also combines customer tags (such as budget, focus, etc.) and the current sales stage (predicted from historical dialogue data and current text, such as demand mining, objection handling, test drive, etc.) to complete customer intent classification, customer intention scoring, and sales response quality scoring, as detailed below.

[0151] enter: "Based on the following dialogue text, customer tags and sales stage, complete the customer intent classification, intent rating and sales response quality rating. The rating is based on a 1-10 point scale, with higher scores indicating stronger intent / better response quality." Speaker 1 (Customer): This product is priced much higher than the competitors, which is a bit beyond our budget; Speaker 2 (Sales): Our products are of better quality and have a high cost-performance ratio; Speaker 1 (Customer): The quality is indeed good, but the price is still a bit expensive. I'll think about it some more. Customer profile: Limited budget, focused on cost-effectiveness; Current sales stage: Objection handling stage (price objection). Output: Customer intent can be categorized as: price objections, expressing hesitation to buy; Customer intention rating: 5 points (1-10 points, the higher the score, the stronger the intention). Sales response quality rating: 3 points (1-10 points, failed to specifically address the customer's budget issue, response too general). Based on the results output by the aforementioned large model, when the sales response quality score is below 4 points (such as 3 points in the above output), or the customer communication willingness score is below 3 points and the intention score continues to decline (such as the customer intention dropping from 7 points to 5 points and then to 3 points), a dialogue bottleneck warning signal is generated. This indicates that the salesperson is unable to effectively handle customer objections and needs script assistance. The warning signal is then transmitted to the script preloading and weight ranking module to supplement the preloaded and adapted objection handling scripts. At the same time, information such as the customer intention score and objection type is transmitted synchronously to provide a basis for subsequent script weight adjustments.

[0152] 2.3, Topic switching judgment.

[0153] The topic switching judgment submodule identifies changes in the conversation topic in real time through a large model. Specifically, it combines a product-related topic library (such as topic tags like product features, price, after-sales service, and competitor comparison) to determine whether the current conversation topic has switched, as detailed below.

[0154] enter: "Based on historical dialogue fragments and product topic database, identify the current dialogue topic tags, determine whether a topic switch has occurred, and output the topic switch indicator and new topic type; " Excerpt from a historical dialogue: The first 3 minutes mainly discussed product features and labor cost savings; Product Topic Library: Features, Pricing, After-Sales Service, Competitor Comparison, Sales Policy, ... Speaker 2 (Sales): The core function of this product is automated management, which can help you save 50% of your labor costs; Speaker 1 (Customer): The features sound great, but what kind of after-sales service do you provide? How long does it take to resolve any problems? Output: Topic switching indicator: Yes; New topic type: After-sales guarantee. When a topic change is detected, a topic change signal is generated. The topic change signal and the new topic type are synchronously transmitted to the script preloading module and the local cache management module. The script preloading module immediately stops loading the scripts for the original topic and starts preloading the scripts for the new topic. The local cache management module clears the original cached scripts and synchronously caches the preloaded scripts for the new topic, thus avoiding the preloading and caching of invalid scripts and improving preloading efficiency and script adaptability.

[0155] Step 3: Preloading and caching of sales scripts locally.

[0156] 3.1 Script Generation: The script preloading module combines the current dialogue content, retrieves relevant knowledge from the knowledge base, and feeds it into the large model to generate scripts.

[0157] 3.2 Local Caching and Dynamic Refresh: The script preloading module transmits the generated scripts to the local cache management module in real time, which then builds a local candidate script cache pool. If a topic switching signal is received, the existing candidate scripts in the cache pool are first cleared, and then the candidate scripts corresponding to the new topic are pushed to ensure that the cached content is consistent with the current topic.

[0158] Step 4: Dual-mode script delivery.

[0159] The script push module uses candidate scripts from the local cache pool to implement a dual-mode push system of passive and active push, ensuring that the scripts are available in real time and accurately adapted, as detailed below.

[0160] 4.1 Passive Trigger Mode.

[0161] The script delivery module monitors the sales trigger signals in real time. These trigger signals include physical button triggers (such as physical buttons on the salesperson's device), virtual button triggers (such as virtual buttons on the mobile app), and specific voice command triggers (such as the salesperson saying "Let me think" or other preset commands). When a sales trigger signal is detected, the module immediately retrieves the cache pool from the local cache management module and quickly matches the current dialogue scenario using a large model (combining the latest dialogue text, customer intent, and topic type) to select the most suitable candidate scripts (usually the 1-2 with the highest weight).

[0162] Since the candidate scripts have been preloaded and stored in the local cache, there is no need to start the script generation process, which can achieve instant matching. After the matching is completed, the scripts are immediately pushed to the sales terminal in a streaming output mode.

[0163] If there is no suitable script for the current scenario in the local cache pool (e.g., the topic suddenly changes and the cache pool has not yet been refreshed), the script push module will immediately trigger the script preloading module, start the emergency script generation process, and simultaneously stream the generated script to the sales terminal in real time.

[0164] 4.2 Proactive intervention model.

[0165] The script delivery module receives trigger warning signals (i.e., the dialogue rhythm reaches the trigger threshold) transmitted by the dialogue status detection and prediction module in real time, and monitors whether the salesperson triggers the script request. When the dialogue status reaches the trigger threshold and the salesperson does not actively trigger the button, the active push process is immediately started to achieve proactive intervention in communication risks.

[0166] The proactive push process is as follows: retrieve high-priority candidate scripts from the local cache pool, combine them with the current communication bottleneck type, select targeted and suitable scripts, and proactively push them to the sales terminal through streaming output to remind sales staff to use the scripts in a timely manner, avoid communication risks, or advance the communication process.

[0167] In some embodiments, the specific applications of the terminal device in providing dialogue assistance to sales personnel through the aforementioned real-time dialogue assistance system in a sales personnel-customer dialogue scenario may include customer silence intervention and price objection intervention, wherein... Customer silence intervention is as follows: If a customer's silence time reaches 2 seconds (trigger threshold) and the salesperson has not triggered the button, the system will proactively push a guiding message: "You can briefly tell me your concerns. Whether it is price or function, I can answer them for you in detail and help you find the most suitable solution." This will encourage the customer to speak, avoid communication interruption, and awaken the customer's willingness to communicate.

[0168] Price objection intervention works as follows: If a customer raises two consecutive price objections, reaching the proactive intervention threshold, and the salesperson does not trigger the button, the system will proactively push a price objection handling script: "Regarding your concerns about pricing, we have a dedicated solution. For the budget needs of small and medium-sized enterprises, we can provide installment payment services. You only need to pay a small amount each month to enjoy full product features and after-sales support." This helps salespersons address the customer's core objections in a targeted manner and improves the success rate of communication.

[0169] In this way, the dual-mode script delivery mechanism achieves full coverage of proactive demand response and passive risk intervention for sales personnel, ensuring that sales personnel can obtain accurate script prompts in a timely manner, respond to customer needs efficiently, and significantly improve sales communication efficiency and conversion rate.

[0170] In this embodiment, a two-dimensional prediction model is constructed by fusing dialogue rhythm and semantic features. This model extracts rhythmic features such as salesperson's and customer's speaking speed and pause duration in real time. Combined with information such as customer intent and objection type obtained from semantic analysis, multi-level warning thresholds are preset to detect potential communication bottlenecks in advance and generate warning signals. This provides a precise basis for script preloading and proactive intervention, solving the problems of lacking predictive capabilities and being unable to detect difficulties in advance. Furthermore, by adopting a pre-processing mechanism of prediction-preloading-local caching, relying on the two-dimensional prediction model, when a warning signal is generated, the script library is immediately retrieved in advance based on dialogue semantics and customer tags. Candidate scripts are generated through a large model and stored in the local cache of the sales terminal, transforming trigger-generated scripts into prediction-preloaded scripts. Moreover, dynamic cache refresh, updating in real time when topics change, solves the problems of response delay and script-context disconnect. Furthermore, the system employs a dual-mode script push mechanism that combines passive triggering and active intervention. In passive push mode, when a salesperson triggers an instruction, the system instantly retrieves the locally cached appropriate script and outputs it in a streaming manner. In active push mode, when the dialogue reaches the trigger threshold and the salesperson does not actively trigger the instruction, the system automatically pushes the appropriate script. This effectively solves the problem that passive responses alone cannot prevent communication failures.

[0171] Please refer to Figure 10 This application also provides a real-time dialogue assistance system, which can implement the above-described real-time dialogue assistance method.

[0172] like Figure 10 As shown, the real-time dialogue assistance system provided in this application embodiment may include: The acquisition module is used to acquire real-time voice dialogue data; The prediction module is used to extract the dialogue rhythm features of the real-time voice dialogue data and generate a dialogue bottleneck warning signal based on the dialogue rhythm features. The preloading module is used to initiate a dialogue preloading process to generate candidate dialogues in response to the dialogue bottleneck warning signal; An auxiliary module is used to provide dialogue assistance based on the candidate dialogues.

[0173] In some embodiments, the prediction module is further configured to compare the dialogue rhythm feature with a preset rhythm warning threshold to obtain a first comparison result; and generate a dialogue bottleneck warning signal if the first comparison result indicates that the dialogue rhythm feature is greater than or equal to the rhythm warning threshold.

[0174] In some embodiments, the prediction module is further configured to compare the dialogue rhythm feature with a preset rhythm trigger threshold to obtain a second comparison result; the rhythm trigger threshold is greater than the rhythm warning threshold; and when the second comparison result is that the dialogue rhythm feature is greater than or equal to the preset rhythm trigger threshold, a dialogue assistance trigger signal is generated. The auxiliary module is also used to actively push the candidate dialogues for dialogue assistance in response to the dialogue assistance trigger signal.

[0175] In some embodiments, the prediction module is further configured to perform semantic analysis on the dialogue text data corresponding to the real-time voice dialogue data to obtain semantic analysis results; and generate a dialogue bottleneck warning signal based on the semantic analysis results.

[0176] In some embodiments, the semantic analysis result includes a dialogue score. The prediction module is further configured to compare the dialogue score with a preset first scoring threshold to obtain a third comparison result; if the third comparison result indicates that the dialogue score is less than the first scoring threshold, a dialogue bottleneck warning signal is generated.

[0177] In some embodiments, the auxiliary module is further configured to proactively push the candidate dialogue for dialogue assistance when the dialogue score is less than a preset second scoring threshold; the second scoring threshold is less than the first scoring threshold; and proactively push the candidate dialogue for dialogue assistance when the number of times the dialogue bottleneck warning signal is generated is greater than or equal to a preset number threshold.

[0178] In some embodiments, the number of candidate statements is greater than one. The preloading module is further configured to adjust the push weight of each of the multiple candidate statements based on the semantic analysis results, so as to obtain multiple candidate statements after adjusting the push weight. The auxiliary module is also used to push a target candidate dialogue from among the multiple candidate dialogues for dialogue assistance; the push weight of the target candidate dialogue is higher than the push weight of other candidate dialogues among the multiple candidate dialogues excluding the target candidate dialogue.

[0179] In some embodiments, the preloading module is further configured to generate new candidate dialogues in response to a topic switching signal; the topic switching signal is generated when a topic switch occurs in the real-time voice dialogue data. The auxiliary module is also used to push the new candidate dialogues to assist in the dialogue.

[0180] It should be noted that the specific implementation of the real-time dialogue assistance system provided in this application is basically the same as the specific implementation of the real-time dialogue assistance method described above, and will not be repeated here.

[0181] Please see Figure 11 This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-mentioned real-time dialogue assistance method.

[0182] In some embodiments, the electronic device may be any smart terminal such as a terminal device, an in-vehicle terminal, an in-vehicle hardware platform (e.g., an in-vehicle computer), a tablet computer, a smartphone, or a wearable device; or, the electronic device may be a vehicle including a memory and a processor.

[0183] like Figure 11 As shown, the electronic device provided in this application embodiment may include: The processor 1101 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1102 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1102 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1102 and is called and executed by the processor 1101 using the real-time dialogue assistance method of the embodiments of this application. Input / output interface 1103 is used to implement information input and output; The communication interface 1104 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1105 transmits information between various components of the device (e.g., processor 1101, memory 1102, input / output interface 1103, and communication interface 1104); The processor 1101, memory 1102, input / output interface 1103 and communication interface 1104 are connected to each other within the device via bus 1105.

[0184] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described real-time dialogue assistance method.

[0185] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0186] This application also provides a computer program product, including a computer program. The steps implemented by the computer program when executed by a processor are basically the same as those in the specific embodiments of the real-time dialogue assistance method described above, and will not be repeated here.

[0187] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0188] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0189] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0190] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0191] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0192] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding factor is divided by the following factor, or that the related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0193] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0194] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0195] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0196] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0197] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A real-time dialogue assistance method, characterized in that, The method includes: Acquire real-time voice dialogue data; Extract the dialogue rhythm features from the real-time voice dialogue data, and generate a dialogue bottleneck warning signal based on the dialogue rhythm features; In response to the dialogue bottleneck warning signal, the script preloading process is initiated to generate candidate scripts; Dialogue assistance is provided based on the candidate dialogues.

2. The method according to claim 1, characterized in that, The generation of a dialogue bottleneck warning signal based on the dialogue rhythm features includes: The dialogue rhythm features are compared with a preset rhythm warning threshold to obtain a first comparison result; If the first comparison result indicates that the dialogue rhythm feature is greater than or equal to the rhythm warning threshold, a dialogue bottleneck warning signal is generated.

3. The method according to claim 2, characterized in that, The method further includes: The dialogue rhythm features are compared with a preset rhythm trigger threshold to obtain a second comparison result; the rhythm trigger threshold is greater than the rhythm warning threshold. If the second comparison result indicates that the dialogue rhythm feature is greater than or equal to a preset rhythm trigger threshold, a dialogue assistance trigger signal is generated. The dialogue assistance based on the candidate dialogues includes: In response to the dialogue assistance trigger signal, the candidate dialogues are actively pushed to assist in the dialogue.

4. The method according to claim 1, characterized in that, The method further includes: Semantic analysis is performed on the dialogue text data corresponding to the real-time voice dialogue data to obtain semantic analysis results; A dialogue bottleneck warning signal is generated based on the semantic analysis results.

5. The method according to claim 4, characterized in that, The semantic analysis results include a dialogue score, and the generation of a dialogue bottleneck warning signal based on the semantic analysis results includes: The dialogue score is compared with a preset first score threshold to obtain a third comparison result; If the third comparison result indicates that the dialogue score is less than the first score threshold, a dialogue bottleneck warning signal is generated.

6. The method according to claim 5, characterized in that, The dialogue assistance based on the candidate dialogues includes at least one of the following: If the dialogue score is less than a preset second score threshold, the candidate dialogues are actively pushed to assist in the dialogue. The second scoring threshold is less than the first scoring threshold; If the number of times the dialogue bottleneck warning signal is generated is greater than or equal to a preset threshold, the candidate dialogue is actively pushed to assist in the dialogue.

7. The method according to claim 4, characterized in that, The number of candidate dialogues is greater than one. After the dialogue bottleneck warning signal is triggered to initiate the dialogue preloading process to generate candidate dialogues, the method further includes: Based on the semantic analysis results, the push weights of each of the multiple candidate dialogues are adjusted to obtain multiple candidate dialogues with adjusted push weights. The dialogue assistance based on the candidate dialogues includes: The target candidate dialogue is pushed from among the multiple candidate dialogues to assist in the dialogue; the push weight of the target candidate dialogue is higher than the push weight of other candidate dialogues among the multiple candidate dialogues excluding the target candidate dialogue.

8. The method according to any one of claims 1 to 7, characterized in that, After initiating the dialogue preloading process to generate candidate dialogues in response to the dialogue bottleneck warning signal, the method further includes: New candidate statements are generated in response to a topic switching signal; the topic switching signal is generated when a topic switch occurs in the real-time voice dialogue data. The dialogue assistance based on the candidate dialogues includes: The new candidate phrases are pushed to assist in the conversation.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the real-time dialogue assistance method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the real-time dialogue assistance method as described in any one of claims 1 to 8.

11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the real-time dialogue assistance method as described in any one of claims 1 to 8.