A voice AI-based intelligent conference management method and system

By collecting and analyzing the voice and gesture data of sales personnel, combined with voice AI and intelligent conference management systems, the problem of multimodal interaction features being difficult to restore in existing technologies is solved, and intelligent management and instant feedback of the sales process are achieved.

CN120475122BActive Publication Date: 2025-09-26BEIJING DEHE SHUNTIAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510961372.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-26
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing technologies are unable to fully restore the multimodal interaction characteristics in real sales scenarios in automobile sales training and actual combat drills, affecting the integrity and real-time nature of process compliance assessments, especially in complex sales scenarios.

Method used

By collecting the voice data and gesture data of sales personnel, using voice AI recognition to generate an ordered text sequence, and extracting a set of keywords and gestures, combined with the intelligent conference management system to compare with the preset sales process node sequence, generate process deviation information, and display the process node status mark in real time through display control instructions.

Benefits of technology

It achieves the simultaneous perception of language expression and body behavior during the sales process, improves the intelligence level of process compliance assessment, enhances the restoration of real sales scenarios and the reliability of process assessment, and ensures instant feedback and visual guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475122B_ABST
    Figure CN120475122B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for intelligent conference management based on voice AI, which collects voice data and gesture data of sales personnel; uses voice AI to generate an ordered text sequence; matches the keywords therein with a preset term library in the field of automobile sales to obtain a set of matching keywords; extracts the first type of actions pointing to vehicle entities and the second type of actions emphasizing sales nodes from the gesture data to generate a gesture action set; based on the intelligent conference management system, compares the matching keyword set and gesture action set with the preset sales process node sequence to generate process deviation information; parses the node position identifiers therein, generates display control instructions, and sends them to the sales speech demonstration interface to realize intelligent guidance and management of sales personnel; the present application realizes monitoring and deviation warning of automobile sales processes, as well as real-time guidance and process standardization management of automobile OEM store sales personnel in speech demonstrations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to a method and system for intelligent conference management based on speech AI. Background Art

[0002] In automotive sales training and real-world practice scenarios, salespeople must master standard dialogue and employ appropriate body language to improve customer communication efficiency and close sales. This places high demands on process standardization during meetings or training, necessitating a technology that can perceive voice and motion information in real time, perform intelligent analysis, and provide process feedback.

[0003] To meet these technical requirements, the current mainstream solution utilizes an intelligent conference assistance system based on single-line speech recognition and keyword matching. This system uses a microphone array to capture the salesperson's voice content, converts the speech into text using automatic speech recognition technology, extracts key information from it, compares it with the standard sales process, identifies situations where the speech is missing or the sequence is out of order, and generates corresponding prompt information for feedback to the user. However, existing solutions have some flaws. For example, since they rely solely on voice information to judge the process, they lack the perception and integration analysis of non-verbal interactive behaviors such as gestures and movements, making it difficult to fully reproduce the multimodal interaction characteristics in real sales scenarios. At the same time, this solution cannot effectively capture the semantic information conveyed by emphasized movements, which in turn affects the integrity and real-time nature of the overall process compliance assessment, limiting its application effectiveness in complex sales scenarios. Summary of the Invention

[0004] This application provides a voice AI-based intelligent conference management method and system to solve the problems in the existing technology that it is difficult to fully restore the multimodal interaction characteristics in real sales scenarios; it affects the integrity and real-time performance of the overall process compliance assessment, and the application effect is poor in complex sales scenarios.

[0005] In the first aspect, the present application provides a method for intelligent conference management based on voice AI, comprising:

[0006] Collect sales staff’s voice data and gesture data;

[0007] Using voice AI to recognize the voice data and generate an ordered text sequence;

[0008] Matching the keywords in the ordered text sequence with a preset term library in the field of automobile sales to obtain a set of matching keywords;

[0009] Extracting a first type of action directed at a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combining the first type of action and the second type of action to generate a gesture action set;

[0010] Based on the intelligent conference management system, the matching keyword set and the gesture action set are compared with a preset sales process node sequence to generate process deviation information;

[0011] The node position identifier in the process deviation information is parsed to generate a display control instruction, and the display control instruction is sent to the sales speech demonstration interface to display the status mark corresponding to each sales process node in real time, thereby realizing intelligent guidance and management of sales personnel.

[0012] Optionally, using voice AI to recognize the voice data and generate an ordered text sequence, including:

[0013] Using voice AI, performing noise reduction processing on the voice data to obtain noise-reduced voice data;

[0014] Performing frame processing on the noise-reduced voice data to obtain a continuous voice stream;

[0015] Segmenting the continuous speech stream according to a preset time window to generate multiple speech segments;

[0016] Extracting acoustic feature vectors from each speech segment and matching candidate syllable texts corresponding to the acoustic feature vectors;

[0017] Based on a preset terminology library in the field of automobile sales, the candidate syllable text is subjected to vocabulary correction processing to generate a corrected text;

[0018] Based on the segmentation order of the speech segments, the corrected text is reorganized in time sequence to generate an ordered text sequence.

[0019] Optionally, based on a preset terminology library in the field of automobile sales, the candidate syllable text is subjected to vocabulary correction processing to generate a corrected text, including:

[0020] Extracting first acoustic feature parameters of each candidate syllable from the candidate syllable text, converting the first acoustic feature parameters into a plurality of first feature vectors, and combining all the first feature vectors into a candidate syllable feature sequence, wherein the first acoustic feature parameters include a fundamental frequency contour, a formant frequency, and a phoneme duration;

[0021] Extracting a second acoustic feature parameter of each standard syllable from the standardized terms in the automobile sales field term library, converting the second acoustic feature parameter into a plurality of second feature vectors, and combining all the second feature vectors into a standardized term feature sequence;

[0022] Calculating the syllable space distance between the candidate syllable feature sequence and the standardized term feature sequence, and selecting the standardized term with the smallest syllable space distance as the target candidate term;

[0023] If the syllable space distance of the target candidate term is lower than the preset tolerance threshold, the complete text of the standardized term is extracted as the corrected text; if the syllable space distance of the target candidate term is higher than the preset tolerance threshold, the candidate syllable text is used as the corrected text.

[0024] Optionally, the keywords in the ordered text sequence are matched with a preset term library in the field of automobile sales to obtain a set of matching keywords, including:

[0025] Performing semantic boundary analysis on the text content in the ordered text sequence to obtain a plurality of candidate keywords;

[0026] Extract vehicle performance parameter keywords and sales plan keywords from the preset automobile sales term library;

[0027] Calculating the similarity between the candidate keywords and the vehicle performance parameter keywords and the sales plan keywords, and selecting candidate keywords that meet a preset matching threshold as valid keywords;

[0028] All valid keywords are aggregated to generate a matching keyword set corresponding to the text content of the ordered text sequence.

[0029] Optionally, extracting a first type of action directed to a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combining the first type of action and the second type of action to generate a gesture action set, including:

[0030] Performing three-dimensional spatial coordinate conversion on the gesture action data to generate a three-dimensional spatial gesture trajectory;

[0031] Based on the preset position coordinates of the vehicle entity, selecting a first trajectory segment that meets the pointing condition from the three-dimensional gesture trajectory, and merging all the first trajectory segments into a first type of action pointing to the vehicle entity;

[0032] Determine a node time window based on time stamps in a preset sales process node sequence, select second trajectory segments within the node time window whose action amplitude values ​​exceed a preset threshold from the three-dimensional gesture trajectory, and merge all second trajectory segments into a second type of action that emphasizes the sales node;

[0033] The first type of actions and the second type of actions are combined to generate a gesture action set.

[0034] Optionally, based on the intelligent conference management system, the matching keyword set and the gesture action set are compared with a preset sales process node sequence to generate process deviation information, including:

[0035] Based on the intelligent conference management system, keyword matching is performed on the matching keyword set with sales process nodes in a preset sales process node sequence to obtain matching results, and status marking is performed on the sales process nodes according to the matching results to generate a voice coverage status for each sales process node;

[0036] Expanding the time stamps of all sales process nodes to generate a node time range, extracting a target gesture action occurring within the node time range from the gesture action set, and determining the action synchronization state of the target gesture action;

[0037] Co-verifying the voice coverage state and the motion synchronization state to identify a target node whose voice coverage is uncovered or whose motion synchronization state is not synchronized;

[0038] The location identifier of the target node is extracted and the deviation type is marked to generate process deviation information.

[0039] Optionally, parsing the node position identifier in the process deviation information to generate a display control instruction includes:

[0040] Extracting a node position identifier and a deviation type identifier of a target node from the process deviation information;

[0041] Convert the node position identifier into the target node area coordinates in the sales talk demonstration interface;

[0042] Matching the deviation type identifier in the process deviation information with a preset state mark type rule library to determine graphic attributes and display parameters corresponding to the deviation type identifier;

[0043] The target node area coordinates, the graphic attributes, and the display parameters are packaged into instructions to generate a display control instruction.

[0044] In the second aspect, this application provides a voice AI-based intelligent conference management system, including:

[0045] The acquisition module is used to collect the salesperson’s voice data and gesture data;

[0046] A recognition module, configured to utilize voice AI to recognize the voice data and generate an ordered text sequence;

[0047] A matching module, configured to match the keywords in the ordered text sequence with a preset term library in the field of automobile sales to obtain a set of matching keywords;

[0048] an aggregation module, configured to extract a first type of action directed to a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combine the first type of action and the second type of action to generate a gesture action set;

[0049] a comparison module, configured to compare the matching keyword set and the gesture action set with a preset sales process node sequence based on the intelligent conference management system to generate process deviation information;

[0050] The marking module parses the node position identifier in the process deviation information, generates a display control instruction, and sends the display control instruction to the sales speech demonstration interface to display the status mark corresponding to each sales process node in real time, thereby realizing intelligent guidance and management of sales personnel.

[0051] In a third aspect, an embodiment of the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a voice AI-based intelligent conference management method as described in the first aspect above.

[0052] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements a voice AI-based intelligent conference management method as described in the first aspect.

[0053] In this application, the voice data and gesture data of the salesperson are collected; the voice data is recognized using voice AI to generate an ordered text sequence; the keywords in the ordered text sequence are matched with a preset terminology library in the field of automobile sales to obtain a set of matching keywords; the first type of action pointing to the vehicle entity and the second type of action emphasizing the sales node are extracted from the gesture data, and the first type of action and the second type of action are combined to generate a gesture set; based on the intelligent conference management system, the matching keyword set and the gesture set are compared with a preset sales process node sequence to generate process deviation information; the node position identifier in the process deviation information is parsed to generate a display control instruction, and the display control instruction is sent to the sales speech demonstration interface to display the status mark corresponding to each sales process node in real time, thereby realizing intelligent guidance and management of sales personnel. The technical solution provided in this application realizes the synchronous perception of language expression and body behavior in the sales process; uses voice AI to recognize voice data and generate an ordered text sequence, which can convert continuous voice content into structured text, improving the accuracy and logic of the speech content analysis; matches the keywords in the text with the terminology library in the automotive sales field, which helps to accurately identify the key speech nodes in the sales process and enhance the system's understanding of industry semantics; extracts the first type of actions pointing to vehicle entities and the second type of actions emphasizing sales nodes from gesture action data, and aggregates them so that non-verbal information such as pointing, emphasizing and other actions can be effectively identified and incorporated into the process evaluation system, thereby enhancing the system's restoration of real sales scenarios; based on the intelligent conference management system, the matching keyword set and the gesture action set are compared with the preset process to realize dynamic monitoring and deviation judgment of the sales process execution, thereby improving the intelligence level of process compliance assessment; finally, the node position identifier in the deviation information is parsed and display control instructions are generated to drive the sales speech demonstration interface to display the process status mark in real time, realizing instant feedback and visual guidance for sales personnel. This application enhances the robustness of speech recognition in complex contexts, especially when dealing with frequently occurring professional terms in sales pitches. It can more accurately restore speech content and avoid process judgment deviations caused by misrecognition. At the same time, by reorganizing the temporal sequence of the speech stream, the generated text is ensured to be consistent with the actual expression order, providing a high-quality language input foundation for subsequent process node matching. This compensates for the shortcomings of existing solutions in speech recognition accuracy and improves the overall system's process evaluation reliability and response speed.

[0054] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0056] Figure 1 A flowchart of a voice AI-based intelligent conference management method provided by this application is shown;

[0057] Figure 2 The following is a schematic diagram of the structure of a voice AI intelligent conference management system provided by the present application;

[0058] Figure 3 A schematic structural diagram of a computing device provided by the present application is shown. DETAILED DESCRIPTION

[0059] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0060] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.

[0061] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0062] In response to the urgent need for process standardization and multimodal interactive feedback in automotive sales training and practical exercises, existing technologies mainly rely on single voice recognition and keyword matching mechanisms, making it difficult to fully perceive the comprehensive performance of sales personnel in terms of speech expression and body movements, resulting in significant limitations in process compliance judgment. To address this problem, this application proposes an intelligent meeting management method based on voice AI. By synchronously collecting voice and gesture data, combined with voice AI recognition, an ordered text sequence is generated, and a set of keywords highly relevant to the sales process is extracted. At the same time, gestures are classified and aggregated to construct a gesture set containing directional and emphatic gestures, thereby realizing the fusion analysis of language and non-verbal information. On this basis, the two types of information are compared with the preset sales process node sequence to generate process deviation information. By parsing the node position identifiers, visual display control instructions are generated, and the status marks of each process node are displayed in real time. Ultimately, a full-process management system is constructed that can fully perceive, intelligently analyze, and instantly guide sales behavior. This method makes up for the shortcomings of existing technologies in multimodal perception and process feedback capabilities, and improves the intelligence level of sales training and practical guidance.

[0063] Figure 1 A flowchart of a voice AI-based intelligent conference management method is provided for the embodiment of this application, such as Figure 1 As shown, the method includes:

[0064] Step 101: Collect the salesperson's voice data and gesture data.

[0065] In this step, voice data refers to the raw audio signals collected by the microphone during the salesperson's presentation, including the human voice, background noise in the exhibition hall, and equipment interference, stored in pulse code modulation format. Gesture data refers to the three-dimensional coordinate sequence of hand joints captured by the depth camera, including position, velocity vector, and rotation angle.

[0066] In an embodiment of the present application, the voice data of the salesperson during the explanation is collected by a directional microphone array, and the gesture action data is captured by a depth camera.

[0067] Step 102: Use voice AI to recognize the voice data and generate an ordered text sequence.

[0068] In this step, the ordered text sequence refers to the text units output by the voice AI arranged in chronological order, and each text unit corresponds to a voice segment.

[0069] In the embodiment of the present application, the voice AI filters out impulse noise (such as the sound of opening and closing doors) through a convolutional noise reduction network, uses frame processing to generate a continuous voice stream to extract acoustic features; matches candidate syllable combinations through a syllable decoder, and then performs forced alignment correction based on the automotive sales field terminology library, and finally reorganizes the text sequence by timestamp to generate an ordered text sequence.

[0070] Step 103: Match the keywords in the ordered text sequence with a preset term library in the field of automobile sales to obtain a set of matching keywords.

[0071] In this step, the pre-defined automotive sales domain terminology database refers to a pre-defined set of automotive sales-specific vocabulary, including performance parameters and sales policies, stored in the form of word vectors. The matching keyword set refers to the keywords successfully matched to the terminology database from the ordered text sequence and their similarity scores.

[0072] In an embodiment of the present application, noun phrases are extracted from an ordered text sequence as candidate keywords, and similarity matching is performed with a preset term library in the field of automobile sales. Similarity = cosine of the angle × 100. The cosine of the angle between the text vector of the keyword and the text vector of the keyword in the preset term library in the field of automobile sales is calculated, and valid keywords are screened out to obtain a set of matching keywords.

[0073] Step 104: extracting a first type of action directed to a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combining the first type of action and the second type of action to generate a gesture action set.

[0074] In this step, the first type of action refers to spatial movements where the hand trajectory continuously points toward the physical vehicle in the showroom. The trajectory segments must be directional and have a duration. The second type of action refers to emphatic gestures where the hand movement amplitude exceeds a threshold within the time window of a sales process node.

[0075] In an embodiment of the present application, a spatiotemporal analysis is performed on gesture action data to identify the intersection of the extended line of the hand trajectory and the vehicle entity. If it lasts for a certain period of time, it is determined to be a valid pointing action and is taken as the first type of action. The peak hand movement speed is detected and occurs within the time window of the sales process node. This action is taken as the second type of action. The first type of action and the second type of action are combined to generate a gesture action set.

[0076] Step 105: Based on the intelligent conference management system, the matching keyword set and the gesture action set are compared with a preset sales process node sequence to generate process deviation information.

[0077] In this step, the intelligent conference management system refers to the software system deployed on the vehicle terminal, which has a built-in sales process engine, multimodal analysis module, and real-time renderer. The preset sales process node sequence refers to a linked list of nodes in the standard automobile sales process, with each node containing an ID, name, time window, and a list of bound keywords. Process deviation information refers to the structured data of the comparison results, recording the deviation node ID, deviation type, and severity level.

[0078] In the embodiment of the present application, the intelligent conference management system calls a preset sales process node sequence; compares the matching keyword set with the node binding term to obtain the node voice coverage status; detects whether there is an emphasis action in the gesture action set within the node time window, and generates the node action synchronization status; identifies abnormal nodes (such as missing voice and unsynchronized action) through state and gate verification, and outputs process deviation information with position identification

[0079] Step 106: parse the node position identifier in the process deviation information, generate a display control instruction, and send the display control instruction to the sales speech demonstration interface to display the status mark corresponding to each sales process node in real time, thereby realizing intelligent guidance and management of sales personnel.

[0080] In this step, the node location identifier refers to the logical identifier of the sales process node in the interface layout, mapped to the user coordinate area. Display control instructions refer to instructions for controlling interface elements, including status mark type, coordinates, and animation parameters. The sales talk demonstration interface refers to the interactive interface displayed on the vehicle's central control screen, which includes a 3D vehicle model, parameter table, process progress bar, and status mark layer. Status marks refer to dynamic graphic elements superimposed on the demonstration interface to indicate process deviations.

[0081] In an embodiment of the present application, the node position identifier in the parsed process deviation information is converted into a pixel coordinate area of ​​the sales speech demonstration interface. According to the deviation type, the status marking rule library is called to generate display control instructions containing coordinates and icons, which are pushed to the sales speech demonstration interface in real time, and dynamic marks are superimposed and displayed next to the vehicle 3D model to achieve intelligent guidance and management of sales personnel.

[0082] The embodiments of the present application solve the three major problems of failure of term recognition in high-noise environments, spatiotemporal disconnection between actions and speech, and delayed feedback on process deviations through multimodal collaborative analysis and precise mapping of sales process nodes.

[0083] This application provides a specific embodiment, step 102, using voice AI to recognize the voice data and generate an ordered text sequence, specifically including the following steps:

[0084] Step 201: Using voice AI, the voice data is subjected to noise reduction processing to obtain noise-reduced voice data.

[0085] In this step, the denoised speech data refers to the speech signal processed by spectral subtraction.

[0086] In an embodiment of the present application, the voice AI adopts a noise reduction algorithm based on spectral subtraction. First, the short-time Fourier transform of the voice data is calculated to obtain a spectrum diagram, and the noise spectrum is generated using the noise estimation module; the original spectrum is subtracted from the noise spectrum and multiplied by the gain coefficient, and then the time domain signal is reconstructed through an inverse Fourier transform to obtain the noise-reduced voice data.

[0087] Step 202: performing frame processing on the noise-reduced speech data to obtain a continuous speech stream.

[0088] In this step, the continued voice stream refers to the voice frame sequence after frame division.

[0089] In an embodiment of the present application, the noise-reduced speech data is windowed and segmented to obtain a continuous time domain signal, which is then converted into a set of time segments to generate a continuous speech stream.

[0090] Step 203: Segment the continuous speech stream according to a preset time window to generate multiple speech segments.

[0091] In this step, the preset time window refers to setting a fixed segment duration to adapt to the noise environment of the car showroom, and adding redundancy based on the maximum duration of the impulse noise. The multiple speech segments refer to a collection of semantic units segmented by time window.

[0092] In an embodiment of the present application, a preset time window is used to aggregate continuous speech streams, and the continuous speech streams within a certain period of time are merged into a speech segment. The semantic units are segmented through endpoint detection, and multiple speech segments are output.

[0093] Step 204: extracting the acoustic feature vector from each speech segment, and matching the candidate syllable text corresponding to the acoustic feature vector.

[0094] In this step, the acoustic feature vector refers to the audio feature array. The candidate syllable text refers to the syllable sequence output by decoding, separated by hyphens.

[0095] In an embodiment of the present application, a feature vector is extracted for each speech segment, and the feature vector is input into a pre-trained hidden Markov model for syllable decoding. The optimal state path is then calculated using the Viterbi algorithm, and the candidate syllable text with the highest probability is output.

[0096] Step 205: Based on a preset terminology library in the field of automobile sales, the candidate syllable text is subjected to vocabulary correction processing to generate a corrected text.

[0097] In this step, the corrected text refers to the Chinese characters or syllables text corrected by the terminology database. The correction rule is that if the edit distance is ≤ 2, it is replaced with the terminology database standard word, otherwise the original syllable is retained.

[0098] In an embodiment of the present application, the candidate syllable text is compared with a preset terminology library in the field of automobile sales, and the syllable sequence editing distance is calculated. If the distance is lower than a preset tolerance threshold and matches the standardized term in the terminology library, it is replaced with the standardized term; otherwise, the original syllable is retained and the corrected text is generated.

[0099] Step 206: Based on the segmentation order of the speech segments, the corrected text is reorganized in time sequence to generate an ordered text sequence.

[0100] In this step, the segment order refers to the temporal index of the speech segments, which are arranged in ascending order of the start time and are used to reorganize the text order.

[0101] In an embodiment of the present application, the corrected text is sorted according to the segmentation order of the speech segments, and the ambiguity of the segmentation is eliminated by sliding window splicing to generate an ordered text sequence.

[0102] The embodiment of the present application improves the impulse noise suppression rate and makes professional vocabulary recognition more accurate by implementing noise reduction processing, domain term correction and time sequence reorganization in the high-noise scene of the car showroom.

[0103] This application provides a specific embodiment, step 205, based on a preset term library in the field of automobile sales, performing vocabulary correction processing on the candidate syllable text to generate a corrected text, specifically comprising the following steps:

[0104] Step 211: extracting the first acoustic feature parameters of each candidate syllable from the candidate syllable text, converting the first acoustic feature parameters into multiple first feature vectors, and combining all the first feature vectors into a candidate syllable feature sequence, wherein the first acoustic feature parameters include fundamental frequency contour, formant frequency and phoneme duration.

[0105] In this step, the first acoustic feature parameter refers to the physical acoustic index parsed from the candidate syllable text, including the fundamental frequency contour, resonance peak frequency, and phoneme duration, which are used to characterize the pronunciation characteristics of the syllable. The fundamental frequency contour refers to the trajectory curve of the fundamental frequency changing with time during the pronunciation of the syllable, measured in Hertz (Hz), reflecting the changing characteristics of the vocal cord vibration frequency. The resonance peak frequency refers to the frequency value corresponding to the peak value of the spectral energy formed by the resonance of the vocal tract, including three main resonance peaks, which are used to distinguish different vowels. The phoneme duration refers to the duration ratio of the consonant transition segment to the vowel steady-state segment in the syllable.

[0106] In an embodiment of the present application, three first acoustic feature parameters of each candidate syllable are first extracted from the candidate syllable text, namely, the fundamental frequency contour refers to the curve of the fundamental frequency changing with time during the pronunciation of the syllable, reflecting the vibration frequency of the vocal cords; the formant frequency refers to the peak energy frequency generated by the resonance of the vocal tract, which is used to distinguish the timbre of vowels; and the phoneme duration refers to the duration ratio of the consonant and vowel parts in the syllable. These three first acoustic feature parameters are respectively converted into multiple first feature vectors, and then all the first feature vectors of the same syllable are spliced ​​in chronological order to obtain the feature representation of the syllable. Finally, the feature representations of all syllables are combined in sequence to generate a candidate syllable feature sequence.

[0107] Step 212: extracting the second acoustic feature parameter of each standard syllable from the standardized terms in the automobile sales field terminology library, converting the second acoustic feature parameter into a plurality of second feature vectors, and combining all the second feature vectors into a standardized term feature sequence.

[0108] In this step, the candidate syllable feature sequence refers to a sequence of feature vectors of all syllables in the candidate syllable text in time order. The standardized term feature sequence refers to the sequence of syllable feature vectors of standard terms in the terminology database, which is aligned with the candidate sequence dimension.

[0109] In an embodiment of the present application, a standardized term is selected from a term library in the field of automobile sales, and the second acoustic feature parameters of each standard syllable thereof are extracted, including the fundamental frequency contour, the resonance peak frequency, and the phoneme duration; the second acoustic feature parameters are converted into multiple second feature vectors, which are spliced ​​according to the syllable time sequence to generate a feature sequence of the standardized term.

[0110] Step 213: Calculate the syllable space distance between the candidate syllable feature sequence and the standardized term feature sequence, and select the standardized term with the smallest syllable space distance as the target candidate term.

[0111] In this step, the syllable space distance refers to the sum of the Euclidean distances of all corresponding position vectors between the two sequences. The preset tolerance threshold refers to the distance critical value set according to the exhibition hall environment, which is used to determine whether to perform term replacement.

[0112] In this embodiment, the Euclidean distance between the vectors at the same position in the candidate syllable feature sequence and each standardized term feature sequence is calculated, and the Euclidean distance values ​​for all positions are summed to obtain the total syllable space distance. After traversing the automotive sales domain terminology library, the standardized term with the smallest syllable space distance is selected as the target candidate term.

[0113] Step 214: If the syllable space distance of the target candidate term is lower than the preset tolerance threshold, the complete text of the standardized term is extracted as the corrected text; if the syllable space distance of the target candidate term is higher than the preset tolerance threshold, the candidate syllable text is used as the corrected text.

[0114] In an embodiment of the present application, it is determined whether the syllable space distance of the target candidate term is lower than a preset tolerance threshold, which is set according to the noise level of the car showroom. If it is lower than the preset tolerance threshold, the complete Chinese character text of the term is extracted as the corrected text; if it is higher than the preset tolerance threshold, the original candidate syllable text is retained as the corrected text.

[0115] The embodiment of the present application improves the correction accuracy and reduces the term correction error rate through an acoustic feature comparison mechanism driven by domain terminology.

[0116] This application provides a specific embodiment, step 103, matching the keywords in the ordered text sequence with a preset term library in the field of automobile sales to obtain a set of matching keywords, specifically including the following steps:

[0117] Step 301: Perform semantic boundary analysis on the text content in the ordered text sequence to obtain multiple candidate keywords.

[0118] In this step, multiple candidate keywords refer to independent semantic units segmented from an ordered text sequence, which must meet the required length and contain no auxiliary words.

[0119] In an embodiment of the present application, semantic boundary analysis is performed on the text content in an ordered text sequence to identify noun phrases, and combined with pause word segmentation, independent semantic units are extracted as candidate keywords. For example, if the text content is "This SUV accelerates to 100 kilometers per hour in only 5 seconds", three candidate keywords, SUV, acceleration to 100 kilometers per hour, and 5 seconds, are extracted after boundary analysis.

[0120] Step 302: extract vehicle performance parameter keywords and sales plan keywords from a preset automobile sales field term library.

[0121] In this step, vehicle performance parameter keywords refer to a collection of professional terms describing a vehicle's technical characteristics, including power parameters, chassis parameters, and range parameters such as torque horsepower, four-wheel drive system, and electric consumption per 100 kilometers. Sales plan keywords refer to a collection of promotional and financial service terms, including financial plans and replacement services, such as manufacturer's suggested price, 24-month interest-free installments, and used car discounts.

[0122] In this embodiment, vehicle performance parameter keywords and keywords describing vehicle technical indicators, such as torque and mileage, are extracted from a pre-defined automotive sales terminology library. Sales plan keywords, which refer to promotional finance terms such as replacement subsidies and low interest rates, are stored as separate sets for subsequent matching.

[0123] Step 303: Calculate the similarity between the candidate keywords and the vehicle performance parameter keywords and the sales plan keywords, and select the candidate keywords that meet the preset matching threshold as valid keywords.

[0124] In this step, similarity refers to the strength of semantic association, measured by the cosine value of the word vector. Higher similarity indicates greater similarity. The preset matching threshold is the critical similarity value for determining keyword validity. It is dynamically adjusted based on the exhibition hall noise level, for example, 0.8 for 65dB noise and 0.9 for 50dB noise. Valid keywords are candidate keywords that exceed the threshold and are included in the terminology database. The record format is keyword text: similarity value.

[0125] In this embodiment of the present application, a cosine similarity algorithm is used to calculate the similarity between the candidate keyword vector and the vehicle performance parameter keyword term library vector and the sales plan keyword term library vector, resulting in a similarity value ranging from -1 to 1. When the similarity value reaches a preset matching threshold, the candidate keyword is marked as a valid keyword.

[0126] Step 304: Aggregate all valid keywords to generate a matching keyword set corresponding to the text content of the ordered text sequence.

[0127] In the embodiment of the present application, all valid keywords are arranged and aggregated according to their order of appearance in the ordered text sequence to generate a structured matching keyword set, which records the keyword text and similarity value.

[0128] The embodiments of the present application achieve three major technical breakthroughs: improving the accuracy of domain keyword recognition, reducing the missed detection rate of policy terms, and reducing the time consumption of matching processing, thereby optimizing the term analysis efficiency in the automobile sales scenario.

[0129] For example, in a certain car sales scenario, the system collects the salesperson's voice data and generates an ordered text sequence through voice AI recognition, such as "autopilot is free for three years." The system then performs semantic boundary analysis on the text content in the ordered text sequence and extracts multiple candidate keywords, including "autopilot" and "free for three years." Next, the system extracts vehicle performance parameter keywords and sales plan keywords from a pre-set car sales terminology library. Vehicle performance parameter keywords include "autopilot" and "range," while sales plan keywords include "free maintenance" and "financial interest subsidies." The system further calculates the similarity between each candidate keyword and the above two categories of keywords. After calculation, the similarity between the candidate keyword "automatic assisted driving" and the vehicle performance parameter keyword "automatic assisted driving" is 0.94; the similarity between the sales plan keyword "three years free" and the sales plan keyword "free maintenance" is 0.89; based on the preset matching threshold (such as 0.8), valid keywords that meet the conditions are screened out, including: automatic assisted driving (similarity 0.94), three years free (similarity 0.89); finally, the system aggregates all valid keywords to generate a matching keyword set corresponding to this segment of ordered text, which is used for subsequent process node matching and intelligent guidance.

[0130] The present application provides a specific embodiment, step 104, extracting a first type of action pointing to a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combining the first type of action and the second type of action to generate a gesture action set, which specifically includes the following steps.

[0131] Step 401: performing three-dimensional space coordinate conversion on the gesture action data to generate a three-dimensional space gesture trajectory.

[0132] In this step, the three-dimensional gesture trajectory refers to the motion path of the hand key points in the three-dimensional coordinate system, including X / Y / Z axis coordinates and timestamps.

[0133] In an embodiment of the present application, the original coordinates of the hand joint points in the gesture action data are obtained through a depth camera, and the three-dimensional space coordinates are converted using the coordinate system transformation matrix to convert the original coordinate system data of the hand joint points into world coordinate system data to generate a continuous three-dimensional space gesture trajectory.

[0134] Step 402: Based on the preset position coordinates of the vehicle entity, select the first trajectory segment that meets the pointing condition from the three-dimensional gesture trajectory, and merge all the first trajectory segments into a first type of action pointing to the vehicle entity.

[0135] In this step, the preset vehicle physical position coordinates refer to the coordinates of the three-dimensional bounding box of the exhibition vehicle calibrated by laser scanning. The first trajectory segment refers to a continuous trajectory segment that meets the vehicle pointing condition.

[0136] In an embodiment of the present application, the preset vehicle entity position coordinates are loaded, and the trajectory segment that meets the pointing condition in the three-dimensional space gesture trajectory is selected. When the angle between the hand trajectory direction vector and the normal vector of the vehicle entity position coordinates meets the condition, it is marked as the first trajectory segment, and all continuous first trajectory segments are merged into the first type of action pointing to the vehicle entity.

[0137] Step 403: Determine a node time window based on the time stamps in a preset sales process node sequence. Within the node time window, select a second trajectory segment whose action amplitude value exceeds a preset threshold from the three-dimensional space gesture trajectory, and merge all the second trajectory segments into a second type of action that emphasizes the sales node.

[0138] In this step, the time stamp refers to the start and end times of the sales process node on the timeline, with an accuracy of 0.1 second. The node time window refers to the motion detection interval extended based on the time stamp. The preset threshold refers to the critical value of hand movement speed, which is dynamically adjusted based on the salesperson's height.

[0139] In this embodiment of the present application, a node time window is determined based on the time stamps in a preset sales process node sequence. For example, the quotation node is expanded to the node start time minus 2 seconds and the node end time plus 3 seconds. Within the node time window, the three-dimensional gesture trajectory is scanned. When the hand movement speed value exceeds a preset threshold, it is marked as a second trajectory segment. All second trajectory segments that meet the conditions are merged into the second type of action that emphasizes the sales node.

[0140] Step 404: Combine the first type of actions and the second type of actions to generate a gesture action set.

[0141] In an embodiment of the present application, the first type of action set and the second type of action set are structurally combined according to the action type index to generate a gesture action set containing spatial trajectory metadata.

[0142] The embodiment of the present application achieves a technical breakthrough by analyzing gestures with dual constraints of space and time sequence, improving the accuracy of vehicle pointing motion recognition, reducing gesture synchronization errors at sales nodes, and improving the filtering rate of invalid gestures.

[0143] This application provides a specific embodiment. In step 105, based on the intelligent conference management system, the matching keyword set and the gesture action set are compared with a preset sales process node sequence to generate process deviation information. The specific steps include:

[0144] Step 501: Based on the intelligent conference management system, keyword matching is performed on the matching keyword set and the sales process nodes in the preset sales process node sequence to obtain matching results, and the sales process nodes are marked with status according to the matching results to generate the voice coverage status of each sales process node.

[0145] In this step, a sales process node refers to a stage unit in the standard automobile sales process, including a unique ID, name, time stamp, bound keyword list, and location index. A matching result refers to a set of keyword matching statuses. Voice coverage status refers to the keyword satisfaction level of a sales process node, calculated based on the matching results.

[0146] In an embodiment of the present application, the intelligent conference management system traverses each sales process node in a preset sales process node sequence to mark the status, and performs similarity matching on the standard keywords bound to each sales process node and the matching keyword set to obtain a matching result. The matching result refers to the matching status of each keyword. If the keyword matching rate required by the node is not met, the voice coverage status of the node is marked as uncovered, otherwise it is marked as covered, and the voice coverage status of each sales process node is obtained; for example, the safety configuration node needs to match the automatic emergency braking system, but actually only matches the automatic emergency braking system and is marked as uncovered.

[0147] Step 502: Expand the time tags of all sales process nodes to generate a node time range, extract the target gesture action occurring within the node time range from the gesture action set, and determine the action synchronization state of the target gesture action.

[0148] In this step, the node time range refers to the motion detection window extended by the time stamp. The target gesture action refers to the gesture trajectory segment within the node time range and with an amplitude exceeding the threshold, including the start frame, end frame, and 3D path data. The action synchronization status is a binary flag indicating the presence of the target gesture action, with presence indicating synchronization and absence indicating non-synchronization.

[0149] In an embodiment of the present invention, a node time range is generated based on the time mark of the sales process node. The node time range is a continuous time period from the start time range to the end time range. The start time range = node start time - (the average time of the salesperson's gesture triggering in advance × 1.5), and the end time range = node end time + (the average duration of the customer's question × 0.6). The target gesture actions occurring within the time range are scanned from the gesture action set. If there is a gesture trajectory segment with an amplitude exceeding the threshold, the action synchronization status is marked as synchronized, otherwise it is marked as not synchronized.

[0150] Step 503: Co-verify the voice coverage status and the motion synchronization status to identify the target node whose voice coverage is uncovered or whose motion synchronization status is not synchronized.

[0151] In this step, the target node refers to the abnormal node identified through collaborative verification, and the node ID and abnormal type combination are recorded.

[0152] In an embodiment of the present application, the voice coverage status and the action synchronization status are logically verified. When the voice coverage status is not covered or the action synchronization status is not synchronized, the sales process node is marked as the target node. The verification process introduces a node weight coefficient. Financial service nodes are directly marked if they are not synchronized, and security configuration nodes must be marked only if they are not covered at the same time.

[0153] Step 504: extract the location identifier of the target node and mark the deviation type to generate process deviation information.

[0154] In this step, the deviation type refers to the classification code defined according to the abnormal combination of speech and action.

[0155] In an embodiment of the present application, the position identifier of the target node, that is, the index number of the node in the process sequence, is extracted, and the deviation type is marked according to the deviation cause, including pure voice missing mark, pure action missing mark, and double missing mark, to generate process deviation information.

[0156] The embodiment of the present application improves the sales process deviation recognition accuracy and response efficiency through a multimodal collaborative decision-making mechanism, and simultaneously optimizes the sales conversion effect.

[0157] This application provides a specific embodiment, step 106, parsing the node position identifier in the process deviation information to generate a display control instruction, specifically including the following steps:

[0158] Step 601: extracting the node position identifier and the deviation type identifier of the target node from the process deviation information.

[0159] In an embodiment of the present application, the intelligent conference management system parses the process deviation information, extracts the node position identifier of the target node, that is, the digital index value of the node in the sales process sequence, and extracts the deviation type identifier, that is, the classification code of the pure voice missing mark, the pure action missing mark, and the double missing mark.

[0160] Step 602: Convert the node position identifier into target node area coordinates in the sales talk demonstration interface.

[0161] In this step, the target node area coordinates refer to the pixel coordinates of the display area assigned to a specific node in the sales talk demonstration interface, which are used to locate the status mark rendering position.

[0162] In an embodiment of the present application, a pre-stored interface layout mapping table is used, which records the pixel area corresponding to each position index. The node position identifier is directly matched with the index and coordinate range using a table lookup method to obtain the target node area coordinates in the sales talk demonstration interface.

[0163] Step 603: Match the deviation type identifier in the process deviation information with a preset state mark type rule library to determine graphic attributes and display parameters corresponding to the deviation type identifier.

[0164] In this step, the preset state marking type rule base refers to a database storing the mapping relationship between deviation types and visualization schemes.

[0165] In an embodiment of the present application, the deviation type identifier is input into a preset status mark type rule library for matching. The rule library stores display schemes corresponding to three types of deviation types. The pure voice missing mark matches the red exclamation mark icon and pulse frequency, the pure action missing mark matches the yellow palm icon and rotation animation, and the double missing mark matches the red and yellow flashing icon. The matching process is a key-value pair precise query to determine the graphic attributes and display parameters corresponding to the deviation type identifier.

[0166] Step 604: Encapsulate the target node region coordinates, the graphic attributes, and the display parameters into an instruction to generate a display control instruction.

[0167] In this step, graphic attributes refer to the static visual feature set of the status mark, including icon shape, basic color, and size. Display parameters refer to the dynamic behavior parameter set of the status mark, including animation type, frequency value, and transparency.

[0168] In an embodiment of the present application, the target node area coordinates, graphic attributes and display parameters are encapsulated, the target node area coordinates are converted into an array, the graphic attributes are converted into icon codes, and the display parameters are converted into animation parameter objects to generate display control instructions.

[0169] The embodiment of the present application realizes accurate identification and real-time feedback of process deviations through a multimodal collaborative verification mechanism.

[0170] Figure 2 The present invention provides a structural diagram of a voice AI intelligent conference management system. Figure 2 As shown, the system includes:

[0171] The acquisition module 21 is used to collect the voice data and gesture data of the salesperson;

[0172] A recognition module 22 is used to use voice AI to recognize the voice data and generate an ordered text sequence;

[0173] A matching module 23 is configured to match the keywords in the ordered text sequence with a preset term library in the field of automobile sales to obtain a set of matching keywords;

[0174] A combining module 24 is configured to extract a first type of action directed to a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combine the first type of action with the second type of action to generate a gesture action set;

[0175] A comparison module 25 is configured to compare the matching keyword set and the gesture action set with a preset sales process node sequence based on the intelligent conference management system to generate process deviation information;

[0176] The marking module 26 parses the node position identifier in the process deviation information, generates a display control instruction, and sends the display control instruction to the sales speech demonstration interface to display the status mark corresponding to each sales process node in real time, thereby realizing intelligent guidance and management of sales personnel.

[0177] Figure 2 The AI-based intelligent conference management system can be implemented Figure 1 The implementation principle and technical effects of the voice-based AI intelligent conference management method described in the illustrated embodiment will not be repeated here. The specific manner in which each module and unit performs operations in the voice-based AI intelligent conference management system in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.

[0178] In one possible design, Figure 2 A voice AI-based intelligent conference management system according to the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0179] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0180] The processing component 32 is used for the above Figure 1 The embodiment provides a voice AI-based intelligent conference management method.

[0181] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.

[0182] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0183] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0184] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0185] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0186] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0187] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The embodiment shown is a voice AI-based intelligent conference management method.

[0188] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0190] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A voice AI-based intelligent conference management method, characterized in that: include: Collect sales staff’s voice data and gesture data; Using voice AI to recognize the voice data and generate an ordered text sequence; Matching the keywords in the ordered text sequence with a preset term library in the field of automobile sales to obtain a set of matching keywords; Extracting a first type of action directed at a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combining the first type of action and the second type of action to generate a gesture action set; Based on the intelligent conference management system, the matching keyword set and the gesture action set are compared with a preset sales process node sequence to generate process deviation information; The node position identifier in the process deviation information is parsed to generate a display control instruction, and the display control instruction is sent to the sales speech demonstration interface to display the status mark corresponding to each sales process node in real time, thereby realizing intelligent guidance and management of sales personnel.

2. The method according to claim 1, characterized in that Using voice AI, the voice data is recognized to generate an ordered text sequence, including: Using voice AI, performing noise reduction processing on the voice data to obtain noise-reduced voice data; Performing frame processing on the noise-reduced voice data to obtain a continuous voice stream; Segmenting the continuous speech stream according to a preset time window to generate multiple speech segments; Extracting acoustic feature vectors from each speech segment and matching candidate syllable texts corresponding to the acoustic feature vectors; Based on a preset terminology library in the field of automobile sales, the candidate syllable text is subjected to vocabulary correction processing to generate a corrected text; Based on the segmentation order of the speech segments, the corrected text is reorganized in time sequence to generate an ordered text sequence.

3. The method according to claim 2, characterized in that Based on a preset terminology database in the field of automobile sales, the candidate syllable text is subjected to vocabulary correction processing to generate a corrected text, including: Extracting first acoustic feature parameters of each candidate syllable from the candidate syllable text, converting the first acoustic feature parameters into a plurality of first feature vectors, and combining all the first feature vectors into a candidate syllable feature sequence, wherein the first acoustic feature parameters include a fundamental frequency contour, a formant frequency, and a phoneme duration; Extracting a second acoustic feature parameter of each standard syllable from the standardized terms in the automobile sales field term library, converting the second acoustic feature parameter into a plurality of second feature vectors, and combining all the second feature vectors into a standardized term feature sequence; Calculating the syllable space distance between the candidate syllable feature sequence and the standardized term feature sequence, and selecting the standardized term with the smallest syllable space distance as the target candidate term; If the syllable space distance of the target candidate term is lower than the preset tolerance threshold, the complete text of the standardized term is extracted as the corrected text; if the syllable space distance of the target candidate term is higher than the preset tolerance threshold, the candidate syllable text is used as the corrected text.

4. The method according to claim 1, wherein The keywords in the ordered text sequence are matched with a preset term library in the field of automobile sales to obtain a set of matching keywords, including: Performing semantic boundary analysis on the text content in the ordered text sequence to obtain a plurality of candidate keywords; Extract vehicle performance parameter keywords and sales plan keywords from the preset automobile sales term library; Calculating the similarity between the candidate keywords and the vehicle performance parameter keywords and the sales plan keywords, and selecting candidate keywords that meet a preset matching threshold as valid keywords; All valid keywords are aggregated to generate a matching keyword set corresponding to the text content of the ordered text sequence.

5. The method according to claim 1, wherein Extracting a first type of action directed to a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combining the first type of action and the second type of action to generate a gesture action set, including: Performing three-dimensional spatial coordinate conversion on the gesture action data to generate a three-dimensional spatial gesture trajectory; Based on the preset position coordinates of the vehicle entity, selecting a first trajectory segment that meets the pointing condition from the three-dimensional gesture trajectory, and merging all the first trajectory segments into a first type of action pointing to the vehicle entity; Determine a node time window based on time stamps in a preset sales process node sequence, select second trajectory segments within the node time window whose action amplitude values ​​exceed a preset threshold from the three-dimensional gesture trajectory, and merge all second trajectory segments into a second type of action that emphasizes the sales node; The first type of actions and the second type of actions are combined to generate a gesture action set.

6. The method according to claim 1, characterized in that Based on the intelligent conference management system, the matching keyword set and the gesture action set are compared with a preset sales process node sequence to generate process deviation information, including: Based on the intelligent conference management system, keyword matching is performed on the matching keyword set with sales process nodes in a preset sales process node sequence to obtain matching results, and status marking is performed on the sales process nodes according to the matching results to generate a voice coverage status for each sales process node; Expanding the time stamps of all sales process nodes to generate a node time range, extracting a target gesture action occurring within the node time range from the gesture action set, and determining the action synchronization state of the target gesture action; Co-verifying the voice coverage state and the motion synchronization state to identify a target node whose voice coverage is uncovered or whose motion synchronization state is not synchronized; The location identifier of the target node is extracted and the deviation type is marked to generate process deviation information.

7. The method according to claim 1, characterized in that Parsing the node position identifier in the process deviation information to generate a display control instruction includes: Extracting a node position identifier and a deviation type identifier of a target node from the process deviation information; Convert the node position identifier into the target node area coordinates in the sales talk demonstration interface; Matching the deviation type identifier in the process deviation information with a preset state mark type rule library to determine graphic attributes and display parameters corresponding to the deviation type identifier; The target node area coordinates, the graphic attributes, and the display parameters are packaged into instructions to generate a display control instruction.

8. A voice AI-based intelligent conference management system, characterized in that: include: The acquisition module is used to collect the salesperson’s voice data and gesture data; A recognition module, configured to utilize voice AI to recognize the voice data and generate an ordered text sequence; A matching module, configured to match the keywords in the ordered text sequence with a preset term library in the field of automobile sales to obtain a set of matching keywords; an aggregation module, configured to extract a first type of action directed to a vehicle entity and a second type of action emphasizing a sales node from the gesture action data, and combine the first type of action and the second type of action to generate a gesture action set; a comparison module, configured to compare the matching keyword set and the gesture action set with a preset sales process node sequence based on the intelligent conference management system to generate process deviation information; The marking module parses the node position identifier in the process deviation information, generates a display control instruction, and sends the display control instruction to the sales speech demonstration interface to display the status mark corresponding to each sales process node in real time, thereby realizing intelligent guidance and management of sales personnel.

9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a voice AI intelligent conference management method as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, a voice AI intelligent conference management method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Service standardization method and device

    CN114022954A

  • Marketing verbal skill analysis and mining method based on deep reinforcement learning

    CN116303936A