Conference management method based on natural language understanding and related device

By using a natural language understanding model to perform intent reasoning and slot pair extraction on user audio data, the problem of insufficient fault tolerance in existing meeting management systems is solved, achieving more efficient and intelligent meeting management and improving user experience.

CN120977308APending Publication Date: 2025-11-18广州市迪士普音响科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511190512.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing voice assistant-based meeting management systems are not good at fault tolerance and are difficult to effectively improve the operational efficiency and intelligence of meeting management.

Method used

A natural language understanding model is used to infer intent from user audio data. Through slot pair extraction and similarity matching, the slot fields of users’ colloquial expressions are automatically corrected to ensure data accuracy. When information is insufficient, guided voice is provided to supplement the data, thus realizing the slot filling operation.

Benefits of technology

It improved the fault tolerance of the meeting management system, enhanced the operational efficiency and intelligence of meeting management, and improved the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977308A_ABST
    Figure CN120977308A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent conferences, and particularly provides a conference management method based on natural language understanding and a related device. Specifically, slot value pair extraction can be performed on user audio data to obtain a target slot name field and a target slot value field. If the target slot name field is the slot name field needing to be corrected, performing similarity matching on the target slot value field and each preset standard slot value field, determining a similar standard field similar to the target slot value field in each standard slot value field, and finally performing slot position filling by adopting the similar standard field and the target slot name field. Therefore, the fault-tolerant capability of the conference management system can be improved, and the user is allowed to perform voice interaction by adopting spoken expression, so that the operation efficiency and the intelligent degree of conference management can be effectively improved, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent meeting technology, and in particular to a meeting management method and related apparatus based on natural language understanding. Background Technology

[0002] Currently, existing technologies typically rely on web-based or mobile applications for meeting management. Taking meeting scheduling as an example, users need to log into the application and manually fill in information such as meeting time, location, topic, and attendees to complete the scheduling process. This method suffers from low operational efficiency and low levels of automation. To improve the convenience and automation of meeting management and enhance user experience, some solutions have proposed voice assistant-based meeting management systems, enabling users to manage functions such as meeting scheduling and inquiry through methods such as voice dialogue.

[0003] However, the inventors' research revealed that current voice assistant-based meeting management systems are not very fault-tolerant and are unable to effectively improve the operational efficiency and intelligence of meeting management. Summary of the Invention

[0004] The purpose of this application is to address at least one of the aforementioned technical deficiencies, particularly the technical deficiencies in the prior art that make it difficult to effectively improve the operational efficiency and intelligence level of meeting management.

[0005] In a first aspect, embodiments of this application provide a meeting management method based on natural language understanding, including:

[0006] Collect user audio data;

[0007] Based on the natural language understanding model and the user audio data, the valid meeting operation type is obtained by inferring the intent.

[0008] Based on the valid meeting operation type, slot-value pairs are extracted from the user audio data to obtain the target slot name field and the target slot value field.

[0009] If the target slot name field is a slot name field that needs to be corrected, then the target slot value field is matched with each preset standard slot value field to determine the similarity standard field;

[0010] The slot filling operation is performed using the similarity standard field and the target slot name field;

[0011] In response to the slot filling operation, if the filled slot data meets the preset operation triggering conditions, then the meeting management operation is performed based on the filled slot data; otherwise, the step of collecting user audio data is performed.

[0012] In some embodiments, the step of performing similarity matching between the target slot value field and each preset standard slot value field to determine the similarity standard field includes:

[0013] Calculate the field similarity between the target slot value field and each of the standard slot value fields;

[0014] The field similarity that is greater than or equal to the preset similarity threshold is taken as the target similarity.

[0015] If the number of target similarities is 1, then the standard slot value field corresponding to the target similarity will be used as the similarity standard field;

[0016] If the number of target similarities is greater than 1, then a first guiding voice is generated and played according to each candidate slot value field, and the step of collecting user audio data is executed; wherein, the candidate slot value field is the standard slot value field corresponding to the target similarity, and the first guiding voice is used to guide the user to select from each of the candidate slot value fields;

[0017] If the number of target similarities is less than 1, then the first guiding voice is generated and played.

[0018] In some embodiments, the step of performing intent reasoning based on a natural language understanding model and the user audio data to obtain a valid meeting operation type includes:

[0019] Based on the natural language understanding model, the user audio data is used to perform intent reasoning to obtain the inference meeting operation type and inference confidence.

[0020] If the inference confidence level is greater than or equal to the preset confidence threshold, then the inference meeting operation type is taken as the valid meeting operation type;

[0021] If the inference confidence is less than the preset confidence threshold, a second guiding voice is generated and played, and the step of collecting user audio data is executed; wherein, the second guiding voice is used to guide the user to re-express meeting operation information.

[0022] In some embodiments, the method further includes:

[0023] If the target slot name field is a time field, then the target slot value field is subjected to time standardization processing to obtain a structured time field, and the slot filling operation is performed using the structured time field and the target slot name field.

[0024] In some embodiments, the time normalization processing of the target slot value field to obtain a structured time field includes:

[0025] If the target slot value field includes a relative date field, then the relative date field is converted into an absolute date field according to the current system date, and the structured time field is obtained based on the absolute date field;

[0026] If the target slot value field includes a relative time field, then the absolute time field is determined according to the relative time field and the preset time correspondence, and the structured time field is obtained according to the absolute time field;

[0027] Time continuity verification is performed based on the filled data, the target slot name field, and the structured time field;

[0028] If the time continuity check fails, the structured time field is modified based on the preset time extension rules.

[0029] In some embodiments, determining whether the filled slot data meets the operation triggering condition includes:

[0030] Determine whether the filled slot data includes all slot value pairs required by the target slot template; wherein, the target slot template is the slot template corresponding to the valid meeting operation type;

[0031] If so, the operation triggering condition is determined to be met; otherwise, the operation triggering condition is determined not to be met.

[0032] In some embodiments, the method further includes:

[0033] If the operation triggering condition is not met, the unfilled slot name field is determined based on the filled slot data and the target slot template;

[0034] A third guiding voice is generated and played based on the unfilled slot name field; wherein the third guiding voice is used to guide the user to answer the slot value information corresponding to the unfilled slot name field.

[0035] Secondly, embodiments of this application provide a meeting management device based on natural language understanding, comprising:

[0036] The audio acquisition module is used to collect user audio data;

[0037] The intent reasoning module is used to perform intent reasoning based on the natural language understanding model and the user audio data to obtain the valid meeting operation type.

[0038] The slot-value pair extraction module is used to extract slot-value pairs from the user audio data based on the valid meeting operation type, and obtain the target slot name field and the target slot value field.

[0039] The similarity matching module is used to perform similarity matching between the target slot value field and each preset standard slot value field if the target slot name field is a slot name field that needs to be corrected, and to determine the similarity standard field.

[0040] The slot filling module is used to perform a slot filling operation using the similarity standard field and the target slot name field;

[0041] The operation execution module is used to respond to the execution of the slot filling operation. If the filled slot data meets the preset operation triggering conditions, the meeting management operation is executed according to the filled slot data; otherwise, the step of collecting user audio data is executed.

[0042] Thirdly, embodiments of this application provide a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the natural language understanding-based meeting management method described in any of the above embodiments.

[0043] Fourthly, embodiments of this application provide a conference assistant system, including: one or more processors, and a memory;

[0044] The memory stores computer-readable instructions, which, when executed by the one or more processors, perform the steps of the natural language understanding-based meeting management method described in any of the above embodiments.

[0045] In the natural language understanding-based meeting management method and related apparatus provided in some embodiments of this application, slot-value pairs can be extracted from user audio data to obtain a target slot name field and a target slot value field. If the target slot name field is a slot name field that needs to be corrected, the target slot value field can be matched with each pre-set standard slot value field for similarity. Among each standard slot value field, a similar standard field that is similar to the target slot value field is determined. Finally, the similar standard field and the target slot name field are used to fill the slots.

[0046] Therefore, this application can automatically replace user-provided slot fields with their corresponding standard slot fields through a similarity matching mechanism, and use the standard slot fields for slot filling to ensure data accuracy. This improves the fault tolerance of the meeting management system, allows users to interact via voice using colloquial expressions, effectively improving the operational efficiency and intelligence of meeting management, and enhancing the user experience. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is one of the flowcharts of a meeting management method based on natural language understanding in some embodiments;

[0049] Figure 2 This is a flowchart illustrating one of some embodiments of a meeting management method based on natural language understanding.

[0050] Figure 3 This is a schematic block diagram of a conference management device based on natural language understanding in some embodiments;

[0051] Figure 4 This is a diagram illustrating the internal structure of the meeting assistant system in some embodiments. Detailed Implementation

[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0053] In some embodiments, such as Figure 1 As shown, this application provides a meeting management method based on natural language understanding, which specifically includes the following steps:

[0054] S102: Collect user audio data.

[0055] In this step, the meeting assistant system can collect the user's voice to obtain user audio data. In some examples, the meeting assistant system can be activated by a specified keyword and enter a wake-up state. In the wake-up state, the meeting assistant system can collect the user's audio data and identify and determine the corresponding meeting management information through the user's audio data, so as to perform corresponding meeting management operations based on the identified meeting management information, thereby supporting users to express their meeting management needs using voice and natural language.

[0056] S104: Based on the natural language understanding model and user audio data, perform intent reasoning to obtain the valid meeting operation type.

[0057] In this step, the meeting assistant system can perform intent reasoning on the user's audio data based on a Natural Language Understanding (NLU) model and obtain a valid intent type. This valid intent type is the valid meeting operation type described in this application, used to represent the meeting management operation that the user wants to perform.

[0058] For example, valid meeting operation types may include, but are not limited to: creating a meeting, querying a meeting, canceling a meeting, notifying participants, screen sharing, projection, controlling meeting room equipment, etc.

[0059] S106: Extract slot-value pairs from user audio data based on valid meeting operation type to obtain target slot name field and target slot value field.

[0060] In this step, upon recognizing a valid intent, the meeting assistant system can identify and extract key information from the user's audio data based on the slot configuration required for the valid meeting operation type, thereby obtaining the target slot name field and the target slot value field. It can be understood that the target slot name field and the target slot value field form a slot-value pair, used to reflect the details of the meeting management operation. Taking meeting creation as an example, when the target slot name field is "room" and the target field is "Meeting Room 203", it indicates that the user wants to hold the meeting in Meeting Room 203.

[0061] It should be noted that different meeting operation types can correspond to different slot configurations. For example, the slot configuration for creating a meeting may include the meeting topic, meeting room, start time, end time, and participants; the slot configuration for querying a meeting may include the meeting room and query time; the slot configuration for canceling a meeting may include the meeting topic, meeting room, and meeting time; the slot configuration for notifying participants may include the notified person, notification content, meeting time, and meeting topic; the slot configuration for screen sharing may include shared device information and shared content type; the slot configuration for projection may include projection device information and projection content; and the slot configuration for controlling meeting room equipment may include the type of equipment to be controlled (such as lights, air conditioning, curtains, etc.), control actions (such as turn on, turn off, adjust, etc.), and control level (such as brightness level, temperature level).

[0062] S108: If the target slot name field is the slot name field that needs to be corrected, then the target slot value field is matched with each preset standard slot value field to determine the similar standard field.

[0063] In this step, considering that users' voice expressions in voice interaction scenarios such as meeting scheduling, meeting inquiry, and meeting notification are often colloquial, non-standard, and incomplete, the meeting assistant system can integrate and correct the extracted slot pairs to ensure accurate execution of meeting management operations according to user instructions. This corrects the colloquial, non-standard, and incomplete information into standard information that can be accurately understood, processed, and executed by the system devices.

[0064] Specifically, given the target slot name field and the target slot value field, the meeting assistant system can determine whether the target slot name field is a slot name field that needs correction. In some examples, the slot name field that needs correction may include non-time fields such as meeting room, meeting title, and device.

[0065] It is understandable that the meeting assistant system can implement the aforementioned judgment in any way. For example, the slot name field to be corrected can be obtained through pre-setting or pre-specification. After obtaining the target slot name field, the meeting assistant system can determine whether the obtained field is a pre-specified field; if so, it determines that the target slot name field is the slot name field to be corrected. Alternatively, the meeting assistant system can use a pre-trained field classification model to determine that the target slot name field is the slot name field to be corrected.

[0066] If the target slot name field is the slot name field that needs to be corrected, the meeting assistant system can compare the extracted target slot value field with each of the pre-set standard slot value fields to determine the similar standard field that is similar to the target slot value field among the standard slot value fields.

[0067] For example, when the target slot field is "Meeting Room 305", similarity matching can find "3F-305 Meeting Room" as a similarity standard field in various standard slot fields. Similarly, when the target slot field is "Project Summary Meeting", "2025 First Half-Year Project Summary Meeting" can be matched as a similarity standard field.

[0068] S110: Perform slot filling operation using the similar standard field and the target slot name field.

[0069] In this step, the meeting assistant system can treat the target slot name field and similar standard fields as a set of slot value pairs and perform a slot filling operation. It should be noted that the slot filling operation described in this application refers to the operation of integrating information from the obtained sets of slot value pairs, which may include, but is not limited to: filling the slot value pairs into the corresponding slot template, combining the information from each set of slot value pairs into a collection, etc.

[0070] S112: In response to the execution of the slot filling operation, if the filled slot data meets the preset operation triggering conditions, then the meeting management operation is executed based on the filled slot data; otherwise, the step of collecting user audio data is executed.

[0071] In this step, after each slot filling operation, the meeting assistant system determines whether the operation triggering conditions are met based on the filled slot data. If the conditions are met, it indicates that a meeting management operation can be triggered. In this case, the meeting assistant system can execute the corresponding meeting management operation based on the valid meeting operation type and the filled slot data. For example, when the valid meeting operation type is "create meeting," the meeting assistant system can automatically execute the meeting creation operation based on the filled fields such as meeting topic, meeting room, meeting time, and participants to create a meeting according to the user's requirements.

[0072] If the operation triggering conditions are not met, it indicates that the currently acquired information is insufficient to perform accurate meeting management operations. In this case, the meeting assistant system can return to step S102 to re-collect user audio data and execute steps S104~S112 based on the re-collected audio.

[0073] It is understood that the specific conditions for triggering the operation can be determined based on the actual situation. In some examples, this application can determine whether the filled slot data includes all the slot value pairs required by the target slot template. If so, the operation triggering condition is determined to be met; otherwise, the operation triggering condition is not met. Here, the target slot template refers to the slot template corresponding to the valid meeting operation type.

[0074] In this example, the corresponding slot template can be dynamically loaded based on the valid meeting operation type, and the completeness of the slot information determines whether to trigger the execution of the meeting management operation. If the slot information is complete, it indicates that the meeting assistant system has obtained enough information to form a complete intent data structure and trigger the execution of the corresponding meeting management operation.

[0075] Furthermore, in some examples, if the operation triggering conditions are not met, this application can determine the unfilled slot name field based on the filled slot data and the target slot template; and generate and play a third guiding voice based on the unfilled slot name field. The third guiding voice is used to guide the user to answer the slot value information corresponding to the unfilled slot name field. In this example, if the slot information is incomplete, this application can generate and play the third guiding voice based on the missing slot information to ask the user supplementary questions. After playing the third guiding voice, this application can proceed to step S102. This ensures that complete instruction data is ultimately formed.

[0076] This application utilizes a similarity matching mechanism to automatically replace user-provided slot fields with their corresponding standard slot fields, and then uses these standard slot fields for slot filling. This effectively solves the problem of inaccurate information caused by user expressions such as aliases, abbreviations, and vague statements, ensuring data accuracy. This application also improves the fault tolerance of the meeting management system, allows users to use colloquial expressions for voice interaction, thereby effectively improving the operational efficiency and intelligence of meeting management, and enhancing the user experience.

[0077] In some embodiments, the target slot value field is matched with each preset standard slot value field to determine the similarity standard fields, including:

[0078] Step A1: Calculate the field similarity between the target slot value field and each standard slot value field;

[0079] Step A3: Use the field similarity that is greater than or equal to the preset similarity threshold as the target similarity;

[0080] Step A5: If the number of target similarities is 1, then the standard slot value field corresponding to the target similarity is used as the similarity standard field;

[0081] Step A7: If the number of target similarities is greater than 1, then generate and play the first guiding voice according to each candidate slot value field, and execute the step of collecting user audio data; wherein, the candidate slot value field is the standard slot value field corresponding to the target similarity, and the first guiding voice is used to guide the user to select from each candidate slot value field;

[0082] Step A9: If the number of similarities to the target is less than 1, generate and play the first guiding voice.

[0083] In this embodiment, during the similarity matching process, if the similarity of all fields is less than the preset similarity threshold, or if there are multiple fields with similarity greater than or equal to the preset similarity threshold, supplementary questions can be asked through the first guiding voice to improve the accuracy of the slot information.

[0084] Specifically, the meeting assistant system can employ fuzzy matching algorithms such as edit distance, Jaccard, and vector semantic similarity to calculate the field similarity between the target slot field and N standard slot fields, thereby obtaining N field similarities, where N is a positive integer greater than 1. Subsequently, the meeting assistant system can filter the N field similarities using a preset similarity threshold, and use the field similarities that are not less than the preset similarity threshold as the target similarity.

[0085] If the target similarity count is 1, it means that only one of the N field similarities has a similarity greater than or equal to the preset similarity threshold. In this case, the meeting assistant system can directly use the standard slot field corresponding to the target similarity as the similarity standard field, and replace the target slot field with the standard slot field corresponding to the maximum field similarity.

[0086] If the number of target similarities is less than 1, it indicates that the similarity of all N fields is less than the preset similarity threshold. In this case, a first guiding voice can be generated and played to ask supplementary questions and guide the user to provide further answers, so that the meeting assistant system can obtain more accurate information.

[0087] If the number of target similarities is greater than 1, it indicates that among the N field similarities, there are multiple field similarities greater than or equal to the preset similarity threshold, meaning the matching items are not unique. In this case, the meeting assistant system can generate and play a first guiding voice based on the standard slot value field corresponding to each target similarity. This allows the user to be guided through the candidate slot value fields by the first guiding voice, thereby improving the accuracy of slot information through secondary confirmation.

[0088] In some embodiments, the method provided in this application may further include the following steps:

[0089] If the target slot name field is a time field, then the target slot value field is time-normalized to obtain a structured time field, and the slot filling operation is performed using the structured time field and the target slot name field.

[0090] In this embodiment, date and time information can be standardized to convert the colloquial target slot field into a structured time field. The structured time field and the target slot name field are then used to fill the slots, ensuring that the meeting assistant system can be used directly afterward, reducing errors and improving the accuracy of meeting management.

[0091] In some embodiments, the target slot value field is subjected to time normalization to obtain a structured time field, including:

[0092] If the target slot value field includes a relative date field, then the relative date field is converted to an absolute date field based on the current system date, and a structured time field is obtained based on the absolute date field;

[0093] If the target slot value field includes a relative time field, then the absolute time field is determined based on the relative time field and the preset time correspondence, and the structured time field is obtained based on the absolute time field.

[0094] Perform time continuity validation based on the populated data, the target slot name field, and the structured time field;

[0095] If the time continuity check fails, the structured time field will be modified based on the preset time extension rules.

[0096] In this embodiment, time standardization processing may include relative date parsing, time period default completion, and time continuity verification. Relative date parsing refers to converting relative date fields such as "tomorrow," "the day after tomorrow," and "last Wednesday" into absolute date fields such as yyyy-mm-dd (year-month-day) based on the current system date. Default time completion refers to converting ambiguous relative time fields such as "afternoon" and "evening" into specific hours and minutes according to a preset correspondence. For example, "afternoon" can correspond to "14:00," and "evening" can correspond to "19:00."

[0097] After obtaining the specific structured time fields including year, month, day, hour, and minute, this application can perform time continuity verification based on the populated data, the target slot name field, and the structured time field. For example, it can verify whether the meeting's end time is earlier than its start time; if so, the structured time field can be extended according to a preset time extension rule. For example, the end time can be increased by 12 hours or extended to the same time the following day.

[0098] Furthermore, in some examples, time standardization processing may also include time expression parsing and multi-time combination context awareness. Time expression parsing refers to supporting multiple time expression formats; for example, expressions such as "14:30," "2:30 PM," and "2:30 PM later" can be uniformly converted into the structured time field "14:30:00." Multi-time combination context awareness refers to maintaining consistency in time context within continuous dialogue and correctly resolving sequential relationships. For example, the date can be extracted first from the dialogue, and then the time can be extracted.

[0099] In some embodiments, performing intent reasoning based on a natural language understanding model and user audio data to determine the valid meeting operation type may include the following sub-steps:

[0100] Step C1: Perform intent reasoning on user audio data based on the natural language understanding model to obtain the inference meeting operation type and inference confidence;

[0101] Step C3: If the inference confidence is greater than or equal to the preset confidence threshold, then the inference meeting operation type is considered a valid meeting operation type;

[0102] Step C5: If the inference confidence is less than the preset confidence threshold, then generate and play the second guiding voice, and perform the step of collecting user audio data; wherein, the second guiding voice is used to guide the user to re-express the meeting operation information.

[0103] In this example, the NLU model can perform intent recognition and intent inference on user audio data, and output the inference intent type (i.e., the inference meeting operation type described in this application) and the inference confidence corresponding to the inference intent type. For example, the meeting assistant system can use speech recognition technology to convert user audio data into audio text, and input the audio text into the NLU model to obtain the inference result output by the NLU model.

[0104] Once the inference results are obtained, the meeting assistant system can compare the inference confidence level output by the model with the preset confidence threshold, and determine whether the inference intent type is a valid intent based on the comparison result.

[0105] Specifically, if the inference confidence is greater than or equal to a preset confidence threshold, the NLU model output can be determined as a valid intent. In this case, the meeting assistant system can determine the inference intent type of the model output as a valid meeting operation type and execute steps S106~S112.

[0106] Conversely, if the inference confidence is less than a preset confidence threshold, the output of the NLU model can be considered an invalid intent. In this case, the conference assistant system can generate and play a second guiding voice to guide the user to re-express the operation they want to perform via voice. After playing the second guiding voice, the conference assistant system can execute step S102 to re-acquire user audio data and execute steps S104~S112 based on the re-acquired user audio data.

[0107] This embodiment can quickly distinguish between valid and invalid intents by comparing the inference confidence level with a preset confidence threshold, thereby improving the response efficiency of the meeting assistant system.

[0108] To facilitate understanding of the solutions in this application, in some embodiments, such as Figure 2 As shown, this application provides a meeting management method based on natural language understanding, which specifically includes the following steps:

[0109] S202: In response to the wake-up keyword, enter the wake-up state;

[0110] S204: Collect user audio data;

[0111] S206: Convert user audio data into audio text;

[0112] S208: Perform intent reasoning on the audio text to obtain the inference intent type and inference confidence;

[0113] S210: Determine whether the inference confidence level is greater than or equal to the preset confidence threshold. If yes, proceed to step S218; otherwise, proceed to step S212.

[0114] S212: Generate feedback text and convert it to audio;

[0115] S214: Play audio;

[0116] S216: Enter hibernation mode;

[0117] S218: Extract slot pairs from the audio text to obtain the target slot pairs;

[0118] S220: Integrate and record slot contents based on the target slot value;

[0119] S222: Determine whether the slot content meets the operation trigger condition. If yes, proceed to step S228; otherwise, proceed to step S224.

[0120] S224: Generate a third guiding voice based on the unfilled slot;

[0121] S226: Play the audio and proceed to step S204;

[0122] S228: Execute the corresponding meeting operation instructions based on the meeting operation type and the recorded slot content.

[0123] This embodiment utilizes speech recognition and natural language understanding technologies to allow users to naturally express their appointment needs via voice. The system can automatically recognize the intent, extract key information, and proactively ask supplementary questions when information is missing, enabling multi-turn dialogue and contextual slot integration. Compared to traditional methods, this solution offers a more natural interaction, a more efficient process, and a more intelligent experience, effectively improving the convenience and intelligence of meeting appointments.

[0124] The following describes the conference management device based on natural language understanding provided in the embodiments of this application. The conference management device based on natural language understanding described below and the conference management method based on natural language understanding described above can be referred to in correspondence.

[0125] In some embodiments, such as Figure 3 As shown, this application provides a conference management device 300 based on natural language understanding, specifically including:

[0126] Audio acquisition module 302 is used to acquire user audio data;

[0127] Intent reasoning module 304 is used to perform intent reasoning based on the natural language understanding model and the user audio data to obtain the valid meeting operation type;

[0128] The slot value extraction module 306 is used to extract slot value pairs from the user audio data based on the valid meeting operation type to obtain the target slot name field and the target slot value field.

[0129] The similarity matching module 308 is used to perform similarity matching between the target slot value field and each preset standard slot value field if the target slot name field is a slot name field that needs to be corrected, and to determine the similar standard fields.

[0130] Slot filling module 310 is used to perform slot filling operation using the similarity standard field and the target slot name field;

[0131] The operation execution module 312 is used to respond to the execution of the slot filling operation. If the filled slot data meets the preset operation triggering conditions, the meeting management operation is executed according to the filled slot data; otherwise, the step of collecting user audio data is executed.

[0132] In some embodiments, the similarity matching module 308 of this application includes:

[0133] A similarity calculation unit is used to calculate the field similarity between the target slot value field and each of the standard slot value fields, respectively.

[0134] A filtering unit is used to select field similarities that are greater than or equal to a preset similarity threshold as target similarities;

[0135] A matching unit is configured to use the standard slot value field corresponding to the target similarity as the similarity standard field if the number of target similarities is 1.

[0136] The first guidance unit is configured to generate and play a first guidance voice based on each candidate slot value field if the number of target similarities is greater than 1, and to execute the step of collecting user audio data; wherein, the candidate slot value field is the standard slot value field corresponding to the target similarity, and the first guidance voice is used to guide the user to select from each of the candidate slot value fields;

[0137] The second guidance unit is used to generate and play the first guidance voice if the number of target similarities is less than 1.

[0138] In some embodiments, the intent reasoning module 304 of this application includes:

[0139] The intent reasoning unit is used to perform intent reasoning on the user audio data based on the natural language understanding model to obtain the inference conference operation type and inference confidence.

[0140] The effective intent determination unit is used to determine the inference meeting operation type as the effective meeting operation type if the inference confidence is greater than or equal to a preset confidence threshold.

[0141] An invalid intent determination unit is used to generate and play a second guiding voice if the reasoning confidence is less than the preset confidence threshold, and to perform the step of collecting user audio data; wherein the second guiding voice is used to guide the user to re-express meeting operation information.

[0142] In some embodiments, the apparatus 300 of this application may further include:

[0143] The time standardization module is used to perform time standardization processing on the target slot value field if the target slot name field is a time field, to obtain a structured time field, and to perform slot filling operation using the structured time field and the target slot name field.

[0144] In some embodiments, the time standardization module of this application includes:

[0145] A date conversion unit is configured to, if the target slot value field includes a relative date field, convert the relative date field into an absolute date field based on the current system date, and obtain the structured time field based on the absolute date field;

[0146] A time conversion unit is used to determine an absolute time field based on the relative time field and a preset time correspondence if the target slot value field includes a relative time field, and to obtain the structured time field based on the absolute time field.

[0147] A continuity verification unit is used to perform time continuity verification based on the filled data, the target slot name field, and the structured time field.

[0148] The time modification unit is used to modify the structured time field based on a preset time extension rule if the time continuity check fails.

[0149] In some embodiments, the operation execution module 312 of this application includes:

[0150] The triggering judgment unit is used to determine whether the filled slot data includes all the slot value pairs required by the target slot template; wherein, the target slot template is the slot template corresponding to the valid meeting operation type; if yes, it is determined that the operation triggering condition is met; otherwise, it is determined that the operation triggering condition is not met.

[0151] In some embodiments, the apparatus 300 of this application may further include:

[0152] The missing value determination module is used to determine the unfilled slot name field based on the filled slot data and the target slot template when the operation triggering condition is not met.

[0153] The supplementary question module is used to generate and play a third guiding voice based on the unfilled slot name field; wherein the third guiding voice is used to guide the user to answer the slot value information corresponding to the unfilled slot name field.

[0154] In one embodiment, this application also provides a storage medium storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the natural language understanding-based meeting management method as described in any embodiment.

[0155] In one embodiment, this application also provides a conference assistant system that stores computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the conference management method based on natural language understanding as described in any embodiment.

[0156] Indicatively, Figure 4 This is a schematic diagram of the internal structure of a meeting assistant system provided in an embodiment of this application. In one example, the meeting assistant system can be a server. (Refer to...) Figure 4 The conference assistant system 900 includes a processing component 902, which further includes one or more processors, and memory resources represented by memory 901 for storing instructions, such as applications, that can be executed by the processing component 902. The applications stored in memory 901 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 902 is configured to execute instructions to perform the steps of the natural language understanding-based conference management method described in any of the above embodiments.

[0157] The conference assistant system 900 may also include a power supply component 903 configured to perform power management of the conference assistant system 900, a wired or wireless network interface 904 configured to connect the conference assistant system 900 to a network, and an input / output (I / O) interface 905. The conference assistant system 900 can operate on an operating system stored in memory 901, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or similar.

[0158] Those skilled in the art will understand that the internal structure of the meeting assistant system shown in this application is merely a block diagram of a portion of the structure related to the solution of this application, and does not constitute a limitation on the meeting assistant system to which the solution of this application is applied. A specific meeting assistant system may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0159] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this document, "a," "an," "the," "the," and "its" may also include plural forms unless the context clearly indicates otherwise. "Multiple" refers to at least two, such as 2, 3, 5, or 8, etc. "And / or" includes any and all combinations of the related listed items.

[0160] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0161] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A meeting management method based on natural language understanding, characterized in that, include: Collect user audio data; Based on the natural language understanding model and the user audio data, the valid meeting operation type is obtained by inferring the intent. Based on the valid meeting operation type, slot-value pairs are extracted from the user audio data to obtain the target slot name field and the target slot value field. If the target slot name field is a slot name field that needs to be corrected, then the target slot value field is matched with each preset standard slot value field to determine the similarity standard field; The slot filling operation is performed using the similarity standard field and the target slot name field; In response to the slot filling operation, if the filled slot data meets the preset operation triggering conditions, then the meeting management operation is performed based on the filled slot data; otherwise, the step of collecting user audio data is performed.

2. The method according to claim 1, characterized in that, The step of performing similarity matching between the target slot value field and each preset standard slot value field to determine the similarity standard field includes: Calculate the field similarity between the target slot value field and each of the standard slot value fields; The field similarity that is greater than or equal to the preset similarity threshold is taken as the target similarity. If the number of target similarities is 1, then the standard slot value field corresponding to the target similarity will be used as the similarity standard field; If the number of target similarities is greater than 1, then a first guiding voice is generated and played according to each candidate slot value field, and the step of collecting user audio data is executed; wherein, the candidate slot value field is the standard slot value field corresponding to the target similarity, and the first guiding voice is used to guide the user to select from each of the candidate slot value fields; If the number of target similarities is less than 1, then the first guiding voice is generated and played.

3. The method according to claim 1, characterized in that, The process of inferring the effective meeting operation type based on the natural language understanding model and the user audio data includes: Based on the natural language understanding model, the user audio data is used to perform intent reasoning to obtain the inference meeting operation type and inference confidence. If the inference confidence level is greater than or equal to the preset confidence threshold, then the inference meeting operation type is taken as the valid meeting operation type; If the inference confidence is less than the preset confidence threshold, a second guiding voice is generated and played, and the step of collecting user audio data is executed; wherein, the second guiding voice is used to guide the user to re-express meeting operation information.

4. The method according to claim 1, characterized in that, The method further includes: If the target slot name field is a time field, then the target slot value field is subjected to time standardization processing to obtain a structured time field, and the slot filling operation is performed using the structured time field and the target slot name field.

5. The method according to claim 4, characterized in that, The step of performing time standardization on the target slot value field to obtain a structured time field includes: If the target slot value field includes a relative date field, then the relative date field is converted into an absolute date field according to the current system date, and the structured time field is obtained based on the absolute date field; If the target slot value field includes a relative time field, then the absolute time field is determined according to the relative time field and the preset time correspondence, and the structured time field is obtained according to the absolute time field; Time continuity verification is performed based on the filled data, the target slot name field, and the structured time field; If the time continuity check fails, the structured time field is modified based on the preset time extension rules.

6. The method according to any one of claims 1 to 5, characterized in that, Determining whether the filled slot data meets the operation triggering condition includes: Determine whether the filled slot data includes all slot value pairs required by the target slot template; wherein, the target slot template is the slot template corresponding to the valid meeting operation type; If so, the operation triggering condition is determined to be met; otherwise, the operation triggering condition is determined not to be met.

7. The method according to claim 6, characterized in that, The method further includes: If the operation triggering condition is not met, the unfilled slot name field is determined based on the filled slot data and the target slot template; A third guiding voice is generated and played based on the unfilled slot name field; wherein the third guiding voice is used to guide the user to answer the slot value information corresponding to the unfilled slot name field.

8. A meeting management device based on natural language understanding, characterized in that, include: The audio acquisition module is used to collect user audio data; The intent reasoning module is used to perform intent reasoning based on the natural language understanding model and the user audio data to obtain the valid meeting operation type. The slot-value pair extraction module is used to extract slot-value pairs from the user audio data based on the valid meeting operation type, and obtain the target slot name field and the target slot value field. The similarity matching module is used to perform similarity matching between the target slot value field and each preset standard slot value field if the target slot name field is a slot name field that needs to be corrected, and to determine the similarity standard field. The slot filling module is used to perform a slot filling operation using the similarity standard field and the target slot name field; The operation execution module is used to respond to the execution of the slot filling operation. If the filled slot data meets the preset operation triggering conditions, the meeting management operation is executed according to the filled slot data; otherwise, the step of collecting user audio data is executed.

9. A storage medium, characterized in that, The storage medium stores computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the natural language understanding-based conference management method as described in any one of claims 1 to 7.

10. A meeting assistant system, characterized in that, include: One or more processors, and memory; The memory stores computer-readable instructions, which, when executed by the one or more processors, perform the steps of the natural language understanding-based meeting management method as described in any one of claims 1 to 7.