Conference summary generation method and device based on artificial intelligence, electronic equipment and computer program product

By identifying and segmenting the speech of target individuals in meeting minutes, and utilizing pre-configured question templates and semantic analysis models, combined with voiceprint recognition technology, logically clear and complete meeting minutes are generated. This solves the accuracy and efficiency problems of large models when generating meeting minutes and avoids hallucination phenomena.

CN121902809APending Publication Date: 2026-04-21BEYONDSOFT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEYONDSOFT CORP
Filing Date
2025-11-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing large-scale modeling techniques suffer from insufficient processing capabilities for extremely long texts, difficulty in understanding complex logic, and frequent hallucinations when generating meeting minutes, resulting in inaccurate meeting minutes.

Method used

By identifying the speech of the target audience in the meeting minutes, a semantic analysis model driven by a pre-configured question template is used to analyze the speech fragments and generate a speech summary. Voiceprint recognition technology is used to accurately locate the speech content of the target audience, and prompt words are adjusted according to the meeting agenda to generate a logically clear and complete meeting minutes.

Benefits of technology

It improved the accuracy and quality of meeting minutes generation, significantly increased generation efficiency, avoided illusions, and ensured the authenticity and reliability of meeting minutes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121902809A_ABST
    Figure CN121902809A_ABST
Patent Text Reader

Abstract

The invention discloses a conference summary generation method and device based on artificial intelligence, electronic equipment and a computer program product. The method comprises the steps that a conference record is acquired, and the conference record is at least used for storing conference speaking of a plurality of speaking objects; identifying a target conference speech of the target object in conference speech of a plurality of speech objects stored in the conference record, the target conference speech including a plurality of speech segments; a pre-configured question template is adopted to drive a preset semantic analysis model to analyze the speaking segments, a speaking summary of each speaking segment is obtained, the question template is used for providing cue words for the semantic analysis model, and the cue words are used for guiding the semantic analysis model to extract target information from the speaking segments, the target information is used for generating a speaking summary; and splicing the speaking summaries corresponding to the plurality of speaking segments to generate a conference summary of the target object. The technical problem that the conference summary generated by the existing model is inaccurate is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and more specifically, to a method, apparatus, electronic device, and computer program product for generating meeting minutes based on artificial intelligence. Background Technology

[0002] In the current technological field, especially in enterprise management and organizational operations, AI-based meeting minutes generation primarily relies on traditional manual compilation methods. This method is not only time-consuming, but also often struggles to guarantee the accuracy and efficiency of information processing, particularly for large meetings. With the development of AI technology, especially the maturity of large-scale modeling, exploring intelligent and automated meeting minutes generation has become possible.

[0003] While existing large-model techniques have made significant progress in the field of natural language processing, their application in specific professional scenarios still faces many challenges. For the application scenario of meeting minutes generation, directly using large-model techniques has the following limitations:

[0004] 1. Insufficient processing capability for extremely long texts: When processing extremely long texts, large models are often limited by the length of the input and cannot receive and understand the entire meeting record at once. This affects their overall grasp of the meeting content and detailed analysis.

[0005] 2. Difficulty in understanding complex logic: In scenarios with many meeting agendas and complex discussion topics, large models may have difficulty accurately capturing the connections and logical structures between various topics, resulting in meeting minutes that lack a clear structure and coherence.

[0006] 3. Frequent hallucination phenomena: When faced with meeting content that is logically complex or has ambiguous information, the large model may exhibit "hallucination" phenomena, that is, generate information that does not conform to the facts, which will directly affect the authenticity and reliability of the meeting minutes.

[0007] In summary, existing large-scale modeling techniques, when applied to the automated and intelligent generation of meeting minutes, suffer from inaccurate generated meeting minutes.

[0008] There is currently no effective solution to the problem of inaccurate meeting minutes generated by the existing models. Summary of the Invention

[0009] This invention provides a method, apparatus, electronic device, and computer program product for generating meeting minutes based on artificial intelligence, in order to at least solve the technical problem of inaccurate meeting minutes generated by existing models.

[0010] According to one aspect of the present invention, an artificial intelligence-based meeting minutes generation method is provided, comprising: acquiring meeting records, wherein the meeting records are used to store meeting speeches by at least a plurality of speakers; identifying a target meeting speech of a target object among the meeting speeches of the plurality of speakers stored in the meeting records, wherein the target meeting speech includes a plurality of speech fragments; using a pre-configured question template to drive a pre-set semantic analysis model to analyze the speech fragments, obtaining a speech summary for each speech fragment, wherein the question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech fragments, the target information is used to generate the speech summary; and concatenating the speech summaries corresponding to the plurality of speech fragments to generate the meeting minutes of the target object.

[0011] Optionally, identifying the target meeting speech of a target object from among the meeting speeches of multiple speaking objects stored in the meeting record includes: when the meeting record is audio information, acquiring pre-collected voiceprint information of the target object, wherein the meeting record includes: multiple audio records, and the voiceprint information is used to describe the unique voice features of each speaking object; identifying at least one target audio record in the meeting record based on the voiceprint information of the target object, wherein the voice features of the target audio record match the voice features of the target object; and converting at least one target audio record into text form to obtain the target meeting speech.

[0012] Optionally, converting at least one of the target audio recordings into text form to obtain the target conference speech includes: identifying the timestamp of each target audio recording during the conference; grouping two target audio recordings whose time interval between the timestamps is less than a preset interval threshold into the same audio set; and converting the target audio recordings in multiple audio sets into text form to obtain the target conference speech, wherein multiple target audio recordings in the same audio set are converted into the same speech segment in the target conference speech.

[0013] Optionally, identifying the target meeting speech of the target object from among the meeting speeches of multiple speaking objects stored in the meeting record includes: when the meeting record is text information, identifying multiple target text records corresponding to the target object in the meeting record, wherein the meeting record includes: multiple text records, and the speaking object corresponding to each text record, each text record being the text conversion result of the speech of the same speaking object, and the target text record being the text record corresponding to the target object; segmenting the multiple target text records according to a preset word count requirement to obtain multiple speech segments; and generating the target meeting speech based on the multiple speech segments of the target object.

[0014] Optionally, before using a pre-configured question template to drive a pre-set semantic analysis model to analyze the speech fragments and obtain a speech summary for each speech fragment, the method further includes: obtaining a pre-configured question sample, wherein the question sample includes at least: fixed prompts and variable prompts, wherein the fixed prompts are used to provide preset guidance to the semantic analysis model, and the variable prompts are used to provide variable guidance to the semantic analysis model according to meeting requirements; configuring the variable prompts in the question sample according to a pre-acquired meeting agenda to obtain at least one question template, wherein the meeting agenda includes at least: the identity information of the target object and the meeting agenda, and the variable prompts are determined at least based on the identity information or the meeting agenda.

[0015] Optionally, configuring the variable prompts in the question sample to obtain the question template based on the pre-acquired meeting agenda includes: adjusting the variable prompts in the question template to preset prompts corresponding to the identity information based on the identity information in the meeting agenda, thereby obtaining the question template corresponding to the target object, wherein the question template is pre-configured with preset prompts corresponding to the identity information of multiple speaking objects.

[0016] Optionally, configuring the variable prompt words in the question sample to obtain the question template based on the pre-acquired meeting agenda includes: determining multiple meeting stages based on the meeting agenda in the meeting agenda, wherein the meeting agenda includes at least: multiple pre-set meeting stages; determining the target meeting stage matching each speech segment, wherein the target meeting stage is determined at least based on the speaking time of the speech segment proposed by the target object; querying the target prompt word adjustment strategy corresponding to each target meeting stage from multiple pre-set preset prompt word adjustment strategies, wherein the preset prompt word adjustment strategy is pre-set according to the meeting task requirements of each meeting stage; adjusting the variable prompt words in the question sample according to the target prompt word adjustment strategy corresponding to each target meeting stage to obtain the question template corresponding to each speech segment.

[0017] According to another aspect of the present invention, an artificial intelligence-based meeting minutes generation device is also provided, comprising: an acquisition module for acquiring meeting records, wherein the meeting records are used to store meeting speeches of at least a plurality of speakers; an identification module for identifying a target meeting speech of a target object among the meeting speeches of the plurality of speakers stored in the meeting records, wherein the target meeting speech includes a plurality of speech fragments; an analysis module for using a pre-configured question template to drive a pre-set semantic analysis model to analyze the speech fragments and obtain a speech summary for each speech fragment, wherein the question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech fragments, and the target information is used to generate the speech summary; and a splicing module for splicing the speech summaries corresponding to the plurality of speech fragments respectively to generate the meeting minutes of the target object.

[0018] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the artificial intelligence-based meeting minutes generation method through the computer program.

[0019] According to another aspect of the present invention, a computer program product is also provided, including computer instructions that, when executed by a processor, implement the steps of the artificial intelligence-based meeting minutes generation method.

[0020] The embodiments described above in this application identify and segment the meeting remarks of the target object in the meeting minutes, obtaining multiple speech fragments of the target object, with each speech fragment having a moderate length to facilitate model analysis. Then, a pre-configured question template is used to provide guiding prompts for the semantic analysis model, enabling the model to focus on extracting key target information and generating a speech summary for each speech fragment. Subsequently, the speech summaries of all speech fragments are spliced ​​together to form a logically clear and complete meeting minutes. This not only improves the accuracy and quality of meeting minutes generation but also significantly enhances generation efficiency, avoids the illusion situation that may occur when large models deal with complex problems, and solves the problem of inaccurate meeting minutes generated by existing models. Attached Figure Description

[0021] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0022] Figure 1 This is a flowchart of a meeting minutes generation method based on artificial intelligence according to an embodiment of the present invention;

[0023] Figure 2 This is a schematic diagram of an artificial intelligence-based meeting minutes generation device according to an embodiment of the present invention;

[0024] Figure 3 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0027] According to an embodiment of the present invention, an embodiment of a meeting minutes generation method based on artificial intelligence is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0028] Figure 1 This is a flowchart of a meeting minutes generation method based on artificial intelligence according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:

[0029] Step S102: Obtain meeting minutes, wherein the meeting minutes are used to store the meeting speeches of at least multiple speakers;

[0030] Step S104: Identify the target meeting speech of the target object among the meeting speeches of multiple speaking objects stored in the meeting records, wherein the target meeting speech includes multiple speech fragments;

[0031] Step S106: The pre-configured question template drives the pre-set semantic analysis model to analyze the speech segments and obtain a speech summary for each speech segment. The question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech segments, and the target information is used to generate the speech summary.

[0032] Step S108: Summarize the speech summaries corresponding to multiple speech fragments to generate the meeting minutes of the target object.

[0033] The embodiments described above in this application identify and segment the meeting remarks of the target object in the meeting minutes, obtaining multiple speech fragments of the target object, with each speech fragment having a moderate length to facilitate model analysis. Then, a pre-configured question template is used to provide guiding prompts for the semantic analysis model, enabling the model to focus on extracting key target information and generating a speech summary for each speech fragment. Subsequently, the speech summaries of all speech fragments are spliced ​​together to form a logically clear and complete meeting minutes. This not only improves the accuracy and quality of meeting minutes generation but also significantly enhances generation efficiency, avoids the illusion situation that may occur when large models deal with complex problems, and solves the problem of inaccurate meeting minutes generated by existing models.

[0034] In step S104 above, the target audience can be specific participants in the meeting, such as the host or a leader. The target audience's remarks at the meeting play an important role in the meeting. For example, the host's remarks can guide the meeting; the leader's remarks can influence the decision-making tendency of the meeting and clarify the meeting objectives. In other words, the target audience's target remarks at the meeting have a significant impact on the entire meeting. Therefore, by identifying the target audience's target remarks at the meeting and based on those remarks, the meeting minutes can be made more consistent with the actual situation of the meeting.

[0035] In step S104 above, the semantic analysis model can be a large AI model.

[0036] In step S106 above, the question template is used to provide prompt words to the semantic analysis model. These prompt words can instruct the semantic analysis model to focus on specific dimensions of the speech fragments in a biased manner, so as to obtain an accurate summary of the speech.

[0037] In step S108 above, the summaries of the speeches corresponding to the multiple speech segments can be pieced together in a way similar to a jigsaw puzzle, such as by first piecing together small pieces and then combining them into a larger picture, thus piecing together multiple speech summaries into meeting minutes.

[0038] To improve the accuracy and effectiveness of meeting minutes generation, this application integrates voiceprint recognition technology to accurately locate the speech content of the target subject.

[0039] As an optional embodiment, identifying the target speaker's speech among multiple speakers stored in the meeting minutes includes: when the meeting minutes are audio information, acquiring pre-collected voiceprint information of the target speaker, wherein the meeting minutes include multiple audio recordings, and the voiceprint information is used to describe the unique voice features of each speaker; identifying at least one target audio recording in the meeting minutes based on the target speaker's voiceprint information, wherein the voice features of the target audio recording match the voice features of the target speaker; and converting at least one target audio recording into text form to obtain the target speech.

[0040] In the embodiments described above, the meeting record storage module maintains the voiceprint information of each speaker. This information is collected and stored before the meeting begins and describes the unique voice characteristics of each participant. When it is necessary to identify the speech of a target speaker, the system first analyzes multiple voice recordings in the meeting minutes based on the stored voiceprint information, and selects at least one target voice recording that matches the voice characteristics of the target speaker. Next, the selected voice recording is converted into text. This process utilizes advanced speech-to-text technology to ensure the accuracy and speed of the conversion. Through accurate voiceprint recognition and efficient speech-to-text conversion, the entire speech of the target speaker can be quickly located and analyzed, providing detailed basic materials for subsequent topic analysis, viewpoint extraction, and decision summarization, ensuring the comprehensiveness and accuracy of the meeting minutes.

[0041] Optionally, the system performance can be further improved by refining the voiceprint recognition algorithm to increase the recognition rate or by optimizing the noise filtering mechanism in the speech-to-text process.

[0042] Optionally, voiceprint information is not limited to collection during the preliminary preparation stage of meeting recording, but can also be updated in real time during the meeting to enhance the system's adaptability and robustness.

[0043] As an optional embodiment, converting at least one target audio recording into text form to obtain a target conference speech includes: identifying the timestamp of each target audio recording during the conference; classifying two target audio recordings whose time interval between timestamps is less than a preset interval threshold into the same audio set; and converting the target audio recordings in multiple audio sets into text form to obtain a target conference speech, wherein multiple target audio recordings in the same audio set are converted into the same speech segment in the target conference speech.

[0044] In the above embodiments of this application, during the process of converting target speech recordings into text, the timestamp of each target speech recording during the meeting can be identified. The timestamp is used to accurately locate the speech recording in time to ensure the accuracy of the conversion. Then, two target speech recordings with a time interval of less than a preset interval threshold are divided into the same speech set. This division strategy utilizes the principle of temporal continuity, which helps to integrate coherent speech content and avoids the incorrect integration of discontinuous content with excessively large time intervals.

[0045] The embodiments described above convert target speech records from multiple speech sets into text form to obtain target meeting speeches. Multiple target speech records from the same speech set are converted into the same speech segment within the target meeting speech. This conversion method not only improves the efficiency of speech-to-text conversion but also ensures the integrity and coherence of the speech content, providing a structured text foundation for subsequent meeting minutes generation. Through this process, speech information from meetings can be accurately and efficiently converted into text, laying a solid foundation for subsequent intelligent analysis and minutes generation.

[0046] Optionally, multiple speech segments can be processed using different time interval thresholds or speech set segmentation strategies to adapt to the meeting recording needs of different scenarios.

[0047] For meeting speeches from multiple speakers stored in meeting minutes, this application can accurately identify multiple target text records corresponding to a target speaker in the text information format of the meeting minutes.

[0048] As an optional embodiment, identifying the target meeting speech of the target object from the meeting speeches of multiple speakers stored in the meeting records includes: when the meeting records are text information, identifying multiple target text records corresponding to the target object in the meeting records, wherein the meeting records include: multiple text records, and a speaker corresponding to each text record, each text record being the text conversion result of the speech of the same speaker, and the target text record being the text record corresponding to the target object; segmenting the multiple target text records according to a pre-set word count requirement to obtain multiple speech fragments; and generating the target meeting speech based on the multiple speech fragments of the target object.

[0049] In the embodiments described above, the meeting minutes consist of multiple text records, each associated with a specific speaker. Specifically, each text record represents the textual result of a speaker's speech converted into text, while the target text record specifically refers to the speech record of the target speaker. Next, to optimize subsequent analysis and processing, these target text records are reasonably segmented according to pre-set word count requirements, generating a series of speech fragments. This method decomposes the target speaker's lengthy speech content, ensuring efficiency and accuracy during large-scale model processing. Finally, based on the multiple speech fragments of the target speaker, a complete target meeting speech content is generated through comprehensive analysis and intelligent integration. This process fully utilizes the semantic understanding and text generation capabilities of the large-scale model, ensuring the quality and completeness of the meeting minutes. This technical solution not only improves the efficiency of meeting minutes generation but also significantly enhances the accuracy and professionalism of the minutes, especially when processing complex meeting content, where it has a more pronounced advantage.

[0050] Alternatively, meeting minutes can also be in non-text formats, such as audio or video recordings. As long as they can be converted into a format suitable for large-scale model analysis using appropriate conversion methods, the automated generation of meeting minutes can also be achieved. Through this flexible processing method, meeting minutes can be generated efficiently and accurately regardless of the initial format of the meeting minutes, meeting diverse meeting management needs.

[0051] As an optional embodiment, before using a pre-configured question template to drive a pre-set semantic analysis model to analyze speech fragments and obtain a speech summary for each speech fragment, the method further includes: obtaining a pre-configured question sample, wherein the question sample includes at least: fixed prompts and variable prompts, wherein the fixed prompts are used to provide preset guidance to the semantic analysis model, and the variable prompts are used to provide variable guidance to the semantic analysis model according to the meeting requirements; configuring the variable prompts in the question sample according to a pre-acquired meeting agenda to obtain at least one question template, wherein the meeting agenda includes at least: the identity information of the target object and the meeting agenda, and the variable prompts are determined at least based on the identity information or the meeting agenda.

[0052] In the embodiments described above, before analyzing the speech segments of the target audience, a pre-configured question sample can be obtained. This question sample includes fixed prompts and variable prompts. The fixed prompts provide pre-defined guidance for the semantic analysis model, while the variable prompts are dynamically adjusted according to meeting needs, providing targeted guidance to the model. By configuring the variable prompts in the question sample according to the meeting agenda, a question template adapted to the characteristics of the meeting can be generated. Using the question template to guide the model in analysis can improve the accuracy of the model analysis, ensuring that the generated meeting minutes are more relevant to the essence of the meeting, have a clearer logical structure, and effectively improve the quality and efficiency of the minutes.

[0053] Specifically, the meeting agenda contains the identity information of the target audience and the meeting agenda. Based on this, the system determines variable prompts to achieve customized analysis for speakers with different identities and for different topics.

[0054] Alternatively, question templates can also be generated in other ways, such as dynamically adjusting fixed and variable prompts based on the attendees' positions and meeting objectives, to provide multi-dimensional guidance for the semantic analysis model and further enhance the flexibility and applicability of minutes generation.

[0055] As an optional embodiment, the process of configuring variable prompts in the question sample to obtain a question template based on the pre-acquired meeting agenda includes: adjusting the variable prompts in the question template to preset prompts corresponding to the identity information based on the identity information in the meeting agenda, thereby obtaining a question template corresponding to the target object. The question template is pre-configured with preset prompts corresponding to the identity information of multiple speaking objects.

[0056] In the embodiments described above, the question template corresponding to the target object can be generated based on the identity information in the meeting agenda. According to the identity information of different speakers in the meeting agenda, the variable prompt words in the question sample can be dynamically replaced with preset prompt words that match the identity information, creating a personalized question template for the target object. Based on this question template, the semantic analysis model can be guided, which can improve the semantic analysis model's understanding of the speech content of a specific speaker (such as the speech fragment of the target object). This is because the preset prompt words are customized according to the speaker's role and possible speaking topics, making the large model more focused and accurate when analyzing meeting records.

[0057] In the embodiments described above, the use of personalized question templates can more effectively extract key information from meeting minutes, optimize the quality and efficiency of information extraction, reduce analysis errors caused by insufficient template universality, and improve the accuracy and professionalism of meeting minutes.

[0058] Optionally, the configuration strategy for the question template can also be flexibly adjusted. For example, prompts can be dynamically configured based on factors such as meeting type and the importance of the topic, further enhancing the system's ability to adapt to different types of meetings.

[0059] As an optional embodiment, configuring variable prompts in the question sample to obtain a question template based on a pre-acquired meeting agenda includes: determining multiple meeting stages based on the meeting agenda in the meeting agenda, wherein the meeting agenda includes at least: multiple pre-set meeting stages; determining the target meeting stage matching each speaking segment, wherein the target meeting stage is determined at least based on the speaking time of the speaking segment proposed by the target object; querying the target prompt adjustment strategy corresponding to each target meeting stage from multiple pre-set preset prompt adjustment strategies, wherein the preset prompt adjustment strategy is pre-set according to the meeting task requirements of each meeting stage; adjusting the variable prompts in the question sample according to the target prompt adjustment strategy corresponding to each target meeting stage to obtain a question template corresponding to each speaking segment.

[0060] In the embodiments described above, multiple meeting stages can be determined based on the content of the meeting agenda. Each meeting stage is associated with a specific meeting task. For each meeting stage's task, a corresponding preset prompt word adjustment strategy can be pre-configured. Since each speaking segment represents a speech made by the target audience at each meeting stage, by matching the target meeting stage corresponding to each speaking segment and adjusting the variable prompt words in the question sample according to the target prompt word adjustment strategy for each target meeting stage, the most suitable question template can be generated for each speaking segment. This technical solution ensures the relevance and accuracy of the questions, improves the efficiency and quality of large model analysis, and intelligently generates question templates based on different meeting stages and tasks, further optimizing the automated generation process of meeting minutes.

[0061] Alternatively, the target meeting phase can also be determined by analyzing the format of the meeting agenda or by identifying keywords, to accommodate a wider range of meeting types and task requirements.

[0062] The present invention also provides a preferred embodiment, which provides a meeting minutes generation method. In the typical scenario of using large models to automatically and intelligently generate meeting minutes, this method employs technologies such as multi-level AI collaboration, dynamic slicing, result aggregation algorithms, structured storage, and one-click generation mechanisms. This method can intelligently and automatically generate high-quality meeting minutes with clear logic and complete content, even in cases of extremely long meeting content, complex meeting agendas, and complex logical structures. This avoids the situation where large models may misinterpret meetings with complex logical structures.

[0063] Optionally, the meeting minutes generation method provided in this application can be applied to a meeting management system, which includes: a meeting record intelligent analysis subsystem, a multi-level AI inference engine, and a structured storage module. When a user triggers a minutes generation command, the system automates the process through the following innovative workflow:

[0064] 1) Standardize the format of the original meeting minutes and extract key content segments;

[0065] 2) Employ a hierarchical recursive large model technique to perform issue analysis, viewpoint extraction, and resolution identification in stages;

[0066] 3) Integrate AI output through a dynamic result aggregation algorithm;

[0067] 4) Automatically extract structured data such as time and attendees;

[0068] 5) Generate standardized meeting minutes and persist them.

[0069] The entire process breaks through the traditional manual compilation mode, achieving an efficiency improvement of over 300% in meeting minutes generation.

[0070] Alternatively, the meeting management system acts like an intelligent meeting secretary. When a user issues an instruction, it automatically completes the following tasks: retrieves the meeting minutes book → extracts key content → consults an AI expert team → organizes the minutes into a formal summary → archives the minutes in a filing cabinet → and finally delivers the final product to the user for review.

[0071] As an optional embodiment, the specific working steps of the conference management system are as follows:

[0072] Users issue commands on the system page;

[0073] Select a specific meeting and click the "Start Intelligent Minute Generation" button (equivalent to notifying the intelligent secretary: "Please compile the minutes of meeting number XX").

[0074] Intelligent back-end processing (secretary workflow).

[0075] As an optional example, background intelligent processing (secretary workflow) includes the following steps:

[0076] The first stage involves organizing the original materials;

[0077] The second phase involves core analysis of the large AI model.

[0078] The third stage is information integration.

[0079] Optionally, the first stage involves organizing the original materials, specifically including:

[0080] S11 is used to retrieve the original audio / text recordings of the meeting and organize the messy content into a standard format (equivalent to printing a messy handwritten draft into a neat document).

[0081] S12, special processing of leaders' speeches, identifies all speech segments and cuts long speeches into smaller segments (like cutting a long video into a short TikTok video).

[0082] Optionally, the second phase involves core analysis of the large AI model, specifically including:

[0083] S21, the system prepares a set of "question templates" (like a list of questions a reporter prepares before an interview).

[0084] It should be noted that these "question templates" are a series of "formulas" or "patterns" designed by the system to analyze leaders' speeches. Specifically, they take the form of large model prompts. The system prepares several prompts in advance according to the characteristics of the meeting type and the common patterns of leaders' speeches. These prompts contain both fixed and variable content. Each time they are called, the system fills in the corresponding dynamic changes based on the specific circumstances of the meeting and the task at hand, combining them to form a complete prompt, which is then provided to the large model.

[0085] S22, First round of AI large model processing (analysis of the core content of the conference).

[0086] Specifically, the leader's speech segments are sent to the AI ​​model in batches; the AI ​​model identifies the core viewpoints of each speech; and the scattered analysis results are pieced together into a complete conclusion (similar to a jigsaw puzzle: first put together small pieces and then combine them into a large picture).

[0087] S23, Second round of AI large model processing (analysis of the meeting agenda).

[0088] Specifically, retrieve the meeting agenda; break down major topics into smaller discussion points; query the AI ​​model for each smaller discussion point; and record the model's conclusions for each topic.

[0089] It should be noted that this approach combines large-scale model capabilities with engineering capabilities. First, an engineering method is used to format the "Meeting Agenda," searching for keywords such as "agenda," "topics," and "speaking order." Using these keyword locations as clues, and combining them with the document format information of the "Meeting Agenda," paragraphs are determined. Then, the semantic understanding and summarizing capabilities of the large-scale model are utilized to analyze each paragraph separately, extracting the core points of each paragraph to form a sequence of "topic points" and "discussion points."

[0090] Optionally, the third stage, information integration, specifically includes:

[0091] Step S31: Extract basic information. The AI ​​model identifies key information such as meeting time and attendee list (similar to automatically extracting expense report elements from a jumble of invoices).

[0092] Step S32, final synthesis.

[0093] Specifically, based on the format requirements of the "Meeting Minutes Template," engineering methods are used to break it down into different modules, such as meeting information (including time, location, attendees, positions, etc.), minutes sections (such as "Meeting Points," "Meeting Clarifications," "Meeting Decisions," and "To-Dos"), etc. Relatively formatted content is directly pieced together from the materials obtained in the previous step using engineering methods. More generalized parts, strongly related to the meeting content, are then processed using prompt word engineering methods. These parts, combined with the meeting minutes format information, word count control, and relational requirements, are submitted to a large model. Leveraging the model's semantic understanding, text generation capabilities, and artificial intelligence, the relevant sections of the "Meeting Minutes" are processed and written to meet the requirements, and then output. Finally, engineering methods are used to format and adjust all content, resulting in the final complete version of the "Meeting Minutes."

[0094] Step S33: Automatic archiving and saving. The system saves the generated minutes into the database according to a standard format (similar to a secretary classifying and placing documents into labeled filing cabinets).

[0095] Step S34: The user views the results. Click "View Meeting Minutes" on the page; the system instantly retrieves the prepared minutes from the database (like saying "Please give me the XX meeting minutes," and the secretary immediately takes them from the filing cabinet).

[0096] According to an embodiment of the present invention, an embodiment of an AI-based meeting minutes generation device is also provided. It should be noted that the AI-based meeting minutes generation device can be used to execute the AI-based meeting minutes generation method in the embodiment of the present invention, and the AI-based meeting minutes generation method in the embodiment of the present invention can be executed in the AI-based meeting minutes generation device.

[0097] Figure 2 This is a schematic diagram of an artificial intelligence-based meeting minutes generation device according to an embodiment of the present invention, such as... Figure 2As shown, the device may include: an acquisition module 22 for acquiring meeting records, wherein the meeting records are used to store meeting speeches from at least multiple speakers; an identification module 24 for identifying a target meeting speech from the multiple speakers stored in the meeting records, wherein the target meeting speech includes multiple speech fragments; an analysis module 26 for using a pre-configured question template to drive a pre-set semantic analysis model to analyze the speech fragments and obtain a speech summary for each speech fragment, wherein the question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech fragments, and the target information is used to generate a speech summary; and a splicing module 28 for splicing the speech summaries corresponding to multiple speech fragments to generate a meeting minutes for the target object.

[0098] It should be noted that the acquisition module 22 in this embodiment can be used to execute step S102 in this application embodiment, the identification module 24 in this embodiment can be used to execute step S104 in this application embodiment, the analysis module 26 in this embodiment can be used to execute step S106 in this application embodiment, and the splicing module 28 in this embodiment can be used to execute step S108 in this application embodiment. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments.

[0099] The embodiments described above in this application identify and segment the meeting remarks of the target object in the meeting minutes, obtaining multiple speech fragments of the target object, with each speech fragment having a moderate length to facilitate model analysis. Then, a pre-configured question template is used to provide guiding prompts for the semantic analysis model, enabling the model to focus on extracting key target information and generating a speech summary for each speech fragment. Subsequently, the speech summaries of all speech fragments are spliced ​​together to form a logically clear and complete meeting minutes. This not only improves the accuracy and quality of meeting minutes generation but also significantly enhances generation efficiency, avoids the illusion situation that may occur when large models deal with complex problems, and solves the problem of inaccurate meeting minutes generated by existing models.

[0100] As an optional embodiment, the recognition module includes: an acquisition unit, configured to acquire pre-collected voiceprint information of a target object when the meeting record is voice information, wherein the meeting record includes: multiple voice records, and the voiceprint information is used to describe the unique voice features of each speaker; a first recognition unit, configured to recognize at least one target voice record in the meeting record based on the voiceprint information of the target object, wherein the voice features of the target voice record match the voice features of the target object; and a conversion unit, configured to convert at least one target voice record into text form to obtain the target meeting speech.

[0101] As an optional embodiment, the conversion unit includes: an identification subunit for identifying the timestamp of each target audio recording during the conference; a division subunit for dividing two target audio recordings whose time interval between timestamps is less than a preset interval threshold into the same audio set; and a conversion subunit for converting the target audio recordings in multiple audio sets into text form to obtain the target conference speech, wherein multiple target audio recordings in the same audio set are converted into the same speech segment in the target conference speech.

[0102] As an optional embodiment, the recognition module includes: a second recognition unit, used to recognize multiple target text records corresponding to a target object in the meeting records when the meeting records are text information, wherein the meeting records include: multiple text records and a speaker corresponding to each text record, each text record being a text conversion result of the speech of the same speaker, and the target text record being the text record corresponding to the target object; a segmentation unit, used to segment the multiple target text records according to a preset word count requirement to obtain multiple speech fragments; and a generation unit, used to generate a target meeting speech based on the multiple speech fragments of the target object.

[0103] As an optional embodiment, the apparatus further includes: an acquisition submodule, configured to acquire a pre-configured question sample before analyzing speech segments using a pre-configured question template-driven pre-set semantic analysis model to obtain a speech summary for each speech segment, wherein the question sample includes at least: fixed prompts and variable prompts, wherein the fixed prompts are used to provide preset guidance to the semantic analysis model, and the variable prompts are used to provide variable guidance to the semantic analysis model according to meeting requirements; and a configuration submodule, configured to configure the variable prompts in the question sample according to a pre-acquired meeting agenda to obtain at least one question template, wherein the meeting agenda includes at least: the identity information of the target object and the meeting agenda, and the variable prompts are determined at least based on the identity information or the meeting agenda.

[0104] As an optional embodiment, the configuration submodule includes: a first configuration unit, used to adjust the variable prompts in the question template to preset prompts corresponding to the identity information based on the identity information in the meeting agenda, so as to obtain the question template corresponding to the target object, wherein the question template is pre-configured with preset prompts corresponding to the identity information of multiple speaking objects respectively.

[0105] As an optional embodiment, the configuration submodule includes: a first determining unit, configured to determine multiple meeting stages based on the meeting agenda in the meeting agenda table, wherein the meeting agenda includes at least: multiple pre-set meeting stages; a second determining unit, configured to determine the target meeting stage matching each speaking segment, wherein the target meeting stage is determined at least based on the speaking time of the speaking segment proposed by the target object; a querying unit, configured to query the target prompt word adjustment strategy corresponding to each target meeting stage from among multiple pre-set preset prompt word adjustment strategies, wherein the preset prompt word adjustment strategy is pre-set according to the meeting task requirements of each meeting stage; and an adjusting unit, configured to adjust the variable prompt words in the question sample according to the target prompt word adjustment strategy corresponding to each target meeting stage, to obtain the question template corresponding to each speaking segment.

[0106] Embodiments of the present invention can provide an electronic device, which can be a computer terminal, and the computer terminal can be any one of a group of computer terminal devices. Optionally, in this embodiment, the computer terminal can also be replaced by a mobile terminal or other terminal device.

[0107] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0108] In this embodiment, the aforementioned computer terminal can execute the program code for the following steps in the AI-based meeting minutes generation method: acquiring meeting records, wherein the meeting records are used to store meeting speeches from at least multiple speakers; identifying the target meeting speech of the target object among the meeting speeches of the multiple speakers stored in the meeting records, wherein the target meeting speech includes multiple speech fragments; using a pre-configured question template to drive a pre-set semantic analysis model to analyze the speech fragments and obtain a speech summary for each speech fragment, wherein the question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech fragments, and the target information is used to generate a speech summary; and concatenating the speech summaries corresponding to multiple speech fragments to generate the meeting minutes of the target object.

[0109] Figure 3 This is a structural block diagram of a computer terminal according to an embodiment of the present invention, such as... Figure 3 As shown, the computer terminal 30 may include one or more (only one is shown in the figure) processors 32 and memory 34.

[0110] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the AI-based meeting minutes generation method and apparatus in this embodiment of the invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned AI-based meeting minutes generation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal 30 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0111] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: acquiring meeting minutes, wherein the meeting minutes are used to store meeting speeches from at least multiple speakers; identifying the target meeting speech of the target object among the meeting speeches of the multiple speakers stored in the meeting minutes, wherein the target meeting speech includes multiple speech fragments; using a pre-configured question template to drive a pre-set semantic analysis model to analyze the speech fragments and obtain a speech summary for each speech fragment, wherein the question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech fragments, and the target information is used to generate a speech summary; concatenating the speech summaries corresponding to multiple speech fragments to generate the meeting minutes of the target object.

[0112] Optionally, the processor may also execute program code for the following steps: when the meeting record is voice information, acquire pre-collected voiceprint information of the target object, wherein the meeting record includes: multiple voice records, and the voiceprint information is used to describe the unique voice features of each speaker; based on the voiceprint information of the target object, identify at least one target voice record in the meeting record, wherein the voice features of the target voice record match the voice features of the target object; convert at least one target voice record into text form to obtain the target meeting speech.

[0113] Optionally, the processor may also execute program code for the following steps: identifying the timestamp of each target audio recording during the conference; classifying two target audio recordings whose time interval between timestamps is less than a preset interval threshold into the same audio set; converting the target audio recordings in multiple audio sets into text form to obtain the target conference speech, wherein multiple target audio recordings in the same audio set are converted into the same speech segment in the target conference speech.

[0114] Optionally, the processor may also execute program code for the following steps: when the meeting minutes are text information, identify multiple target text records corresponding to the target object in the meeting minutes, wherein the meeting minutes include: multiple text records, and the speaker corresponding to each text record, each text record being the text conversion result of the speech of the same speaker, and the target text record being the text record corresponding to the target object; segment the multiple target text records according to a pre-set word count requirement to obtain multiple speech fragments; and generate a target meeting speech based on the multiple speech fragments of the target object.

[0115] Optionally, the processor may also execute program code that performs the following steps: obtaining pre-configured question samples, wherein the question samples include at least: fixed prompts and variable prompts, wherein the fixed prompts are used to provide pre-defined guidance to the semantic analysis model, and the variable prompts are used to provide variable guidance to the semantic analysis model according to the meeting requirements; configuring the variable prompts in the question samples according to the pre-obtained meeting agenda to obtain at least one question template, wherein the meeting agenda includes at least: the identity information of the target object and the meeting agenda, and the variable prompts are determined at least based on the identity information or the meeting agenda.

[0116] Optionally, the processor may also execute program code that performs the following steps: based on the identity information in the meeting agenda, adjust the variable prompts in the question template to the preset prompts corresponding to the identity information, thereby obtaining the question template corresponding to the target object, wherein the question template is pre-configured with preset prompts corresponding to the identity information of multiple speaking objects.

[0117] Optionally, the processor may also execute program code for the following steps: determining multiple meeting stages based on the meeting agenda in the meeting agenda table, wherein the meeting agenda includes at least: multiple pre-set meeting stages; determining the target meeting stage matching each speaking segment, wherein the target meeting stage is determined at least based on the speaking time of the speaking segment proposed by the target object; querying the target prompt word adjustment strategy corresponding to each target meeting stage from among multiple pre-set preset prompt word adjustment strategies, wherein the preset prompt word adjustment strategy is pre-set according to the meeting task requirements of each meeting stage; adjusting the variable prompt words in the question sample according to the target prompt word adjustment strategy corresponding to each target meeting stage to obtain the question template corresponding to each speaking segment.

[0118] Those skilled in the art will understand that Figure 3 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a mobile internet device (MID), a PAD, and other terminal devices. Figure 3This does not limit the structure of the aforementioned electronic device. For example, computer terminal 30 may also include components that are more advanced than those described above. Figure 3 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 3 The different configurations shown.

[0119] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a computer program instructing the hardware related to the terminal device. The computer program can be stored in a non-volatile medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0120] Embodiments of the present invention also provide a non-volatile storage medium. Optionally, in this embodiment, the aforementioned non-volatile storage medium can be used to store the program code executed by the AI-based meeting minutes generation method provided in the above embodiments.

[0121] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0122] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: acquiring meeting records, wherein the meeting records are used to store meeting speeches of at least multiple speakers; identifying the target meeting speech of the target object among the meeting speeches of the multiple speakers stored in the meeting records, wherein the target meeting speech includes multiple speech fragments; using a pre-configured question template to drive a pre-set semantic analysis model to analyze the speech fragments and obtain a speech summary for each speech fragment, wherein the question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech fragments, and the target information is used to generate a speech summary; concatenating the speech summaries corresponding to multiple speech fragments to generate the meeting minutes of the target object.

[0123] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: when the meeting record is voice information, acquiring pre-collected voiceprint information of the target object, wherein the meeting record includes: multiple voice records, and the voiceprint information is used to describe the unique voice features of each speaker; based on the voiceprint information of the target object, identifying at least one target voice record in the meeting record, wherein the voice features of the target voice record match the voice features of the target object; converting at least one target voice record into text form to obtain the target meeting speech.

[0124] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: identifying the timestamp of each target voice recording during the conference; classifying two target voice recordings whose time interval between timestamps is less than a preset interval threshold into the same voice set; converting the target voice recordings in multiple voice sets into text form to obtain the target conference speech, wherein multiple target voice recordings in the same voice set are converted into the same speech segment in the target conference speech.

[0125] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: when the meeting record is text information, identify multiple target text records corresponding to the target object in the meeting record, wherein the meeting record includes: multiple text records, and a speaker corresponding to each text record, each text record being the text conversion result of the speech of the same speaker, and the target text record being the text record corresponding to the target object; segment the multiple target text records according to a preset word count requirement to obtain multiple speech fragments; and generate a target meeting speech based on the multiple speech fragments of the target object.

[0126] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a pre-configured question sample, wherein the question sample includes at least: fixed prompts and variable prompts, wherein the fixed prompts are used to provide preset guidance to the semantic analysis model, and the variable prompts are used to provide variable guidance to the semantic analysis model according to the meeting requirements; configuring the variable prompts in the question sample according to a pre-acquired meeting agenda to obtain at least one question template, wherein the meeting agenda includes at least: the identity information of the target object and the meeting agenda, and the variable prompts are determined at least based on the identity information or the meeting agenda.

[0127] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: adjusting the variable prompts in the question template to preset prompts corresponding to the identity information based on the identity information in the meeting agenda, thereby obtaining the question template corresponding to the target object, wherein the question template is pre-configured with preset prompts corresponding to the identity information of multiple speaking objects respectively.

[0128] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: determining multiple meeting stages based on the meeting agenda in the meeting agenda table, wherein the meeting agenda includes at least: multiple pre-set meeting stages; determining the target meeting stage matching each speaking segment, wherein the target meeting stage is determined at least based on the speaking time of the speaking segment proposed by the target object; querying the target prompt word adjustment strategy corresponding to each target meeting stage from among multiple pre-set preset prompt word adjustment strategies, wherein the preset prompt word adjustment strategy is pre-set according to the meeting task requirements of each meeting stage; adjusting the variable prompt words in the question sample according to the target prompt word adjustment strategy corresponding to each target meeting stage to obtain the question template corresponding to each speaking segment.

[0129] Embodiments of the present invention also provide a computer program product, including a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the steps of the artificial intelligence-based meeting minutes generation method provided in the above embodiments.

[0130] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0131] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0132] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0133] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0134] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0135] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned non-volatile storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0136] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for generating meeting minutes based on artificial intelligence, characterized in that, include: Obtain meeting minutes, wherein the meeting minutes are used to store the meeting speeches of at least multiple speakers; Among the meeting speeches of multiple speaking subjects stored in the meeting records, the target meeting speech of the target subject is identified, wherein the target meeting speech includes multiple speech segments; The pre-configured question template drives a pre-set semantic analysis model to analyze the speech segments and obtain a speech summary for each speech segment. The question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech segments, and the target information is used to generate the speech summary. By piecing together the summaries of the speeches corresponding to the multiple speech fragments, a meeting summary of the target object is generated.

2. The method according to claim 1, characterized in that, Among the meeting speeches of the multiple speakers stored in the meeting minutes, the target meeting speech for identifying the target speaker includes: When the meeting record is audio information, the voiceprint information of the target object is obtained in advance, wherein the meeting record includes: multiple audio records, and the voiceprint information is used to describe the unique voice features of each speaker; Based on the voiceprint information of the target object, at least one target voice record in the meeting minutes is identified, wherein the voice features of the target voice record match the voice features of the target object; At least one of the target audio recordings is converted into text form to obtain the target conference speech.

3. The method according to claim 2, characterized in that, Converting at least one of the target speech records into text form to obtain the target conference speech includes: Identify the timestamp of each of the target audio recordings during the meeting; Two target speech records whose time interval between the timestamps is less than a preset interval threshold are grouped into the same speech set; The target speech records in multiple speech sets are converted into text form to obtain the target conference speech, wherein multiple target speech records in the same speech set are converted into the same speech segment in the target conference speech.

4. The method according to claim 1, characterized in that, Among the meeting speeches of the multiple speakers stored in the meeting minutes, the target meeting speech for identifying the target speaker includes: When the meeting record is text information, multiple target text records corresponding to the target object in the meeting record are identified. The meeting record includes: multiple text records and the speaking object corresponding to each text record. Each text record is the text conversion result of the speech of the same speaking object. The target text record is the text record corresponding to the target object. The multiple target text records are segmented according to a pre-set word count requirement to obtain multiple speech fragments; The target conference speech is generated based on multiple speech fragments of the target object.

5. The method according to claim 1, characterized in that, Before analyzing the speech segments using a pre-configured question template-driven pre-set semantic analysis model to obtain a speech summary for each speech segment, the method further includes: Obtain pre-configured question samples, wherein the question samples include at least: fixed prompt words and variable prompt words, wherein the fixed prompt words are used to provide preset guidance to the semantic analysis model, and the variable prompt words are used to provide variable guidance to the semantic analysis model according to the needs of the meeting; Based on a pre-acquired meeting agenda, the variable prompts in the question samples are configured to obtain at least one question template. The meeting agenda includes at least the identity information of the target object and the meeting agenda, and the variable prompts are determined based at least on the identity information or the meeting agenda.

6. The method according to claim 5, characterized in that, Based on the pre-acquired meeting agenda, the variable prompts in the question sample are configured to obtain the question template, which includes: Based on the identity information in the meeting agenda, the variable prompts in the question template are adjusted to preset prompts corresponding to the identity information to obtain the question template corresponding to the target object. The question template is pre-configured with preset prompts corresponding to the identity information of multiple speaking objects.

7. The method according to claim 5, characterized in that, Based on the pre-acquired meeting agenda, the variable prompts in the question sample are configured to obtain the question template, which includes: Based on the meeting agenda in the meeting agenda sheet, multiple meeting stages are determined, wherein the meeting agenda includes at least: multiple pre-set meeting stages; Determine the target meeting stage that matches each of the said speech segments, wherein the target meeting stage is determined at least based on the speaking time when the target object proposed the speech segment; Among a number of preset prompt word adjustment strategies, the target prompt word adjustment strategy corresponding to each target meeting stage is queried, wherein the preset prompt word adjustment strategy is preset according to the meeting task requirements of each meeting stage; The variable prompts in the question sample are adjusted according to the target prompts adjustment strategy corresponding to each target meeting stage to obtain the question template corresponding to each speech segment.

8. A meeting minutes generation device based on artificial intelligence, characterized in that, include: An acquisition module is used to acquire meeting minutes, wherein the meeting minutes are used to store the meeting speeches of at least multiple speakers; The identification module is used to identify a target meeting speech of a target object among the meeting speeches of multiple speaking objects stored in the meeting records, wherein the target meeting speech includes multiple speech segments; An analysis module is used to drive a pre-set semantic analysis model to analyze the speech segments using a pre-configured question template, and obtain a speech summary for each speech segment. The question template is used to provide prompt words for the semantic analysis model, the prompt words are used to guide the semantic analysis model to extract target information from the speech segments, and the target information is used to generate the speech summary. The splicing module is used to splice together the speech summaries corresponding to multiple speech fragments to generate the meeting minutes of the target object.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the AI-based meeting minutes generation method according to any one of claims 1 to 7 through the computer program.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the AI-based meeting minutes generation method according to any one of claims 1 to 7.