Intelligent conference recording method and apparatus, intelligent device and storage medium
By capturing real-time video frame sequences and audio from meetings, identifying speaking trigger gestures and associating them with speaking content, the efficiency and accuracy issues of traditional meeting recording are resolved, enabling efficient and accurate recording and sharing of meeting information.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SHENZHEN HONGHE INNOVATION INFORMATION TECH CO LTD
- Filing Date
- 2025-06-04
- Publication Date
- 2026-07-02
AI Technical Summary
Traditional meeting minutes are labor-intensive, inaccurate, and difficult to retrieve and share quickly, especially in large or multilingual meetings where the accuracy and completeness of the minutes cannot be guaranteed.
By acquiring real-time video frame sequences and audio from the meeting, the system identifies the target speaker's gestures, determines the speaker and their identity information, and associates the audio content with the identity information to generate meeting minutes.
It improves the accuracy and completeness of meeting minutes, saves human resources, shortens recording time, and facilitates subsequent information retrieval and sharing.
Smart Images

Figure CN2025098991_02072026_PF_FP_ABST
Abstract
Description
Intelligent meeting recording methods, devices, smart devices, and storage media
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411981439.7, filed on December 27, 2024, entitled “Smart Meeting Recording Method, Apparatus, Smart Device and Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of smart device technology, and in particular to a smart meeting recording method, apparatus, smart device, and storage medium. Background Technology
[0004] In today's era of globalization and rapid information development, meetings have become an indispensable means of communication, decision-making, and collaboration across various fields. Meeting minutes, as a written representation of the meeting process and outcomes, are of paramount importance. They are a crucial carrier of information, comprehensively recording key discussion points, decision-making processes, task assignments, and the expression of viewpoints by all parties. High-quality meeting minutes help ensure the effective implementation of meeting outcomes, avoid errors or duplication of effort due to information transmission discrepancies or omissions, improve overall work efficiency and collaboration, and promote communication and collaborative development within and between organizations.
[0005] Traditional meeting minutes typically require a dedicated recorder to document the meeting in real time, including detailed notes on each speaker's identity and remarks. This manual recording method not only consumes additional human resources, but the recorder's speed and focus can also lead to omissions or errors. Even in meetings where minutes are reviewed afterward, extra time is needed to organize and extract key points, significantly increasing meeting-related time costs and reducing overall efficiency. This is especially true in large, lengthy meetings or those involving multilingual communication, where the difficulty and workload of manual recording increase dramatically, and accuracy and completeness are difficult to guarantee. Furthermore, traditional meeting minutes hinder rapid information retrieval and sharing. Reviewing a specific speaker's viewpoint or a particular topic often requires manual searching through lengthy documents, a tedious and time-consuming process.
[0006] Therefore, how to accurately and efficiently complete meeting minutes and facilitate the retrieval and sharing of meeting information is a technical problem that urgently needs to be solved. Summary of the Invention
[0007] This application provides an intelligent meeting recording method, apparatus, intelligent device, and storage medium, which can accurately and efficiently complete meeting recording and facilitate the retrieval and sharing of meeting information.
[0008] In a first aspect, embodiments of this application provide an intelligent meeting recording method, including:
[0009] Real-time acquisition of video frame sequences and audio from the conference venue;
[0010] Identify whether a target speech triggers a gesture in a video frame sequence image;
[0011] Based on the identified target speaking trigger gesture and the meeting participant information record, determine the current speaker and their identity information;
[0012] The speech content corresponding to the on-site audio collected from the moment the target speech is detected and the speaker's identity information will be associated with it.
[0013] Meeting minutes are generated based on the speeches of all speakers identified at the meeting and their associated identity information.
[0014] In one possible implementation of the first aspect, identifying whether a target speech-triggered gesture exists in a sequence of video frame images includes:
[0015] Individual participants in the video frame sequence images are distinguished and located to obtain personnel region sequence images of each participant in the meeting venue;
[0016] Input the personnel region sequence image into the bone joint recognition model to obtain the coordinates of key bone joints of the participants;
[0017] Based on the coordinates of key bone joints corresponding to the personnel area sequence images, determine the motion trajectory of the key bone joints;
[0018] Determine whether the posture corresponding to the movement trajectory of key bones and joints is the posture that triggered the target's speech;
[0019] If so, then it is determined that there is a target speaking trigger posture in the video frame sequence image.
[0020] In one possible implementation of the first aspect, the current speaker and their identity information are determined based on the identified target speaking trigger gesture and the meeting participant information record, including:
[0021] The participant who presents the target speaking trigger gesture will be identified as the current speaker;
[0022] The speaker's target facial image is matched with the facial images of attendees in the meeting attendee information record, and the speaker's identity information is determined based on the identity information of the matched attendees.
[0023] In one possible implementation of the first aspect, the participant who presents the target speaking trigger gesture is identified as the current speaker, including:
[0024] Obtain candidate face images of attendees who will present the target speaking gesture;
[0025] Centered on the video frame image corresponding to the candidate face image, a preset number of video frame images are selected forward and backward from the acquired video frame sequence image as the associated image sequence of the candidate face image;
[0026] Based on associated image sequences, determine the variation features of the mouth contour in candidate face images;
[0027] Based on the changing features of the mouth contour, determine whether the attendee corresponding to the candidate's face image is speaking;
[0028] If someone is speaking, the participant who is in the target speaking trigger pose will be identified as the current speaker.
[0029] In one possible implementation of the first aspect, the current speaker and their identity information are determined based on the identified target speaking trigger gesture and the meeting participant information record, including:
[0030] If multiple participants are identified to simultaneously exhibit the target speaking trigger posture, the direction of the sound source is determined based on the current on-site audio.
[0031] Determine the location of the participants corresponding to the speaking trigger gestures of each target;
[0032] By comparing the location with the direction of the sound source, the participants whose location and the direction of the sound source are consistent within the preset location error range are identified as the current speakers;
[0033] The speaker's target facial image is matched with the facial images of attendees in the meeting attendee information record, and the speaker's identity information is determined based on the identity information of the matched attendees.
[0034] In one possible implementation of the first aspect, before real-time acquisition of video frame sequences and audio from the conference venue, the method further includes:
[0035] Obtain facial images and identity information of attendees;
[0036] Based on facial images and their identity information, a record of information about the participants in this meeting is generated.
[0037] In one possible implementation of the first aspect, obtaining the facial images and identity information of the attendees includes:
[0038] At the meeting, facial images and identity information of attendees are recorded sequentially; or,
[0039] Obtain the list of attendees for this meeting;
[0040] Based on the list of attendees, retrieve the facial images and identity information of the attendees from the meeting management system.
[0041] Secondly, embodiments of this application provide an intelligent meeting recording device, comprising:
[0042] The information acquisition unit is used to acquire video frame sequences and audio from the conference venue in real time.
[0043] The action recognition unit is used to identify whether there is a target speaking trigger posture in the video frame sequence image;
[0044] The speaker identification unit is used to determine the current speaker and their identity information based on the identified target speaking trigger gesture and the information record of meeting participants.
[0045] The content association unit is used to associate the speech content corresponding to the on-site audio collected from the moment the target speech is detected with the speaker and their identity information.
[0046] The record generation unit is used to generate meeting minutes based on the speeches of all speakers identified at the meeting and their associated identity information.
[0047] Thirdly, embodiments of this application provide an intelligent device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent meeting recording method as described in the first aspect above.
[0048] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the intelligent meeting recording method as described in the first aspect above.
[0049] Fifthly, embodiments of this application provide a computer program product that, when run on a smart device, causes the smart device to execute the smart meeting recording method as described in the first aspect above.
[0050] In this embodiment, by real-time acquisition of video frame sequences and audio from the meeting venue, and based on the target speaking trigger posture identified in the video frame sequence and the meeting participant information, the current speaker and their identity information are accurately determined. The audio content acquired from the moment the target speaking trigger posture is identified is then associated with the speaker and their identity information, greatly improving the accuracy of the correspondence between speakers and their speech content in the meeting minutes. Based on the speech content associated with all speakers and their identities identified at the meeting, meeting minutes are intelligently and quickly generated. This solution abandons the traditional method of manual recording by hand, saving human resources, significantly reducing the time cost of meeting minutes, and avoiding information omissions and errors caused by personal factors in manual recording. It effectively ensures the accuracy and completeness of meeting minutes. Furthermore, the precise association between speakers, their identity information, and their speech content allows for convenient retrieval of specific speakers' viewpoints or discussion topics when reviewing or sharing meeting information later, facilitating the retrieval and sharing of meeting information. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 is a flowchart illustrating the implementation of the intelligent meeting recording method provided in an embodiment of this application;
[0053] Figure 2 is a flowchart of a specific implementation of step S102 in the intelligent meeting recording method provided in the embodiment of this application;
[0054] Figure 3 is a schematic diagram of key human skeletal joints in the intelligent meeting recording method provided in the embodiments of this application;
[0055] Figure 4 is a flowchart of a specific implementation of the intelligent meeting recording method provided in this application for determining the current speaker;
[0056] Figure 5 is a flowchart of a specific implementation of step S103 in the intelligent meeting recording method provided in the embodiment of this application;
[0057] Figure 6 is a structural block diagram of the intelligent meeting recording device provided in an embodiment of this application;
[0058] Figure 7 is a schematic diagram of the smart device provided in an embodiment of this application.
[0059] The following is a detailed list of the reference numerals used in the above figures: 61. Information acquisition unit; 62. Action recognition unit; 63. Speaker determination unit; 64. Content association unit; 65. Record generation unit; 7. Smart device; 70. Processor; 71. Memory; 72. Computer program. Detailed Implementation
[0060] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0061] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0062] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0063] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0064] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0065] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0066] The intelligent meeting recording method provided in this application is applicable to intelligent devices that need to record various types of data during meetings. Specifically, intelligent devices may include mobile phones, tablets, wearable devices, laptops, ultra-mobile personal computers (UMPCs), desktop computers, and servers. This application does not impose any limitations on the specific type of intelligent device.
[0067] Figure 1 illustrates the implementation flow of the intelligent meeting recording method provided in this application embodiment. The method flow includes steps S101 to S105. The specific implementation principle of each step is as follows:
[0068] Step S101: Real-time acquisition of video frame sequence images and on-site audio from the conference venue.
[0069] In this embodiment, image acquisition equipment and professional audio acquisition equipment are used to simultaneously capture visual and auditory information of the meeting venue, ensuring a comprehensive record of the meeting scene and providing a rich data foundation for subsequent analysis and generation of meeting minutes. The image acquisition equipment must have high resolution and low-light performance to adapt to meeting environments under different lighting conditions, and the audio acquisition equipment must be able to accurately capture clear, noise-free on-site sound, with a sampling frequency meeting industry standard requirements.
[0070] A video frame sequence is composed of a series of continuous, static video frames arranged in a specific order. In a meeting setting, image acquisition equipment (usually a high-definition camera) continuously captures photos at preset time intervals (e.g., 30 times per second), and each captured image is a video frame. These video frames are closely linked, with only extremely subtle time differences and image changes between each frame, reflecting the states of participants at different moments. These video frames are arranged sequentially according to the order in which they were captured to form a video frame sequence. For example, high-definition cameras are strategically installed at key locations in the meeting room, such as at the four corners, to ensure coverage of the entire meeting area without blind spots. In some possible implementations, additional cameras can be placed at both ends and the center of the conference table to capture facial expressions of participants.
[0071] The video frame sequence is dynamically updated. Until the meeting concludes, the image acquisition equipment continuously captures footage of the meeting venue. As time progresses, newly captured frames are added to the sequence, forming a continuous stream of images that reflects the real-time changes in the actions, expressions, and postures of the participants. Each new moment is recorded as a new video frame and added to the existing sequence, ensuring that the entire video frame sequence remains synchronized with the real-time situation at the meeting.
[0072] On-site audio can be captured using professional microphone arrays. Microphone arrays typically consist of multiple microphone units arranged in a specific geometry, commonly circular or linear layouts. For example, a microphone array can be installed in the center of the conference room ceiling or above the center of the conference table to receive sound from all directions. Microphone arrays utilize beamforming technology to precisely focus on the sound source based on the direction of sound propagation and the time and phase differences between the microphones, effectively reducing ambient noise interference and ensuring clear and pure audio capture.
[0073] In this embodiment, the smart device is communicatively connected to the image acquisition device and the audio acquisition device. The smart device acquires video frame sequences captured by the image acquisition device and on-site audio captured by the audio acquisition device in real time. The video frame sequences captured by the image acquisition device and the on-site audio captured by the audio acquisition device need to be transmitted in real time through a high-speed and stable transmission channel.
[0074] In some possible implementations, the smart device is connected to the image acquisition device and the audio acquisition device via a wired connection. The wired connection method provides stable and reliable data transmission, and the bandwidth can meet the real-time transmission requirements of high-definition video and audio, effectively avoiding problems such as data packet loss and latency.
[0075] In some possible implementations, considering the ease of wiring in the conference room, the smart device can be wirelessly connected to the image acquisition device and the audio acquisition device to ensure that, even when multiple acquisition devices are transmitting data simultaneously in a complex conference environment, the transmission characteristics of low latency and high bandwidth can still be maintained. The image acquisition device and the audio acquisition device will send the acquired video frame sequence images and on-site audio to the smart device for processing in real time and without loss.
[0076] As one possible implementation of this application, before the meeting begins, facial images and identity information of the participants are obtained, and a meeting participant information record is generated based on the facial images and identity information.
[0077] In some possible implementations, before the meeting begins, the facial images and identity information of the attendees are sequentially recorded at the meeting venue. In this embodiment, a human-computer interaction interface guides attendees to sequentially record their facial images and identity information at the meeting venue. A high-definition camera is used for facial image recording to ensure that the facial image resolution is not lower than a preset resolution (e.g., 1080p), and the recording process follows facial image acquisition standards. Before recording begins, the lighting conditions for facial recording are adjusted according to the ambient light at the meeting venue, and the facial pose is standardized and adjusted through the human-computer interaction interface during the recording process.
[0078] After attendees enter the meeting venue, their facial images and identity information are collected. The images obtained represent the attendees' most authentic state at that time, avoiding discrepancies caused by time differences or temporary personnel changes, and ensuring the completeness and accuracy of the entered information.
[0079] In some possible implementations, before the meeting begins, a list of meeting participants is obtained; based on the participant list, facial images and their identity information are retrieved from the meeting management system. In this embodiment, before the meeting begins, a list of meeting participants is obtained from the official meeting management platform provided by the meeting organizer; based on the participant list, facial images and their identity information are retrieved from the meeting management system, following the standard data interface protocol opened by the meeting management system to ensure that the retrieved data format is compatible with the local processing system.
[0080] For large conferences or conferences with fixed attendees who have been registered in advance in the conference management system, directly retrieving the attendees' facial images and their identity information is more efficient and can save manpower and time costs for on-site data entry.
[0081] Step S102: Identify whether there is a target speaking trigger posture in the video frame sequence image.
[0082] The target speaking trigger posture is a preset posture that triggers speaking. This posture can be set according to needs. For example, raising your hand can be set as the target speaking trigger posture, or nodding can be set as the target speaking trigger posture.
[0083] In this embodiment, each video frame in the video frame sequence carries important information, capturing the posture, gestures, and facial expressions of participants at a specific moment. By acquiring video frame sequences from the meeting, the start, process, and end of actions such as raising hands and nodding can be effectively tracked, thereby identifying whether a target speaking trigger posture exists. This provides basic data for subsequent operations such as accurately identifying speakers and recording speaking content, contributing to the accurate generation of meeting minutes.
[0084] For example, in a meeting, a video frame freezes on participant A, who is leaning slightly forward, attentively listening to others. The next frame might capture participant B gently raising their right hand with fingers extended, clearly showing their intention to speak. By analyzing this series of video images, the start, process, and end of participants' hand-raising, nodding, and other actions can be fully tracked, thereby identifying whether a target speaking-triggered posture exists.
[0085] As one possible implementation of this application, Figure 2 illustrates a specific implementation flow of step S102 in the intelligent meeting recording method provided in the embodiment of this application, which is described in detail below:
[0086] A1: Individually distinguish and locate the participants in the video frame sequence images to obtain the personnel area sequence images of each participant in the meeting venue.
[0087] In this embodiment, a target detection algorithm is used to quickly and accurately identify each participant in the video image. Precise positioning is achieved through bounding box selection, resulting in personnel region images of each participant in the meeting venue. This process is repeated for each frame of the video frame sequence, thereby obtaining a series of personnel region images for each participant. Based on this series of personnel region images, a personnel region sequence image for each participant is constructed.
[0088] A2: Input the personnel region sequence image into the bone joint recognition model to obtain the coordinates of the key bone joints of the participants.
[0089] The bone and joint recognition model is a pre-trained neural network model used to identify the coordinates of human bone and joints. The training phase is crucial for the model to achieve accurate recognition capabilities. To adapt the joint recognition model to the bone and joint recognition needs of various meeting scenarios, a massive amount of personnel image data covering different dimensions of features was collected, including gender, age group, body type, and posture. Generally, there are differences between men and women in skeletal structure and limb proportions; children, adults, and the elderly have different ranges of motion and appearance of their bone and joints; and there are differences between thin and obese individuals in limb thickness and joint prominence. Postures include common meeting gestures, such as raising hands, nodding, and turning sideways to communicate, comprehensively covering human movements in meetings. The personnel image data used for training annotates key bone and joints, as shown in Figure 3. After iterative training with a large amount of personnel image data over a long period, the model learns the appearance features and positional patterns of bone and joints in different postures. When a sequence of personnel region images is input into the trained bone and joint recognition model, the model automatically extracts image features and outputs the coordinates of key bone and joints in each personnel region image in the sequence of personnel region images. The coordinates of key bones and joints can reflect the position of the participants' bones and joints in each frame of the image.
[0090] A3: Determine the motion trajectory of key bones and joints based on the coordinates of the key bones and joints corresponding to the personnel area sequence images;
[0091] A personnel region image sequence comprises consecutive frames of personnel region images. Each image sequence corresponds to the coordinates of key joints at consecutive time points. By fitting the coordinates of these key joints at these consecutive time points, the motion trajectory of the joint coordinates can be determined. Fitting methods can employ mathematical methods such as linear interpolation. For example, for the action of raising a hand, by tracking the coordinate changes of the shoulder, elbow, and wrist joints, an interpolation algorithm can be used to connect discrete coordinate points into a smooth curve, visually representing the movement path of the joints from their initial to final positions, thus forming the motion trajectory of the joints.
[0092] A4: Determine whether the posture corresponding to the movement trajectory of the key bone joint is the posture triggered by the target's speech.
[0093] A5: If so, then it is determined that there is a target speaking trigger posture in the video frame sequence image.
[0094] In this embodiment, the key bone joint motion trajectory features of the target speech trigger posture are predefined. The motion trajectory of the identified key bone joints is compared with the key bone joint motion trajectory features of the target speech trigger posture. Based on the comparison result, it is determined whether the posture corresponding to the motion trajectory of the key bone joint is the target speech trigger posture. Generally, if the similarity between the identified key bone joint motion trajectory and the key bone joint motion trajectory features of the target speech trigger posture reaches a preset similarity threshold, it is determined that the identified key bone joint motion trajectory and the key bone joint motion trajectory features of the target speech trigger posture are the same, and the posture corresponding to the motion trajectory of the key bone joint is the target speech trigger posture. Conversely, if the similarity between the identified key bone joint motion trajectory and the key bone joint motion trajectory features of the target speech trigger posture does not reach the preset similarity threshold, it is determined that the identified key bone joint motion trajectory and the key bone joint motion trajectory features of the target speech trigger posture are different, and the posture corresponding to the motion trajectory of the key bone joint is not the target speech trigger posture.
[0095] For example, if the target speaking trigger posture is raising a hand, and the similarity between the motion trajectory of a participant's key joints in the identified video frame sequence and the predefined motion trajectory features of raising a hand reaches a preset similarity threshold, then the posture corresponding to the identified key joint motion trajectory is determined to be raising a hand, and the target speaking trigger posture exists in the video frame sequence. If the similarity between the motion trajectory of a participant's key joints in the identified video frame sequence and the predefined motion trajectory features of raising a hand does not reach the preset similarity threshold, then the posture corresponding to the identified key joint motion trajectory is determined not to be raising a hand, and the target speaking trigger posture does not exist in the video frame sequence.
[0096] In this embodiment of the application, if the target speaking trigger posture is not identified in the video frame sequence image, then after the video frame sequence image is updated, the above steps A1 to A5 are repeated for the newly added video image.
[0097] Step S103: Based on the identified target speaking trigger gesture and the meeting participant information record, determine the current speaker and their identity information.
[0098] The spokesperson's identity information includes their name and organization, and may also include their position. The specific content of the identity information can be customized according to needs.
[0099] In this embodiment, the participant exhibiting the target speaking trigger posture is identified as the current speaker. The speaker's target facial image is matched with the facial images of participants in the meeting participant information record. The speaker's identity information is determined based on the identity information of the matched participants. Specifically, the speaker's facial region is located in the video image where the target speaking trigger posture is identified, and the speaker is identified. The facial image of this region is extracted as the target facial image and matched with the facial images of participants in the meeting participant information record to find the speaker's identity information.
[0100] As one possible implementation of this application, Figure 4 illustrates a specific implementation flow of determining the current speaker in the intelligent meeting recording method provided in the embodiments of this application, which is described in detail below:
[0101] B1: Obtain the candidate face image of the participant who presented the target speaking trigger pose.
[0102] B2: Taking the video frame image corresponding to the candidate face image as the center, select a preset number of video frame images forward and backward from the acquired video frame sequence image as the associated image sequence of the candidate face image.
[0103] B3: Based on the associated image sequence, determine the variation features of the mouth contour in the candidate face image.
[0104] B4: Based on the changing features of the mouth contour, determine whether the attendee corresponding to the candidate's face image is speaking.
[0105] B5: If speaking, the participant who presents the target speaking trigger posture will be identified as the current speaker.
[0106] Considering that attendees may exhibit a target speaking-triggered posture but not necessarily intend to speak, to avoid misjudgment, in this embodiment, upon identifying the target speaking-triggered posture, the attendee exhibiting that posture is immediately identified, and their facial image is captured as a candidate facial image. The video frame containing the candidate facial image is used as the reference video frame. A preset number of video frames are selected forward and backward from the video frame sequence as the associated image sequence for the candidate facial image. The preset number of frames selected forward and backward can be estimated based on common video frame rates and normal speaking speed. For example, selecting 8 frames forward and backward roughly covers 1-2 seconds, capturing the changes in the mouth's contour during the start, process, and short pauses of speech. Based on these changes, it can be determined whether the attendee corresponding to the candidate facial image is speaking. For example, changes in the mouth's contour include changes in the area enclosed by the mouth and changes in the distance between the lips. By setting a reasonable threshold for change, when the changes in the mouth contour characteristics in several consecutive video frames match the frequency and amplitude range of mouth opening and closing during normal speech, it can be determined that the participant in the candidate's face image is speaking. Through correlation analysis of mouth contour changes across multiple video frames, the presence of speaking behavior can be determined, avoiding misidentification of the speaker and greatly improving the accuracy of speaker identification. This ensures that meeting records are accurately linked to the actual speaker.
[0107] As one possible implementation of this application, Figure 5 illustrates a specific implementation flow of step S103 of the intelligent meeting recording method provided in the embodiment of this application, which is described in detail below:
[0108] C1: If multiple participants are detected simultaneously exhibiting the target speaking trigger posture, the direction of the sound source is determined based on the current ambient audio. The determination of the sound source direction can be achieved using microphone array technology to process the ambient audio, which will not be elaborated upon here.
[0109] C2: Determine the location of the participants corresponding to the speaking trigger posture of each target.
[0110] In this embodiment, after identifying a target speaking trigger posture, an image-based spatial coordinate positioning algorithm can be used to determine the location of the participant corresponding to each target speaking trigger posture. The location is determined by analyzing the pixel positions of the participants exhibiting the target speaking trigger posture in the image.
[0111] C3: Compare the location with the direction of the sound source, and determine the participants whose location and the direction of the sound source are consistent within the preset location error range as the current speaker.
[0112] C4: Match the target facial image of the speaker with the facial images of the attendees in the meeting attendee information record, and determine the speaker's identity information based on the identity information of the matched attendees.
[0113] Considering the situation where multiple participants simultaneously present the target speaking trigger posture, by judging the complex situation of multiple people competing to speak at the same time, the true speaker is accurately screened by combining sound localization and visual orientation recognition, avoiding confusion and misjudgment, greatly improving the accuracy of speaker identification, and ensuring that the meeting minutes accurately point to the actual speaker.
[0114] Step S104: Associate the speech content corresponding to the on-site audio collected from the moment the target speech is detected with the speaker and their identity information.
[0115] In this embodiment, the collected on-site audio is input into a speech recognition model in real time. The speech recognition model can be a pre-trained deep learning-based network model. This model quickly parses the speech signal, converting sound into text information, and thus obtaining the corresponding speech content from the on-site audio. The speech content obtained from the moment the target speech trigger gesture is recognized is associated with the identified speaker and their identity information, ensuring that each speech segment is clearly assigned to its corresponding speaker, avoiding information confusion. This precise correspondence between speech content and speaker and identity information allows for rapid and accurate acquisition of the required information, whether for generating meeting minutes, reviewing meeting highlights, or querying the views of a specific speaker, greatly improving the effectiveness of meeting minutes.
[0116] In some possible implementations, the collected on-site audio is input into a speech recognition model in real time to obtain the corresponding text information. This text information is then cleaned using preset text cleaning rules, and semantic recognition is performed on the cleaned text. After error correction based on the context, the spoken content is obtained.
[0117] On the one hand, due to the complex environment at the meeting site, some error characters, repeated words, or meaningless pause symbols caused by noise may be introduced during the speech recognition process. Use the preset text cleaning rules to remove this interference information. For example, use regular expressions to match and delete the same characters that appear multiple times consecutively, such as changing "我我我" to "我"; identify and remove words like "嗯" and "啊" that have no actual semantic meaning and are only caused by speaking pause habits. If these words appear no more than the preset number of times (e.g., 3 times), they are directly deleted to streamline the text content. On the other hand, the speech recognition model may occasionally have the situation of confusing homophonic words, such as misidentifying "形式" as "形势". When a suspected incorrect word is found, analyze the context keywords and sentence structure, and replace the word that does not conform to the semantic logic with the correct word. At the same time, supplement some missing words caused by too fast speech speed or unclear pronunciation to make the text表意完整连贯.
[0118] In some possible implementation manners, in combination with the industry field associated with this meeting, standardize the professional terms and abbreviations in the speech content. Meetings in different industry fields involve a large number of professional terms and specific abbreviations. To ensure the accuracy and generality of the speech content, establish a professional term library and an abbreviation explanation table corresponding to the industry field. When a special word appears in the speech content, replace or expand the explanation of the special word according to the professional term library and abbreviation explanation table corresponding to the industry field associated with this meeting. For example, in a medical meeting, uniformly replace "CT" with "计算机断层扫描"; in an IT industry meeting, explain "AI" in detail as "人工智能" so that those who are not familiar with this field can also understand accurately.
[0119] In some possible implementation manners, organize the speech content in a structured way. Exemplarily, it can be organized according to the structure of "Speaker (identity information) - Speech time - Speech content". For example, "张三(技术总监)-10:15:20-关于新产品研发,我们需要重点关注以下几个方面……", making the speech content well-organized, facilitating the subsequent generation of a standardized meeting record, and ensuring the accurate restoration of the全貌 of the meeting speech.
[0120] Step S105: Generate a meeting record based on the speech content associated with all the speakers determined at the meeting site and their identity information.
[0121] When the meeting ends, the intelligent device comprehensively sorts and integrates the massive information accumulated during the meeting. Traverse the speech content associated with all the speakers determined in the meeting and their identity information to ensure that nothing is missed.
[0122] In this embodiment, meeting minutes can be generated based on a preset standardized template. This template follows industry-standard practices as well as the personalized needs of the meeting organizers, listing each speaker and their identification information (name, organization, and position), speaking time, and content in the order they speak. Each segment of the speech is neatly formatted and clearly organized, facilitating reading and subsequent analysis.
[0123] As can be seen from the above, in this embodiment, by real-time acquisition of video frame sequences and audio from the meeting venue, and based on the target speaking trigger posture identified in the video frame sequence and the meeting participant information, the current speaker and their identity information are accurately determined. Furthermore, the speech content corresponding to the audio collected from the moment the target speaking trigger posture is identified is associated with the speaker and their identity information, greatly improving the accuracy of the correspondence between speakers and speech content in the meeting minutes. Based on the speech content associated with all speakers and their identity information determined at the meeting venue, meeting minutes are intelligently and quickly generated. This solution abandons the traditional method of having a dedicated person handwrite the minutes in real time, saving human resources, significantly reducing the time cost of meeting minutes, and avoiding information omissions and errors caused by personal factors in manual recording. It effectively ensures the accuracy and completeness of meeting minutes information. Simultaneously, based on the precise association between the speaker and their identity information and the speech content, it allows for convenient retrieval of specific speakers' viewpoints or topic discussion content when reviewing or sharing meeting information later, facilitating the retrieval and sharing of meeting information.
[0124] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0125] Corresponding to the intelligent meeting recording method in the above embodiments, Figure 6 shows a structural block diagram of the intelligent meeting recording device provided in the embodiments of this application. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0126] Referring to Figure 6, the intelligent meeting recording device includes: an information acquisition unit 61, an action recognition unit 62, a speaker identification unit 63, a content association unit 64, and a record generation unit 65, wherein:
[0127] Information acquisition unit 61 is used to acquire video frame sequence images and on-site audio in real time at the conference venue;
[0128] Action recognition unit 62 is used to identify whether there is a target speaking trigger posture in the video frame sequence image;
[0129] Speaker identification unit 63 is used to identify the current speaker and their identity information based on the identified target speaking trigger posture and the information record of meeting participants.
[0130] The content association unit 64 is used to associate the speech content corresponding to the on-site audio collected from the moment the target speech is triggered with the speaker and their identity information.
[0131] The record generation unit 65 is used to generate meeting minutes based on the speeches of all speakers and their associated identity information identified at the meeting.
[0132] As one possible implementation of this application, the action recognition unit 62 includes:
[0133] The joint coordinate acquisition module is used to input personnel region sequence images into the bone and joint recognition model to obtain the coordinates of key bone and joints of the participants;
[0134] The trajectory determination module is used to determine the motion trajectory of key bones and joints based on the coordinates of the key bones and joints corresponding to the personnel area sequence image.
[0135] The posture determination module is used to determine whether the posture corresponding to the motion trajectory of the key bone joint is the target speech trigger posture; if so, it determines that the target speech trigger posture exists in the video frame sequence image.
[0136] As one possible implementation of this application, the speaker determination unit 63 includes:
[0137] The first speaker determination module is used to identify the participant who presents the target speaking trigger gesture as the current speaker.
[0138] The first information matching module is used to match the target facial image of the speaker with the facial images of the attendees in the meeting attendee information record, and determine the speaker's identity information based on the identity information of the matched attendees.
[0139] As one possible implementation of this application, the first speaker determination module is specifically used for:
[0140] Obtain candidate face images of attendees who will present the target speaking gesture;
[0141] Centered on the video frame image corresponding to the candidate face image, a preset number of video frame images are selected forward and backward from the acquired video frame sequence image as the associated image sequence of the candidate face image;
[0142] Based on associated image sequences, determine the variation features of the mouth contour in candidate face images;
[0143] Based on the changing features of the mouth contour, determine whether the attendee corresponding to the candidate's face image is speaking;
[0144] If someone is speaking, the participant who is in the target speaking trigger pose will be identified as the current speaker.
[0145] As one possible implementation of this application, the speaker determination unit 63 includes:
[0146] The sound source direction determination module is used to determine the sound source direction based on the current on-site audio if multiple participants are simultaneously identified as exhibiting the target speaking trigger posture.
[0147] The location determination module is used to determine the location of the participants corresponding to the speaking trigger posture of each target.
[0148] The second speaker determination module is used to compare the location with the direction of the sound source and determine the participants whose location and the direction of the sound source are consistent within the preset location error range as the current speaker.
[0149] The second information matching module is used to match the target facial image of the speaker with the facial images of the attendees in the meeting attendee information record, and determine the speaker's identity information based on the identity information of the matched attendees.
[0150] In one possible implementation of this application, the intelligent meeting recording device further includes:
[0151] The information acquisition unit is used to acquire facial images and identity information of the participants;
[0152] The information recording generation unit is used to generate information records of the meeting participants based on facial images and their identity information.
[0153] As one possible implementation of this application, the information acquisition unit includes:
[0154] The first information acquisition module is used to sequentially record the facial images and identity information of attendees at the meeting venue; or,
[0155] The second information acquisition module is used to obtain the list of attendees for this meeting; based on the list of attendees, it retrieves the facial images and identity information of the attendees from the meeting management system.
[0156] As can be seen from the above, in this embodiment, by real-time acquisition of video frame sequences and audio from the meeting venue, and based on the target speaking trigger posture identified in the video frame sequence and the meeting participant information, the current speaker and their identity information are accurately determined. Furthermore, the speech content corresponding to the audio collected from the moment the target speaking trigger posture is identified is associated with the speaker and their identity information, greatly improving the accuracy of the correspondence between speakers and speech content in the meeting minutes. Based on the speech content associated with all speakers and their identity information determined at the meeting venue, meeting minutes are intelligently and quickly generated. This solution abandons the traditional method of having a dedicated person handwrite the minutes in real time, saving human resources, significantly reducing the time cost of meeting minutes, and avoiding information omissions and errors caused by personal factors in manual recording. It effectively ensures the accuracy and completeness of meeting minutes information. Simultaneously, based on the precise association between the speaker and their identity information and the speech content, it allows for convenient retrieval of specific speakers' viewpoints or topic discussion content when reviewing or sharing meeting information later, facilitating the retrieval and sharing of meeting information.
[0157] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0158] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the intelligent meeting recording methods shown in Figures 1 to 5.
[0159] This application also provides an intelligent device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of any of the intelligent meeting recording methods shown in Figures 1 to 5.
[0160] This application also provides a computer program product that, when run on a smart device, causes the smart device to execute steps to implement any of the smart meeting recording methods shown in Figures 1 to 5.
[0161] Figure 7 is a schematic diagram of a smart device provided in an embodiment of this application. As shown in Figure 7, the smart device 7 of this embodiment includes: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70. When the processor 70 executes the computer program 72, it implements the steps in the various smart conference recording method embodiments described above, such as steps S101 to S105 shown in Figure 1. Alternatively, when the processor 70 executes the computer program 72, it implements the functions of each module / unit in the various device embodiments described above, such as the functions of units 61 to 65 shown in Figure 6.
[0162] For example, computer program 72 may be divided into one or more modules / units, one or more of which are stored in memory 71 and executed by processor 70 to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of computer program 72 in smart device 7.
[0163] The smart device 7 may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will understand that FIG7 is merely an example of the smart device 7 and does not constitute a limitation on the smart device 7. It may include more or fewer components than shown, or combine certain components, or different components. For example, the smart device 7 may also include input / output devices, network access devices, buses, etc.
[0164] The processor 70 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0165] The memory 71 can be an internal storage unit of the smart device 7, such as a hard drive or RAM in the smart device 7. The memory 71 can also be an external storage device of the smart device 7, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the smart device 7. Furthermore, the memory 71 can include both internal and external storage units of the smart device 7. The memory 71 is used to store computer programs and other programs and data required by the smart device. The memory 71 can also be used to temporarily store data that has been output or will be output.
[0166] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0167] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0168] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / smart device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0169] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0170] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A smart meeting recording method, comprising: Real-time acquisition of video frame sequences and audio from the conference venue; Identify whether a target speech trigger gesture exists in the video frame sequence image; Based on the identified target speaking trigger gesture and the meeting participant information record, determine the current speaker and their identity information; The speech content corresponding to the on-site audio collected from the moment the target's speech was triggered is associated with the speaker and their identity information; Meeting minutes are generated based on the speeches of all speakers identified at the meeting and their associated identity information.
2. The method according to claim 1, wherein, The step of identifying whether a target speech trigger gesture exists in the video frame sequence image includes: Individual participants in the video frame sequence images are distinguished and located to obtain a personnel area sequence image of each participant in the meeting venue; The personnel region sequence image is input into the bone joint recognition model to obtain the coordinates of the key bone joints of the participants; Based on the coordinates of the key bone joints corresponding to the personnel area sequence image, the motion trajectory of the key bone joints is determined; Determine whether the posture corresponding to the motion trajectory of the key bone joint is the target speech trigger posture; If so, it is determined that a target speaking trigger posture exists in the video frame sequence image.
3. The method according to claim 1, wherein, The process of determining the current speaker and their identity information based on the identified target speaking trigger posture and meeting participant information records includes: The participant who presents the target speaking trigger gesture will be identified as the current speaker; The speaker's target facial image is matched with the facial images of attendees in the meeting attendee information record, and the speaker's identity information is determined based on the identity information of the matched attendees.
4. The method according to claim 3, wherein, The process of identifying the participant who presents the target speaking trigger gesture as the current speaker includes: Obtain candidate face images of participants who present the target speaking trigger posture; Centered on the video frame image corresponding to the candidate face image, a preset number of video frame images are selected forward and backward from the acquired video frame sequence image as the associated image sequence of the candidate face image. Based on the associated image sequence, determine the variation features of the mouth contour in the candidate face image; Based on the changing features of the mouth contour, determine whether the participant corresponding to the candidate face image is speaking; If someone is speaking, the participant who presents the target speaking trigger gesture will be identified as the current speaker.
5. The method according to claim 1, wherein, The process of determining the current speaker and their identity information based on the identified target speaking trigger posture and meeting participant information records includes: If multiple participants are identified to simultaneously exhibit the target speaking trigger posture, the direction of the sound source is determined based on the current on-site audio. Determine the location of the participants corresponding to each of the target speaking trigger postures; By comparing the location with the direction of the sound source, the participants whose location and the direction of the sound source are consistent within the preset location error range are determined as the current speakers; The speaker's target facial image is matched with the facial images of attendees in the meeting attendee information record, and the speaker's identity information is determined based on the identity information of the matched attendees.
6. The method according to any one of claims 1 to 5, further comprising, before the real-time acquisition of video frame sequence images and on-site audio of the conference venue: Obtain facial images and identity information of attendees; Based on the facial images and their identity information, a record of the meeting attendees is generated.
7. The method according to claim 6, wherein, The acquisition of the facial images and identity information of the attendees includes: At the meeting venue, facial images and identity information of attendees are recorded sequentially; or, Obtain the list of attendees for this meeting; Based on the list of attendees, retrieve the facial images and identity information of the attendees from the meeting management system.
8. A smart meeting recording device, comprising: The information acquisition unit is used to acquire video frame sequences and audio from the conference venue in real time. An action recognition unit is used to identify whether a target speech triggering posture exists in the video frame sequence image; The speaker identification unit is used to determine the current speaker and their identity information based on the identified target speaking trigger gesture and the information record of meeting participants. The content association unit is used to associate the speech content corresponding to the on-site audio collected from the moment the target speech is triggered with the speaker and their identity information; The record generation unit is used to generate meeting records based on the speeches of all speakers identified at the meeting and their associated identity information.
9. The apparatus according to claim 8, wherein, The action recognition unit includes: The joint coordinate acquisition module is used to input personnel region sequence images into the bone and joint recognition model to obtain the coordinates of key bone and joints of the participants; The trajectory determination module is used to determine the motion trajectory of the key bone joint based on the coordinates of the key bone joint corresponding to the personnel area sequence image; The posture determination module is used to determine whether the posture corresponding to the motion trajectory of the key bone joint is the target speech trigger posture. If so, it is determined that the target speech trigger posture exists in the video frame sequence image.
10. The apparatus according to claim 8, wherein, The speaker determination unit includes: The first speaker determination module is used to determine the participant who presents the target speaking trigger posture as the current speaker; The first information matching module is used to match the target face image of the speaker with the face images of the participants in the meeting participant information record, and determine the identity information of the speaker based on the identity information of the matched participants.
11. The apparatus according to claim 10, wherein, The first speaker determination module is specifically used for: Obtain candidate face images of participants who present the target speaking trigger posture; Centered on the video frame image corresponding to the candidate face image, a preset number of video frame images are selected forward and backward from the acquired video frame sequence image as the associated image sequence of the candidate face image. Based on the associated image sequence, determine the variation features of the mouth contour in the candidate face image; Based on the changing features of the mouth contour, determine whether the participant corresponding to the candidate face image is speaking; If someone is speaking, the participant who presents the target speaking trigger gesture will be identified as the current speaker.
12. The apparatus according to claim 8, wherein, The speaker determination unit includes: The sound source direction determination module is used to determine the sound source direction based on the current on-site audio if multiple participants are simultaneously identified as exhibiting the target speaking trigger posture. The orientation determination module is used to determine the orientation of the participants corresponding to each of the target speaking trigger postures; The second speaker determination module is used to compare the location with the direction of the sound source and determine the participants whose location and the direction of the sound source are consistent within a preset location error range as the current speaker. The second information matching module is used to match the target face image of the speaker with the face images of the participants in the meeting participant information record, and determine the identity information of the speaker based on the identity information of the matched participants.
13. The apparatus according to any one of claims 8 to 12, wherein the intelligent meeting recording apparatus further comprises: The information acquisition unit is used to acquire facial images and identity information of the participants; The information recording generation unit is used to generate a record of the meeting participants' information based on the facial image and their identity information.
14. The apparatus according to claim 13, wherein, The information acquisition unit includes: The first information acquisition module is used to sequentially record the facial images and identity information of the participants at the meeting venue; or... The second information acquisition module is used to obtain the list of attendees for this meeting, and retrieve the facial images and identity information of the attendees from the meeting management system based on the list of attendees.
15. A smart device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the smart meeting recording method as described in any one of claims 1 to 7.
16. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the intelligent meeting recording method as described in any one of claims 1 to 7.
17. A computer program product, which, when run on a smart device, causes the smart device to perform the smart meeting recording method as described in any one of claims 1 to 7.