Conference recording method and device, intelligent earphone and storage medium

By using smart headphones to simultaneously collect and mark the timestamps of video and audio streams during meetings, the speaking time can be automatically determined and associated with the speaker's image. This solves the problem of difficulty in associating speaker identity in existing technologies, realizes synchronized audio and video meeting recording, and improves the user experience.

CN121967631APending Publication Date: 2026-05-01GEER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GEER TECH CO LTD
Filing Date
2026-01-05
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing meeting recording methods cannot automatically and accurately associate each speech with its corresponding speaker, making it difficult for users to distinguish which specific sentence was spoken by which participant when reviewing meeting minutes, resulting in a poor user experience.

Method used

Using smart headphones equipped with image capture and audio acquisition components, the system simultaneously captures video and audio streams from a first-person perspective while the user is wearing the device, timestamps them synchronously, automatically determines the speaking time period, associates it with the speaker's image, and generates meeting minutes.

Benefits of technology

It enables the automatic capture and precise association of speaking content and speaker images during natural meeting participation, allowing users to directly obtain complete meeting records with synchronized audio and video and clear speaker identification, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967631A_ABST
    Figure CN121967631A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent earphones, and discloses a conference recording method and device, an intelligent earphone and a storage medium, and the method is applied to the intelligent earphone provided with an image shooting part and an audio collection part. The method comprises the following steps: under the condition that a user wears the intelligent earphone, acquiring a video stream of the user in a first person view angle through an image shooting component, and acquiring an audio stream through an audio acquisition component; performing timestamp synchronization marking on the video stream and the audio stream; determining a speaking time period in which speaking audio exists in the audio stream, and determining a spokesman image existing in the video stream in the speaking time period according to a timestamp synchronization marking result; and associating the speech audio with the spokesman image to generate a conference recording result. According to the method, synchronous collection and intelligent association of the first-person sound and picture are realized through the intelligent earphone, so that when the conference record is generated, a record result which corresponds to the sound and the picture and is clear in speaking affiliation can be directly obtained, and user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart headphones, and more particularly to a meeting recording method, apparatus, smart headphones, and storage medium. Background Technology

[0002] With the widespread adoption of remote collaboration and hybrid work models, online and offline meetings have become an important part of daily work. To improve the efficiency of post-meeting summaries and information review, various meeting platforms have generally integrated "automatic meeting minutes" functions. Specifically, this involves: first, capturing meeting audio through the microphone of the terminal device; then, using automatic speech recognition technology to convert the audio stream into text; and finally, performing post-processing such as segmentation and punctuation on the text to form preliminary meeting minutes.

[0003] However, current methods for generating meeting minutes only produce a continuous stream of text, failing to automatically and accurately associate each statement with its corresponding speaker's identity (such as name or role). Consequently, when users review the meeting minutes, they often struggle to distinguish which specific statement came from which participant amidst the continuous text paragraphs, resulting in a poor user experience. Summary of the Invention

[0004] The main purpose of this application is to provide a meeting recording method that aims to solve the technical problem of how to improve the user experience during meeting recording.

[0005] To achieve the above objectives, this application proposes a meeting recording method, which is applied to a smart headset equipped with an image capturing component and an audio acquisition component; The method includes: When the user is wearing the smart headphones, the video stream from the user's first-person perspective is captured by the image capturing component, and the audio stream is captured by the audio capturing component. The video stream and the audio stream are time-stamped and synchronized. Determine the speaking time period in the audio stream where the speech audio exists, and determine the speaker image in the video stream within the speaking time period based on the timestamp synchronization mark result; The audio of the speech is associated with the speaker's image to generate a meeting record.

[0006] In one embodiment, the step of determining the speaking time period in the audio stream where the spoken audio exists, and determining the speaker image in the video stream within the speaking time period based on the timestamp synchronization mark result, includes: The audio frame data in the audio stream is detected to obtain an audio event stream, which includes human voice detection results that characterize whether human voices exist in the audio frame data. The video frame data in the video stream is detected to obtain a video event stream, which includes images of the participants and temporary markers assigned to the images of the participants; Based on the voice detection results, the speaking time period of the spoken audio in the audio stream is determined; The target temporary marker corresponding to the speaking period is determined based on the timestamp synchronization marker result, and the participant image corresponding to the target temporary marker is used as the speaker image.

[0007] In one embodiment, the step of determining the speaking time segment of the spoken audio in the audio stream based on the human voice detection result includes: The audio event stream is detected through a preset decision window, and the human voice detection result within the preset decision window is determined to be the first cumulative duration of the presence of human voice. The video event stream is detected through the preset decision window to determine the second cumulative duration of the temporary marker within the preset decision window; When the first cumulative duration reaches the first preset duration threshold and the second cumulative duration reaches the second preset duration threshold, the speaking time period of the spoken audio in the audio stream is determined based on the human voice detection result.

[0008] In one embodiment, the video event stream also includes orientation detection results indicating whether the participant is facing the user; The step of determining the speaking time period of the spoken audio in the audio stream based on the voice detection result when the first cumulative duration reaches a first preset duration threshold and the second cumulative duration reaches a second preset duration threshold includes: Based on the detection results, determine the third cumulative duration of the participant's interaction with the user within the preset decision window; If the first cumulative duration reaches the first preset duration threshold, the second cumulative duration reaches the second preset duration threshold, and the third cumulative duration reaches the third preset duration threshold, the speaking time period of the spoken audio in the audio stream is determined based on the voice detection result.

[0009] In one embodiment, the step of detecting video frame data in the video stream to obtain a video event stream includes: The face region in the video data frame of the video stream is determined, and the key points of the face region are extracted. The feature extraction results are processed to obtain the current yaw angle and the current pitch angle; Based on the current yaw angle and the current pitch angle, an orientation detection result representing whether the participant is facing the user is obtained, and a video event stream is generated based on the orientation detection result.

[0010] In one embodiment, the audio event stream further includes the azimuth angle of the sound source when human voice is present in the audio frame data; The step of using the participant image corresponding to the target temporary marker as the speaker image includes: Determine the number of temporary markers for the target; When the number of markers is at least two, determine the image azimuth angle of the participant image for each of the target temporary markers, and determine the azimuth angle difference between the sound source and each of the image azimuth angles; Based on the azimuth difference values, the target participant image is selected as the speaker image from the participant image corresponding to the target temporary marker.

[0011] In one embodiment, the video event stream further includes face bounding box coordinates, which correspond to the image of the participant; The step of using the participant image corresponding to the target temporary marker as the speaker image includes: The corresponding face bounding box coordinates are determined based on the target temporary marker; The video frame data is cropped according to a preset expansion ratio, centered on the coordinates of the face bounding box. The cropped result will be used as the speaker's image.

[0012] Furthermore, to achieve the above objectives, this application also proposes a meeting recording device, the device comprising: The acquisition module is used to acquire a video stream from the user's first-person perspective through an image capturing component and an audio stream through an audio acquisition component when the user is wearing smart headphones. The tagging module is used to perform timestamp synchronization tagging on the video stream and the audio stream; The determination module is used to determine the speaking time period in the audio stream where there is spoken audio, and to determine the speaker image in the video stream within the speaking time period based on the timestamp synchronization mark result; The association module is used to associate the spoken audio with the speaker's image to generate meeting record results.

[0013] In addition, to achieve the above objectives, this application also proposes a smart earphone, which includes: an image capturing component and an audio acquisition component; The smart headset further includes: a memory, a processor, and a conference recording program stored in the memory and executable on the processor, the conference recording program being configured to implement the steps of the conference recording method as described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a meeting recording program is stored, and when the meeting recording program is executed by a processor, it implements the steps of the meeting recording method described above.

[0015] This application proposes a meeting recording method, apparatus, smart headset, and storage medium. The method is applied to a smart headset equipped with an image capturing component and an audio capturing component. The method includes: when a user is wearing the smart headset, acquiring a video stream from the user's first-person perspective through the image capturing component, and acquiring an audio stream through the audio capturing component; synchronizing the video stream and the audio stream with timestamps; determining the speaking time period in the audio stream where spoken audio exists, and determining the speaker's image in the video stream within the speaking time period based on the timestamp synchronization result; associating the spoken audio with the speaker's image to generate a meeting recording result.

[0016] This application incorporates an image capture component and an audio acquisition component, both mounted on the same smart headset, to simultaneously capture video and audio streams from the user's first-person perspective and timestamp both streams. In actual meetings, the device automatically identifies speaking segments in the audio stream and synchronously associates the corresponding speaker's image based on the timestamp. Compared to existing meeting recordings that rely on external fixed equipment and struggle to automatically associate audio and video information, this application, through first-person wearable capture and timestamp synchronization, can automatically capture and accurately associate speaking content with speaker images during natural meeting participation. Consequently, when reviewing meetings, users can directly obtain a complete meeting record with synchronized audio and video and clearly identified speakers, enhancing the user experience. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1This is a block diagram of the smart earphone structure proposed in the first embodiment of this application; Figure 2 This is a flowchart of the first embodiment of the meeting recording method proposed in this application; Figure 3 This is a flowchart of a second embodiment of the meeting recording method proposed in this application; Figure 4 This is a flowchart of the third embodiment of the meeting recording method proposed in this application; Figure 5 A diagram of a meeting recording device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a smart earphone suitable for implementing the embodiments of this application.

[0020] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0023] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0024] Understandably, with the widespread adoption of remote collaboration and hybrid work models, online and offline meetings have become an important part of daily work. To improve the efficiency of post-meeting summaries and information review, various meeting platforms have generally integrated "automatic meeting minutes" functions. Specifically, this involves: first, capturing meeting audio through the microphone of the terminal device; then, using automatic speech recognition technology to convert the audio stream into text; and finally, performing post-processing such as segmentation and punctuation on the text to form preliminary meeting minutes.

[0025] However, current methods for generating meeting minutes only produce a continuous stream of text, failing to automatically and accurately associate each statement with its corresponding speaker's identity (such as name or role). Consequently, when users review the meeting minutes, they often struggle to distinguish which specific statement came from which participant amidst the continuous text paragraphs, resulting in a poor user experience.

[0026] Therefore, to address the technical challenge of improving user experience during meeting recording, this embodiment proposes a meeting recording method. This method is applied to a smart headset equipped with an image capturing component and an audio acquisition component. The method includes: when the user is wearing the smart headset, acquiring a video stream from the user's first-person perspective through the image capturing component and acquiring an audio stream through the audio acquisition component; synchronizing the video stream and audio stream with timestamps; determining the speaking time period in the audio stream where the speaker's audio is present, and determining the speaker's image in the video stream within the speaking time period based on the timestamp synchronization result; associating the speaking audio with the speaker's image to generate a meeting recording result.

[0027] This embodiment includes an image capture unit and an audio acquisition unit, both mounted on the same smart headset to simultaneously capture video and audio streams from the user's first-person perspective, and timestamp both streams. In a meeting, the device automatically determines the speaking time segment in the audio stream and synchronously associates the speaker's image with the corresponding time segment based on the timestamp. Compared to existing meeting recordings that rely on external fixed equipment and struggle to automatically associate audio and video information, this embodiment, through first-person wearable capture and timestamp synchronization, can automatically capture and accurately associate speaking content with speaker images during natural meeting participation. Therefore, when reviewing the meeting, users can directly obtain a complete meeting record with synchronized audio and video and clearly identified speakers, improving the user experience.

[0028] For ease of understanding, the following is combined with Figures 1 to 6 The meeting recording method provided in the embodiments of this application, as well as the meeting recording method, apparatus, smart earphone, and storage medium provided in the following embodiments, will be described in detail.

[0029] This application provides a meeting recording method, which is applied to a smart headset. (See reference...) Figure 1 , Figure 1 This is a structural block diagram of the smart earphone proposed in the first embodiment of this application. Figure 1 As shown, the aforementioned smart earphones are equipped with an image capturing component and an audio acquisition component.

[0030] It is understood that the aforementioned image capturing component can refer to a component used to acquire optical image information and convert it into a digital video stream, such as a camera, image sensor, or camera module integrated into the front or side of the device, used to acquire visual data from the user's first-person perspective. The aforementioned audio acquisition component can refer to a hardware module used to capture ambient sound and convert it into a digital audio signal, such as a microphone, microphone array, digital microphone module, or audio input unit with acoustic signal acquisition capabilities, used to acquire voice information during a meeting.

[0031] Reference Figure 2 , Figure 2 This is a flowchart of the first embodiment of the meeting recording method proposed in this application.

[0032] like Figure 1 As shown, the method includes: Step S10: With the user wearing the smart headphones, the video stream from the user's first-person perspective is acquired by the image capturing component, and the audio stream is acquired by the audio capturing component.

[0033] It should be noted that the executing entity in this embodiment can be a multifunctional machine device with meeting recording capabilities, such as a smart headset integrating a camera and microphone, or a device capable of performing the aforementioned functions. This embodiment uses a smart headset (hereinafter referred to as the device) equipped with an image capturing component and an audio acquisition component for illustration, but this embodiment is not specifically limited thereto.

[0034] Understandably, the aforementioned first-person perspective refers to the field of vision simulated by the user's eyes. Considering that the user will wear the aforementioned headphones, when someone speaks, the wearer will subconsciously look at the speaker, thus allowing the device to directly capture a video stream from the wearer's first-person perspective.

[0035] Furthermore, it can be understood that the aforementioned video stream refers to a sequence of images continuously acquired and output in chronological order by the image capturing component. The aforementioned audio stream refers to a sequence of digital audio data continuously acquired and output in chronological order by the audio capturing component. In a specific implementation, after the user wears the device and activates the conference recording mode, the image capturing component continuously acquires a video stream from the user's first-person perspective, while the audio capturing component continuously acquires an audio stream from the environment.

[0036] Step S20: Timestamp synchronization is performed on the video stream and the audio stream.

[0037] Understandably, the timestamps mentioned above can be time stamps used to identify the moment data is generated or processed, such as millisecond-precise time values ​​generated based on Coordinated Universal Time, system clocks, or device local clocks. The synchronization markers mentioned above can refer to the operation or process of assigning timestamps based on the same time base to image frames in a video stream and audio packets in an audio stream, thereby aligning the two types of data on the timeline.

[0038] In its implementation, the aforementioned device generates and attaches a corresponding timestamp to each frame of video image and each audio data packet when acquiring video and audio streams respectively, so that the video data and audio data are aligned in time and have a unified timing reference.

[0039] Step S30: Determine the speaking time period in the audio stream where the speaking audio exists, and determine the speaker image in the video stream within the speaking time period based on the timestamp synchronization mark result.

[0040] It should be noted that the aforementioned audio recordings can be audio streams containing human voices and conveying semantic intent. The aforementioned speaking time segment can be a continuous time interval within the audio stream from the start to the end of the speech, typically defined by a start timestamp and an end timestamp. The aforementioned timestamp synchronization result can refer to the data association result with a unified temporal reference formed after the video stream and audio stream have undergone timestamp alignment processing. The aforementioned speaker image can be a facial image identified from the video frame corresponding to the speaking time segment, matching the spoken audio, based on the timestamp synchronization relationship.

[0041] In practical applications, the aforementioned device first analyzes the audio stream to identify the speaking period containing valid human voices. Then, using the previously established timestamp synchronization relationship between the video stream and the audio stream, it locates the corresponding video frame sequence within the speaking period and detects and determines the image of the person speaking from these video frames as the speaker image.

[0042] Step S40: Associate the spoken audio with the speaker's image to generate meeting record results.

[0043] It is understandable that the aforementioned meeting minutes can refer to a structured data output, the content of which at least includes correlated audio recordings of speeches and their corresponding speaker images. The format can be a text summary with facial markers, an interactive timeline, or a structured database entry. In practical use, the aforementioned device, based on a unified timestamp, binds audio data segments that overlap or correspond in time with speaker image data, and encapsulates this binding relationship along with other metadata (such as time information and temporary speaker IDs) to form structured meeting minutes data.

[0044] To facilitate understanding, the following example illustrates a specific scenario using headphones as the device, but does not impose any specific limitations on this embodiment: A user attends a meeting wearing smart headphones with an integrated side-facing camera. When colleague B speaks, the headphones simultaneously record environmental video and meeting audio, adding a unified timestamp. The headphones analyze the audio stream to identify the time period during which colleague B speaks (e.g., 14:20:15.200 to 14:20:22.500), and simultaneously, in the video stream corresponding to that time period, they use face detection and posture analysis to locate the facial image of colleague B speaking directly to the camera. After the meeting, the headphones associate and bind this audio (or transcribed text) with the extracted facial image of colleague B, generating a structured meeting minutes with facial identifiers, marking the speaker for each segment.

[0045] Furthermore, in order to obtain the speaker's image, the step of determining the speaking time period in the audio stream where the speaking audio exists, and determining the speaker's image in the video stream within the speaking time period based on the timestamp synchronization mark result, includes: Step S31: Detect the audio frame data in the audio stream to obtain an audio event stream, wherein the audio event stream includes human voice detection results that characterize whether human voice exists in the audio frame data.

[0046] It should be noted that the aforementioned audio frame data can refer to the smallest processing unit formed by dividing a continuous audio stream according to a preset duration (e.g., 20 milliseconds) or number of samples. The aforementioned detection can refer to the process of extracting acoustic features from audio frame data and applying a speech activity detection algorithm for analysis and judgment. The aforementioned audio event stream can refer to a data sequence consisting of a series of detection results arranged in chronological order, where each result corresponds to one or a group of audio frames. The aforementioned human voice detection result can refer to the core information field contained in the audio event stream, whose value is binary (e.g., "yes / no") or a probability value, used to indicate whether human voice exists within the corresponding time period.

[0047] In its implementation, the device segments the acquired audio stream into frames of fixed duration, generating continuous audio frame data. Then, the device analyzes each audio frame in real time, calculating its acoustic characteristics and comparing them with preset thresholds or models to determine whether the frame contains human voice. The device then organizes the information, including timestamps and the determination results, into an audio event stream in chronological order.

[0048] Step S32: Detect the video frame data in the video stream to obtain a video event stream, which includes participant images and temporary markers assigned to the participant images.

[0049] It is understandable that the aforementioned video frame data can refer to a single still image constituting a continuous video stream, which is a frame in a series of images acquired in chronological order. The aforementioned video event stream can refer to a serialized data structure containing specific visual event information (such as detected faces, assigned IDs, facial orientation, etc.) output after real-time analysis of continuous video frames. The aforementioned temporary tag can refer to a temporary identifier (ID) dynamically generated by the system and assigned to an image to uniquely and continuously track the same participant in the video stream; this identifier is typically valid within a single conference session.

[0050] In its specific implementation, the aforementioned device performs frame-by-frame analysis on the acquired video stream, detects the presence of a human face in each video frame, and assigns a unique temporary marker to each detected independent face region image and maintains it in subsequent frames. At the same time, the information containing the image and its marker is organized into a video event stream in chronological order.

[0051] Step S33: Determine the speaking time period of the spoken audio in the audio stream based on the human voice detection results; Step S34: Determine the target temporary marker corresponding to the speaking time period based on the timestamp synchronization marker result, and use the participant image corresponding to the target temporary marker as the speaker image.

[0052] It should be noted that the aforementioned voice detection result can refer to the judgment conclusion on the existence of human voices within a specific time period after analyzing audio data through a voice activity detection algorithm. The aforementioned speaking period can refer to a time interval containing valid speech with clearly defined start and end timestamps, determined based on continuous voice detection results. The aforementioned timestamp synchronization mark result can refer to the correspondence between the video stream and audio stream, established with a unified time reference after time alignment processing. The aforementioned target temporary mark can refer to the temporary identifier (ID) corresponding to the participant most likely to be the current speaker, selected from the video event stream based on spatiotemporal correlation analysis within the determined speaking period, such as "T20251217143000_01". The first 14 digits can represent a timestamp, and the last 2 digits can represent the intra-frame detection order. Of course, other forms can also be used, and this embodiment does not limit this. The aforementioned participant image can refer to the face or person image of the participant extracted from the video frame and bound to the target temporary mark.

[0053] In its implementation, the device first determines the start and end times of spoken audio in the audio stream based on the detection results of continuously occurring human voices, thus identifying the speaking period. Then, using the timestamp synchronization results, the device maps the speaking period onto the synchronized video event stream and analyzes multiple participants appearing within that period to determine who the speaker is; the temporary marker corresponding to that speaker is then designated as the target temporary marker. Finally, the device uses all participant images (or representative images selected from them) associated with the target temporary marker as the speaker image for that speaking period.

[0054] For ease of understanding, the following example illustrates the concept, but does not impose specific limitations on this embodiment: Assume that after analyzing the audio event stream, the device determines that the voice detection result is consistently "present" for 5 seconds from 15:10:05.000 to 15:10:10.000, thus identifying this interval as a speaking period. Through timestamp synchronization, the device finds records in the video event stream for the same time period (15:10:05.000 to 15:10:10.000). If the video event stream for this period shows one participant (i.e., a temporary marker P_A), then P_A is used as the target temporary marker, and the cropped face image corresponding to P_A is used as the speaker image to associate with the speaking audio segment. Assume that the video event stream for this period shows three participants (temporarily marked P_A, P_B, and P_C) appearing in the frame. The device further analyzes the visual behavior of the three individuals (such as whether they are facing the wearer) or combines this with sound source direction information, ultimately determining that P_B is the speaker. Therefore, P_B is identified as the target temporary marker, and all face cropped images associated with P_B during that time period stored in the system are used as speaker images to associate with the audio of that speech segment.

[0055] Furthermore, it should be emphasized that this embodiment can also maintain the continuity of the same face ID. For example, it can detect the continuity of temporary markers, that is, determine whether the temporary marker of the same speaker is continuously detected within a preset number of frames. If so, it can be considered a valid acquisition; if not, it can be considered an invalid acquisition, and a new temporary marker will be reassigned to it in subsequent frames. For example, if a face is not detected for 3 consecutive frames, it is determined that the person has left the field of vision, and its temporary marker is marked as invalid. If it is detected again in subsequent frames, a new temporary marker will be assigned.

[0056] This embodiment includes an image capture unit and an audio acquisition unit, both mounted on the same smart headset to simultaneously capture video and audio streams from the user's first-person perspective, and timestamp both streams. In a meeting, the device automatically determines the speaking time segment in the audio stream and synchronously associates the speaker's image with the corresponding time segment based on the timestamp. Compared to existing meeting recordings that rely on external fixed equipment and struggle to automatically associate audio and video information, this embodiment, through first-person wearable capture and timestamp synchronization, can automatically capture and accurately associate speaking content with speaker images during natural meeting participation. Therefore, when reviewing the meeting, users can directly obtain a complete meeting record with synchronized audio and video and clearly identified speakers, improving the user experience.

[0057] Based on the first embodiment, in the second embodiment, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 , Figure 3 This is a flowchart of a second embodiment of the meeting recording method proposed in this application. Further, to achieve more accurate meeting recording, this embodiment requires both video and audio to meet certain requirements before recording begins. Therefore, the step of determining the speaking time segment of the spoken audio in the audio stream based on the voice detection result includes: Step S331: Detect the audio event stream through a preset decision window, and determine that the human voice detection result in the preset decision window is the first cumulative duration of the presence of human voice; Step S332: Detect the video event stream through the preset decision window and determine the second cumulative duration of the temporary marker within the preset decision window; Step S333: When the first cumulative duration reaches the first preset duration threshold and the second cumulative duration reaches the second preset duration threshold, determine the speaking time period of the spoken audio in the audio stream based on the human voice detection result.

[0058] It should be noted that the aforementioned preset decision window can refer to a fixed-length (e.g., 1 second) sliding analysis interval defined on the timeline, used for segmented statistical analysis of audio and video event streams. The aforementioned first cumulative duration can refer to the sum of all time segments in the audio event stream where the voice detection result is "present voice" within the time range corresponding to the preset decision window.

[0059] Furthermore, the aforementioned second cumulative duration can refer to the sum of all time segments in the video event stream corresponding to a specific temporary marker that exist (i.e., are continuously tracked and not lost) within the time range corresponding to the preset decision window. The aforementioned first preset duration threshold can be a time threshold (e.g., 800 milliseconds) used to determine whether a person's voice is continuously valid. The aforementioned second preset duration threshold can be a time threshold (e.g., 300 milliseconds) used to determine whether a participant continuously appears in the field of vision.

[0060] In its implementation, the device defines a preset decision window and slides it along the timeline. For each window position, the device calculates the cumulative duration of human voices in the audio event stream within that window as the first cumulative duration, and simultaneously calculates the cumulative duration of a specific participant (identified by a temporary marker) continuously appearing in the video event stream within that window as the second cumulative duration. When the device determines that the first cumulative duration exceeds a first preset duration threshold and the second cumulative duration exceeds a second preset duration threshold, it triggers a valid speech detection event and precisely defines the speech period of the audio based on the human voice detection results within and before / after that window.

[0061] For ease of understanding, the following example illustrates the concept, but does not impose specific limitations on this embodiment: Assume the device uses a preset decision window of 1 second for analysis. Within the window from t=10.0s to t=11.0s, the device analyzes the audio event stream and finds that the total length of the segments containing human voices is 850 milliseconds, i.e., the first cumulative duration is 850 milliseconds (exceeding the first preset duration threshold of 800 milliseconds). Simultaneously, the device analyzes the video event stream and finds that the participant temporarily tagged with ID_03 has been continuously tracked for 500 milliseconds within this window, i.e., the second cumulative duration is 500 milliseconds (exceeding the second preset duration threshold of 300 milliseconds). Since both duration conditions are met, the device determines that a valid speech trigger has occurred, and uses the human voice detection results within this window as a clue to expand forward and backward to determine a complete speech period, for example, from t=9.8s to t=11.5s.

[0062] Furthermore, considering that eye contact usually takes place during meetings, a speech is only considered valid when the speaker is facing the user. The video event stream also includes a face detection result indicating whether the participant is facing the user. The step of determining the speaking time period of the spoken audio in the audio stream based on the voice detection result when the first cumulative duration reaches a first preset duration threshold and the second cumulative duration reaches a second preset duration threshold includes: Step S3331: Based on the detection results, determine the third cumulative duration of the participant facing the user within the preset decision window; Step S3332: When the first cumulative duration reaches the first preset duration threshold, the second cumulative duration reaches the second preset duration threshold, and the third cumulative duration reaches the third preset duration threshold, the speaking time period of the spoken audio in the audio stream is determined based on the human voice detection result.

[0063] It is understandable that the aforementioned orientation detection result can refer to a judgment conclusion calculated by analyzing the facial feature points of the participants in the video frame, representing whether their heads are facing the device wearer (user). The aforementioned third cumulative duration can refer to the total duration for which the participants face the user within the time range corresponding to the preset decision window. The aforementioned third preset duration threshold can be a time threshold value (e.g., 200 milliseconds) used to determine whether the participants continuously face the user.

[0064] In its implementation, the device, in addition to counting the duration of human voices and the duration of participants' presence, also counts the cumulative time a specific participant spends facing the user within a preset decision window, as a third cumulative duration. The device only determines a speaking event and the speaking period based on the following conditions: the first cumulative duration reaches the first preset duration threshold, the second cumulative duration reaches the second preset duration threshold, and the third cumulative duration reaches the third preset duration threshold.

[0065] For ease of understanding, the following example illustrates the concept, but does not impose specific limitations on this embodiment: Continuing with the 1-second decision window from the previous example (t=10.0s to t=11.0s), assume the device analysis determines that the first cumulative duration (voice duration) is 850 milliseconds, and the second cumulative duration (duration of participant ID_03's appearance) is 500 milliseconds. Furthermore, the device, through head posture analysis, determines that the third cumulative duration of ID_03 facing the user within this window is 400 milliseconds. If the three preset thresholds are 800 milliseconds, 300 milliseconds, and 200 milliseconds respectively, then all conditions are met (850>800, 500>300, 400>200). Therefore, the device confirms this is a valid speech trigger and, based on this window and the preceding and following voice detection results, defines a precise speech period (e.g., from t=9.8s to t=11.5s).

[0066] Furthermore, in order to obtain the detection results, the step of detecting video frame data in the video stream to obtain the video event stream includes: Step S321: Determine the face region in the video data frame of the video stream, and extract features from the key points of the face region; Step S322: Calculate the feature extraction results to obtain the current yaw angle and the current pitch angle; Step S323: Based on the current yaw angle and the current pitch angle, obtain the orientation detection result representing whether the participant is facing the user, and generate a video event stream based on the orientation detection result.

[0067] It is understood that the aforementioned video data frame can refer to a single static image constituting the video stream. The aforementioned face region can refer to a rectangular bounding box region containing a face identified and located in the video data frame by a face detection algorithm (such as a deep learning-based detector). The aforementioned key points can be indicators of feature points on the face with clear anatomical or geometric significance, such as the corners of the eyes, the tip of the nose, and the corners of the mouth.

[0068] Furthermore, it should be noted that the aforementioned current yaw angle can refer to the calculated angle of the face's left-right rotation around the vertical axis (Y-axis), used to represent the left-right shift of the face's orientation. The aforementioned current pitch angle can refer to the calculated angle of the face's up-down rotation around the horizontal axis (X-axis), used to represent the up-down shift of the face's orientation. The aforementioned facing detection result can refer to the binary judgment of "facing" or "not facing" generated based on whether the calculated yaw angle and pitch angle fall within a preset facing angle threshold range (e.g., the absolute value of the yaw angle is less than 15 degrees and the absolute value of the pitch angle is less than 10 degrees).

[0069] In practical use, the aforementioned device first processes each video data frame in the video stream, locates the face region within it, and extracts the coordinates of facial key points within that region. Next, the device uses the correspondence between these two-dimensional key point coordinates and the three-dimensional face model to calculate the yaw and pitch angles of the face in the current frame. Then, the device compares these two angle values ​​with a preset facing angle threshold to determine whether the participant is facing the user, and integrates this facing detection result along with other facial information into the video event stream record for that frame.

[0070] For ease of understanding, the following example illustrates the concept, but does not impose specific limitations on this embodiment: The device processes a video frame with a timestamp of 14:40:01.000. A face is detected in this frame, and its bounding box is located. Next, the device locates 68 facial key points within the face region, such as the left corner of the eye (x1, y1), the right corner of the eye (x2, y2), and the tip of the nose (x3, y3). Using the two-dimensional coordinates of these points, combined with a standard three-dimensional face model, the PnP algorithm calculates the face's current yaw angle as +5 degrees (slight right turn) and current pitch angle as -2 degrees (slight head tilt). Since the absolute values ​​of both angles are less than preset facing thresholds (assuming a yaw angle threshold of 15 degrees and a pitch angle threshold of 10 degrees), the device determines that this person is "facing the user." Finally, the device generates a video event stream record.

[0071] Based on the first and second embodiments, in the third embodiment, the content that is the same as or similar to that in Embodiments 1 and 2 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 , Figure 3 This is a flowchart of a third embodiment of the meeting recording method proposed in this application. Further, considering the presence of at least two faces, the audio event stream also includes the azimuth angle of the sound source when human voices are present in the audio frame data. The step of using the participant image corresponding to the target temporary marker as the speaker image includes: Step S341: Determine the number of temporary markers for the target; Step S342: When the number of markers is at least two, determine the image azimuth angle of the participant image for each of the target temporary markers, and determine the azimuth angle difference between the sound source and each of the image azimuth angles; Step S343: Select the target participant image as the speaker image from the participant images corresponding to the target temporary marker based on the azimuth angle difference values.

[0072] It should be noted that the aforementioned number of markers may refer to the number of temporary target markers obtained through initial screening. The aforementioned image azimuth angle of the participant image may refer to the horizontal angle of the participant relative to the device wearer, calculated based on the two-dimensional pixel coordinates of the participant image (face) in the video frame, combined with the camera's viewpoint and imaging model (e.g., the leftmost edge of the image corresponds to -30 degrees, and the rightmost edge corresponds to +30 degrees). The aforementioned sound source azimuth angle may refer to the angle relative to the device, estimated by processing audio signals through a microphone array, representing the direction of the sound source emitting the speech signal. The aforementioned azimuth angle difference may refer to the absolute value of the angle difference between the sound source azimuth angle and the azimuth angle of a participant image. The aforementioned target participant image may refer to the participant image selected after comparing azimuth angle differences, which best matches the sound source direction; that is, the face image ultimately identified as the speaker.

[0073] In its implementation, after obtaining the target temporary marker, the device first determines the number of target temporary markers. If only one exists, subsequent association processing can proceed directly. If the device determines that multiple initially selected target temporary markers exist within a speaking time period, it calculates the image azimuth angle of the participant image corresponding to each marker. Simultaneously, the device acquires the sound source azimuth angle recorded in the audio event stream during that time period. Next, the device calculates the difference between the sound source azimuth angle and the azimuth angle of each participant image. Finally, the device selects the image corresponding to the participant with the smallest difference as the final speaker image.

[0074] Furthermore, in order to obtain a complete image of the speaker, the video event stream also includes face bounding box coordinates, which correspond to the image of the participant; The step of using the participant image corresponding to the target temporary marker as the speaker image includes: Step S344: Determine the coordinates of the corresponding face bounding box based on the target temporary marker; Step S345: Using the coordinates of the face bounding box as the center, crop the video frame data according to a preset expansion ratio; Step S346: Use the cropped result as the speaker's image.

[0075] It is understood that the aforementioned face bounding box coordinates can refer to the coordinates of the rectangular box used to define the speaker's facial region in a video frame, typically expressed as the pixel coordinates of the top-left and bottom-right corners of the box (e.g., (x1, y1, x2, y2)). The aforementioned preset expansion ratio can refer to the percentage (e.g., 10%) set to expand outwards from the original bounding box to include more context (such as hair, shoulders) or ensure facial integrity when cropping the image. The specific preset expansion ratio can be set according to actual conditions, and this embodiment does not impose any limitations on it. The aforementioned cropping result can refer to the new image obtained after performing the cropping operation, with adjusted size and content.

[0076] In its implementation, after determining the temporary target marker corresponding to the speaker, the device retrieves the coordinates of the face bounding box corresponding to that marker in the current video frame from the video event stream. Then, based on these coordinates, the device calculates a new, larger cropping box coordinate according to a preset expansion ratio. Next, the device extracts the corresponding image region in real time from the stored original video frame data according to the new coordinates. Finally, the device uses the extracted image region as the final speaker image.

[0077] For ease of understanding, the following example illustrates the concept, but does not limit the scope of this embodiment: The device determines that the target temporary marker corresponding to the speaker is ID_07. The system searches for the record of ID_07 in the video event stream and finds that the data entry corresponding to the timestamp 15:00:00.500 contains the face bounding box coordinates (100, 150, 200, 250) and other information. When it is necessary to generate a speaker image, the system reads the video frame data of frame 15:00:00.500 from the stored original video file. Based on a preset expansion ratio of 10%, the system calculates the new cropping box coordinates to be approximately (95, 145, 205, 255). Then, the system performs a cropping operation on the frame data based on this new coordinate to obtain a new image. This cropping result is saved as the speaker image of speaker ID_07 at that time point, thus obtaining a complete face image.

[0078] It is important to emphasize that after cropping according to the preset expansion ratio, in order to highlight the speaker while protecting privacy, this embodiment can apply a Gaussian blur (e.g., using a blur kernel size of 5×5) to the cropped face area, preserving the facial outline but hiding facial details. Alternatively, a highlighted border (e.g., a solid red line with a width of 2px) can be added to the cropped face area without altering the face itself, thus highlighting the speaker while protecting privacy. Of course, other methods can also be used, and this embodiment does not limit them.

[0079] It should also be emphasized that the user-related data (e.g., facial data, audio data) involved in this embodiment are all obtained with the user's permission or consent; that is, when this embodiment is applied to a specific product or technology, user permission is required to obtain and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations and regulatory standards of the relevant countries and regions.

[0080] This embodiment also provides a first embodiment of a conference recording device; please refer to [reference needed]. Figure 5 , Figure 5 A diagram of a meeting recording device provided in this application embodiment is shown. The meeting recording device includes: The acquisition module is used to acquire a video stream from the user's first-person perspective through an image capturing component and an audio stream through an audio acquisition component when the user is wearing smart headphones. The tagging module is used to perform timestamp synchronization tagging on the video stream and the audio stream; The determination module is used to determine the speaking time period in the audio stream where there is spoken audio, and to determine the speaker image in the video stream within the speaking time period based on the timestamp synchronization mark result; The association module is used to associate the spoken audio with the speaker's image to generate meeting minutes. The determining module is further configured to detect audio frame data in the audio stream to obtain an audio event stream, the audio event stream including a voice detection result indicating whether a human voice exists in the audio frame data; detect video frame data in the video stream to obtain a video event stream, the video event stream including participant images and temporary markers assigned to the participant images; determine the speaking time period of the speaking audio in the audio stream based on the voice detection results; determine the target temporary marker corresponding to the speaking time period according to the timestamp synchronization marker result, and use the participant image corresponding to the target temporary marker as the speaker image.

[0081] Referring to the first embodiment of the meeting recording device, this embodiment also proposes a second embodiment of the meeting recording device. The contents that are the same as or similar to those in the first embodiment of the meeting recording device can be referred to the above description, and will not be repeated hereafter.

[0082] The determining module is further configured to detect the audio event stream through a preset decision window, and determine a first cumulative duration of human voice detection within the preset decision window; detect the video event stream through the preset decision window, and determine a second cumulative duration of the temporary marker within the preset decision window; and, when the first cumulative duration reaches a first preset duration threshold and the second cumulative duration reaches a second preset duration threshold, determine the speaking time period of the spoken audio in the audio stream based on the human voice detection result. The determining module is further configured to determine, based on the detection results, the third cumulative duration of the participant facing the user within the preset decision window; If the first cumulative duration reaches the first preset duration threshold, the second cumulative duration reaches the second preset duration threshold, and the third cumulative duration reaches the third preset duration threshold, the speaking time period of the spoken audio in the audio stream is determined based on the human voice detection result. The determining module is further configured to determine the face region of the video data frame in the video stream, and extract features from the key points of the face region; calculate the feature extraction results to obtain the current yaw angle and the current pitch angle; obtain the orientation detection result representing whether the participant is facing the user based on the current yaw angle and the current pitch angle, and generate a video event stream based on the orientation detection result.

[0083] Referring to the first embodiment and the second embodiment of the meeting recording device, this embodiment also proposes a third embodiment of the meeting recording device. The contents that are the same as or similar to the first embodiment and the second embodiment of the meeting recording device can be referred to the above description, and will not be repeated hereafter.

[0084] The determining module is further configured to determine the number of the target temporary markers; when the number of markers is at least two, determine the image azimuth angle of the participant image of each target temporary marker, and determine the azimuth angle difference between the sound source and each image azimuth angle; and select the target participant image as the speaker image from the participant images corresponding to the target temporary marker based on each azimuth angle difference. The determining module is further configured to determine the corresponding face bounding box coordinates based on the target temporary marker; crop the video frame data according to a preset expansion ratio with the face bounding box coordinates as the center; and use the cropping result as the speaker image.

[0085] The meeting recording device provided in this embodiment, employing the meeting recording method described in the above embodiments, can solve the technical problem of how to improve the user experience during meeting recording. Compared with the prior art, the beneficial effects of the meeting recording device provided in this embodiment are the same as those of the meeting recording method described in the above embodiments, and other technical features in the meeting recording device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0086] This embodiment provides a smart headset, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the meeting recording method in Embodiment 1 above.

[0087] The following is for reference. Figure 6 , Figure 6 This is a schematic diagram of the structure of a smart earphone suitable for implementing the embodiments of this application. Figure 6 The smart earphones shown are merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0088] like Figure 6 As shown, the smart headset includes an image capturing unit and an audio acquisition unit, and also includes a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the smart headset. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the smart headset to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows smart headsets with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0089] Specifically, according to this embodiment, the process described above with reference to the flowchart can be implemented as a computer software program. For example, this embodiment includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the disclosed embodiments of this embodiment.

[0090] The smart headset provided in this embodiment, employing the meeting recording method described in the above embodiments, can solve the technical problem of how to improve the user experience during meeting recording. Compared with the prior art, the beneficial effects of the smart headset provided in this embodiment are the same as those of the meeting recording method provided in the above embodiments, and other technical features of the smart headset are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0091] It should be understood that the various parts disclosed in this embodiment can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0092] The above description is merely a specific implementation of this embodiment, but the protection scope of this embodiment is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this embodiment should be included within the protection scope of this embodiment. Therefore, the protection scope of this embodiment should be determined by the protection scope of the claims.

[0093] This embodiment provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the meeting recording method in the above embodiment.

[0094] The computer-readable storage medium provided in this embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0095] The aforementioned computer-readable storage medium may be included in the smart headphones; or it may exist independently and not assembled into the smart headphones.

[0096] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the smart headset, cause the smart headset to: take meeting minutes.

[0097] Computer program code for performing the operations of this embodiment can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this embodiment. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0099] The modules described in this embodiment can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0100] The readable storage medium provided in this embodiment is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described meeting recording method, thereby solving the technical problem of how to improve the user experience during meeting recording. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this embodiment are the same as those of the meeting recording method provided in the above embodiments, and will not be repeated here.

[0101] The above descriptions are only some embodiments and do not limit the patent scope of this embodiment. All equivalent structural transformations made based on the technical concept of this application and the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included within the patent protection scope of this application.

Claims

1. A method for taking meeting minutes, characterized in that, The method is applied to smart headphones equipped with an image capturing component and an audio acquisition component; The method includes: When the user is wearing the smart headphones, the video stream from the user's first-person perspective is captured by the image capturing component, and the audio stream is captured by the audio capturing component. The video stream and the audio stream are time-stamped and synchronized. Determine the speaking time period in the audio stream where the speech audio exists, and determine the speaker image in the video stream within the speaking time period based on the timestamp synchronization mark result; The audio of the speech is associated with the speaker's image to generate a meeting record.

2. The method as described in claim 1, characterized in that, The step of determining the speaking time period in the audio stream where the spoken audio exists, and determining the speaker's image in the video stream within the speaking time period based on the timestamp synchronization mark result, includes: The audio frame data in the audio stream is detected to obtain an audio event stream, which includes human voice detection results that characterize whether human voices exist in the audio frame data. The video frame data in the video stream is detected to obtain a video event stream, which includes images of the participants and temporary markers assigned to the images of the participants; Based on the voice detection results, the speaking time period of the spoken audio in the audio stream is determined; The target temporary marker corresponding to the speaking period is determined based on the timestamp synchronization marker result, and the participant image corresponding to the target temporary marker is used as the speaker image.

3. The method as described in claim 2, characterized in that, The step of determining the speaking time segment of the spoken audio in the audio stream based on the human voice detection result includes: The audio event stream is detected through a preset decision window, and the human voice detection result within the preset decision window is determined to be the first cumulative duration of the presence of human voice. The video event stream is detected through the preset decision window to determine the second cumulative duration of the temporary marker within the preset decision window; When the first cumulative duration reaches the first preset duration threshold and the second cumulative duration reaches the second preset duration threshold, the speaking time period of the spoken audio in the audio stream is determined based on the human voice detection result.

4. The method as described in claim 3, characterized in that, The video event stream also includes detection results indicating whether the participants are facing the user; The step of determining the speaking time period of the spoken audio in the audio stream based on the voice detection result when the first cumulative duration reaches a first preset duration threshold and the second cumulative duration reaches a second preset duration threshold includes: Based on the detection results, determine the third cumulative duration of the participant's interaction with the user within the preset decision window; If the first cumulative duration reaches the first preset duration threshold, the second cumulative duration reaches the second preset duration threshold, and the third cumulative duration reaches the third preset duration threshold, the speaking time period of the spoken audio in the audio stream is determined based on the human voice detection result.

5. The method as described in claim 4, characterized in that, The step of detecting video frame data in the video stream to obtain a video event stream includes: The face region in the video data frame of the video stream is determined, and the key points of the face region are extracted. The feature extraction results are processed to obtain the current yaw angle and the current pitch angle; Based on the current yaw angle and the current pitch angle, an orientation detection result representing whether the participant is facing the user is obtained, and a video event stream is generated based on the orientation detection result.

6. The method as described in claim 2, characterized in that, The audio event stream also includes the azimuth angle of the sound source when human voice is present in the audio frame data; The step of using the participant image corresponding to the target temporary marker as the speaker image includes: Determine the number of temporary markers for the target; When the number of markers is at least two, determine the image azimuth angle of the participant image for each of the target temporary markers, and determine the azimuth angle difference between the sound source and each of the image azimuth angles; Based on the azimuth difference values, the target participant image is selected as the speaker image from the participant image corresponding to the target temporary marker.

7. The method as described in claim 2, characterized in that, The video event stream also includes face bounding box coordinates, which correspond to the participant's image; The step of using the participant image corresponding to the target temporary marker as the speaker image includes: The corresponding face bounding box coordinates are determined based on the target temporary marker; The video frame data is cropped according to a preset expansion ratio, centered on the coordinates of the face bounding box. The cropped result will be used as the speaker's image.

8. A meeting recording device, characterized in that, The device includes: The acquisition module is used to acquire a video stream from the user's first-person perspective through an image capturing component and an audio stream through an audio acquisition component when the user is wearing smart headphones. The tagging module is used to perform timestamp synchronization tagging on the video stream and the audio stream; The determination module is used to determine the speaking time period in the audio stream where there is spoken audio, and to determine the speaker image in the video stream within the speaking time period based on the timestamp synchronization mark result; The association module is used to associate the spoken audio with the speaker's image to generate meeting record results.

9. A smart earphone, characterized in that, The smart earphones include: an image capturing component and an audio acquisition component; The smart headset further includes: a memory, a processor, and a conference recording program stored in the memory and executable on the processor, wherein the conference recording program, when executed by the processor, implements the steps of the conference recording method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a conference recording program, which, when executed by a processor, implements the steps of the conference recording method as described in any one of claims 1 to 7.