Voice recording method, device, equipment and medium
By acquiring the wearer's eye-tracking data and environmental images through smart glasses, and dynamically segmenting speech segments, this technology solves the problem of inaccurate speech recording segmentation in existing technologies, and achieves precise segmentation of speech segments and quantification of attention.
Patent Information
- Application Number
- CN202511439595.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
AI Technical Summary
Existing smart glasses cannot segment ambient sounds, resulting in inaccurate voice recording.
By acquiring the wearer's eye movement data and environmental images, the fixation duration of the gaze area is determined, and the attention weight of the speech segment is determined based on the fixation duration and confidence level. The speech segments are dynamically divided and labeled with attention weights and tags.
It enables dynamic and seamless segmentation of ambient audio signals based on the wearer's gaze, quantifies the attention level of speech segments, and improves the accuracy and usability of voice recording.
Smart Images

Figure CN120909437A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio processing, and more particularly, to a speech recording method and device, equipment and medium. BACKGROUND
[0002] At present, in the scenarios of conferences, academic lectures and the like, it has become a trend to use lightweight intelligent glasses to record and organize the collected speech.
[0003] However, the existing intelligent glasses can only record and convert the environmental sound into text, and cannot divide the environmental sound into paragraphs. SUMMARY
[0004] An object of the present application is to provide a new technical solution for speech recording.
[0005] According to a first aspect of the present application, a speech recording method is provided, applied to an intelligent head-mounted device, comprising: obtaining eye movement data of a wearer during audio collection, an environmental image under the visual angle of the wearer, and an environmental audio signal, the environmental audio signal being an audio signal to be recorded, and the eye movement data, the environmental image and the environmental audio signal being associated with a time stamp representing the corresponding collection time; determining a gaze dwell period of a gaze area of the wearer according to the eye movement data and the environmental image; determining the environmental audio signal obtained during a target gaze dwell period as an independent speech segment, the target gaze dwell period being a gaze dwell period with a duration greater than a preset duration; For any independent speech segment, determining a focus weight of the independent speech segment according to the corresponding duration of the gaze dwell period of the independent speech segment, and marking the corresponding focus weight for the independent speech segment.
[0006] Optionally, the determining of the gaze dwell period of the gaze area of the wearer according to the eye movement data and the environmental image comprises: determining the gaze content type and the gaze dwell period of the gaze area of the wearer according to the eye movement data and the environmental image; Before the determining of the environmental audio signal obtained during the target gaze dwell period as an independent speech segment, the method further comprises: determining a gaze dwell period with a duration greater than the preset duration and a gaze content type corresponding to the gaze area as a target gaze dwell period.
[0007] Optionally, the determining of the focus weight of the independent speech segment according to the corresponding duration of the gaze dwell period of the independent speech segment comprises: determine a confidence of the independent voice segment according to the independent voice segment; determine a focus weight of the independent voice segment according to a corresponding time length of the independent voice segment, the confidence of the independent voice segment, and a sum of the corresponding time lengths of the gaze dwell time segments of all the independent voice segments.
[0008] Optionally, after the environmental audio signal acquired during the target gaze dwell time segment is determined as an independent voice segment, the method further comprises: determine a keyword in the independent voice segment for any of the independent voice segments; determine a target label of the independent voice segment according to the gaze content type corresponding to the independent voice segment, the keyword, and a preset mapping relationship, wherein the preset mapping relationship reflects a corresponding relationship among the gaze content type, the keyword, and the label; add the target label to the independent voice segment.
[0009] Optionally, the determining of the target label of the independent voice segment according to the gaze content type corresponding to the independent voice segment, the keyword, and the preset mapping relationship comprises: determine at least one candidate label of the independent voice segment according to the gaze content type corresponding to the independent voice segment, the keyword, and the preset mapping relationship; determine a candidate label with the highest priority in the at least one candidate label as the target label.
[0010] Optionally, the determining of the gaze dwell time segment of the gaze region of the wearer according to the eye movement data and the environmental image comprises: determine an initial gaze point coordinate of the wearer according to the eye movement data; acquire head posture information of the wearer during the acquisition of the eye movement data; calibrate the initial gaze point coordinate according to the head posture information to obtain a target gaze point coordinate; determine the gaze region of the wearer according to the target gaze point coordinate and the environmental image.
[0011] Optionally, the acquiring of the eye movement data of the wearer, the environmental image from the perspective of the wearer, and the environmental audio signal during the audio acquisition comprises: acquire an environmental sub-audio signal collected by each microphone in a microphone array; synthesize the environmental sub-audio signal collected by each microphone in the microphone array according to a weight of each microphone to obtain the environmental audio signal; The method further comprises: In a case where the gaze point target coordinate of the wearer is determined, a weight of an environmental sub-audio signal collected by a microphone in a gaze direction corresponding to the gaze point target coordinate is set to be greater than a weight of an environmental sub-audio signal collected by the microphone outside the gaze direction.
[0012] According to a second aspect of the present application, a voice recording device applied to a smart head-mounted device is provided, and the voice recording device comprises: The acquisition module is configured to acquire eye movement data of a wearer during audio collection, an environment image under a visual angle of the wearer, and an environment audio signal, the environment audio signal being an audio signal to be recorded, and the eye movement data, the environment image, and the environment audio signal being associated with a timestamp representing a corresponding collection time point; The first determination module is configured to determine a gaze staying time period of a gaze region of the wearer according to the eye movement data and the environment image. The second determination module is configured to determine an environment audio signal acquired during a target gaze staying time period as an independent voice segment, the target gaze staying time period being a gaze staying time period with a time length greater than a preset time length. The marking module is configured to, for any independent voice segment, determine an attention weight of the independent voice segment according to a time length corresponding to a gaze staying time period of the independent voice segment, and mark the independent voice segment with the corresponding attention weight.
[0013] According to a third aspect of the present application, a smart head-mounted device is provided, and the smart head-mounted device comprises the device according to the second aspect; or The smart head-mounted device comprises a memory and a processor, the memory is configured to store computer instructions, and the processor is configured to call the computer instructions from the memory to execute the method according to any one of the first aspect.
[0014] According to a fourth aspect of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, the computer program is executed by a processor to implement the method according to any one of the first aspect.
[0015] The application provides a voice recording method, applied to a smart head-mounted device, comprising: acquiring eye movement data of a wearer during audio acquisition, an environmental image under a visual angle of the wearer, and an environmental audio signal, the environmental audio signal being an audio signal to be recorded, and the eye movement data, the environmental image, and the environmental audio signal being associated with a timestamp representing a corresponding acquisition time; determining a gaze dwell period of a gaze area of the wearer according to the eye movement data and the environmental image; determining the environmental audio signal acquired during a target gaze dwell period as an independent voice segment, the target gaze dwell period being a gaze dwell period with a time length greater than a preset time length; for any independent voice segment, determining an attention weight of the independent voice segment according to a time length corresponding to the gaze dwell period of the independent voice segment, and marking the independent voice segment with the corresponding attention weight. The method can realize dynamic and unobtrusive segment division of the environmental audio signal according to the gaze of the wearer, and can also realize attention quantification of the independent voice segment.
[0016] Other features of the present application, and their advantages, will become apparent from the following detailed description of illustrative embodiments of the present application, when considered in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments of the present application and, together with the description, serve to explain the principles of the present application.
[0018] Figure 1 is a flowchart of a voice recording method provided by the present application; Figure 2 is a structural schematic diagram of a voice recording device provided by the present application; Figure 3 is a structural schematic diagram of a smart head-mounted device provided by the present application. DETAILED DESCRIPTION
[0019] Various illustrative embodiments of the present application will now be described in detail with reference to the accompanying figures. It should be noted that the relative arrangements of the components and steps set forth in the embodiments, numerical expressions, and numerical values, unless specifically stated otherwise, do not limit the scope of the present application.
[0020] The following description of at least one illustrative embodiment is merely exemplary in nature and is in no way intended to limit the present application or its application or uses.
[0021] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail herein, but should be considered as part of the specification, where appropriate.
[0022] In all of the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments can have different values.
[0023] It should be noted that like reference numerals and letters refer to like items throughout the attached drawings, and once an item is defined in one drawing, it is not necessary to discuss it further in subsequent drawings.
[0024] The present application provides a voice recording method, which is applied to a smart head-mounted device. The smart head-mounted device is exemplarily AR glasses.
[0025] As shown in the Figure 1 The voice recording method provided by the present application includes the following steps S1100 to S1400.
[0026] In step S1100, the eye movement data of the wearer during audio acquisition, the environmental image under the visual angle of the wearer, and the environmental audio signal are acquired.
[0027] The environmental audio signal is the audio signal to be recorded, and the eye movement data, the environmental image, and the environmental audio signal are associated with the time stamp of the corresponding acquisition time.
[0028] In an embodiment of the present application, the above step S1100 can be executed at the trigger of the wearer. For example, the wearer triggers the above step S1100 by voice, specific button pressing, or specific touch action.
[0029] The smart head-mounted device is provided with an eye movement tracking module, a microphone (or microphone array), and an image acquisition module. The eye movement tracking module is used to acquire eye movement data and associate the corresponding acquisition time with the eye movement data; the microphone array is used to pick up the environmental audio signal and associate the corresponding acquisition time with the environmental audio signal; and the image acquisition module is used to acquire the environmental image under the visual angle of the wearer and associate the corresponding acquisition time with the environmental image.
[0030] In step S1200, the gaze dwell period of the gaze area of the wearer is determined according to the eye movement data and the environmental image.
[0031] In the present embodiment, the eye movement data in the above step S1200 is the eye movement data within the time period when the gaze point can be accurately determined. In addition, the environmental image in the above step S1200 is the environmental image within the acquisition time corresponding to the eye movement data in step S1200. The environmental image within the acquisition time corresponding to the eye movement data in the above step S1200 can be determined according to the time stamp associated with the eye movement data and the time stamp associated with the environmental image.
[0032] Furthermore, the specific implementation of step S1200 is as follows: determine the coordinates of the user's gaze point based on eye-tracking data; determine the location of the user's gaze point coordinates in the environmental image as the gaze point; determine the area of a preset size centered on the gaze point in the environmental image as the gaze region; and count the gaze duration corresponding to the same gaze region. The gaze region is also called the area of interest (AOI).
[0033] In one embodiment of this application, the gaze coordinates determined from eye-tracking data are often inaccurate when the wearer's head posture changes. Therefore, to obtain accurate gaze coordinates, this application includes an operation to calibrate the gaze coordinates. In this regard, step S1200 above includes at least the steps of the wearer's gaze area as shown in steps S1210 to S1213 below.
[0034] Step S1210: Determine the initial coordinates of the wearer's gaze point based on eye-tracking data.
[0035] In this embodiment, the gaze coordinates of the wearer determined by eye-tracking data according to conventional methods are recorded as the initial gaze coordinates.
[0036] Step S1211: Obtain the wearer's head posture information during eye-tracking data acquisition.
[0037] In this embodiment, the smart head-mounted device is equipped with a 6DOF sensor, which can collect 6DOF information as the wearer's head posture information.
[0038] Step S1212: Based on the head posture information, calibrate the initial coordinates of the gaze point to obtain the target coordinates of the gaze point.
[0039] For step S1212 above, the head pose matrix is first determined based on the head pose information. Further, the initial coordinates of the gaze point are calibrated using the following formula to obtain the target coordinates of the gaze point.
[0040] (Formula 1) in, The coordinates of the gaze point target; This is the sliding attenuation factor, set empirically. This is the head pose matrix. These are the initial coordinates of the gaze point.
[0041] Step S1213: Determine the wearer's gaze area based on the coordinates of the gaze point target and the environmental image.
[0042] The gaze point target coordinate obtained through the step S1212 can be used to more accurately determine the gaze area of the wearer from the environmental image.
[0043] It can be understood that the gaze area of the wearer is usually changing. For example, when the wearer is attending a meeting, the wearer first gazes at the meeting speaker, and at this time, the gaze area is the area where the meeting speaker is located. When the meeting speaker is narrating the PPT of the screen projection, the wearer gazes from the meeting speaker to the PPT of the screen projection, and at this time, the gaze area is the area where the PPT of the screen projection is located. That is, the gaze stay period of at least one gaze area can be obtained through the step S1200.
[0044] In step S1300, the environmental audio signal obtained during the target gaze stay period is determined as an independent speech segment.
[0045] The target gaze stay period is a gaze stay period with a duration greater than a preset duration. The preset duration is the shortest duration of the wearer in the gaze state, which can be set according to experience. In addition, the environmental audio signal in the target gaze stay period can be determined from the environmental audio signal obtained based on the step S1100 according to the time stamp associated with the environmental audio signal.
[0046] Corresponding to one target gaze stay period, since the target gaze stay period corresponds to a duration greater than a preset duration, the wearer is in a gaze state during the target gaze stay period, that is, focuses on the gaze content. On this basis, the environmental audio signal collected in the target gaze stay period, that is, the independent speech segment in the step S1300, can reflect the gaze content that the wearer focuses on. In this way, dynamic and non-intrusive segment division of the environmental audio signal according to the gaze condition of the wearer can be realized.
[0047] In step S1400, for any independent speech segment, the attention weight of the independent speech segment is determined according to the duration corresponding to the gaze stay period of the independent speech segment, and the corresponding attention weight is marked for the independent speech segment.
[0048] In an embodiment of the present application, for any independent speech segment, the proportion of the duration corresponding to the independent speech segment in the sum of the durations corresponding to all independent speech segments can be determined as the attention weight of the independent speech segment.
[0049] In an embodiment of the present application, for any independent speech segment, the attention weight of the independent speech segment can be determined according to the following steps S1410 and S1411.
[0050] In step S1410, the confidence of the independent speech segment is determined according to the independent speech segment.
[0051] In one embodiment of this application, step S1410 can be implemented as follows: The confidence level of each sentence in an independent speech segment is calculated using an Automatic Speech Recognition (ASR) algorithm. The weighted sum of the confidence levels of all sentences in the independent speech segment is then determined as the confidence level of that independent speech segment. It should be noted that the ASR algorithm outputs the confidence level of each recognized sentence while recognizing speech. Furthermore, when calculating the aforementioned weighted sum, the weight can be determined based on the duration of each sentence in the independent speech segment.
[0052] Step S1411: Determine the attention weight of an independent speech segment based on the duration of the gaze dwell time corresponding to the independent speech segment, the confidence level of the independent speech segment, and the sum of the durations of the gaze dwell times of all independent speech segments.
[0053] Specifically, step S1411 above is implemented through the following formula two.
[0054] (Formula 2) in, This refers to the sequence number of an independent speech segment. For the first The attention weight of each individual speech segment For the first Confidence of each individual speech segment The number of independent speech segments, This represents the duration of the gaze pause for the i-th independent speech segment.
[0055] After obtaining the attention weight for each independent speech segment using the above methods, the corresponding attention weight is then assigned to each independent speech segment. This allows for the quantification of attention to independent speech segments.
[0056] The application provides a voice recording method applied to a smart head-mounted device, including: acquiring eye movement data of a wearer during audio collection, an environmental image under a visual angle of the wearer, and an environmental audio signal, the environmental audio signal being an audio signal to be recorded, and the eye movement data, the environmental image, and the environmental audio signal being associated with a timestamp representing a corresponding collection time; determining a gaze dwell period of a gaze area of the wearer according to the eye movement data and the environmental image; determining an environmental audio signal acquired during a target gaze dwell period as an independent voice segment, the target gaze dwell period being a gaze dwell period with a length greater than a preset length; for any independent voice segment, determining an attention weight of the independent voice segment according to a length corresponding to the gaze dwell period of the independent voice segment, and marking the independent voice segment with the corresponding attention weight. The method can realize dynamic and unobtrusive segment division of the environmental audio signal according to the gaze of the wearer, and can also realize attention quantification of the independent voice segment.
[0057] In an embodiment of the application, the step S1200 is specifically implemented through the following step S1220.
[0058] In the step S1220, the gaze content type and the gaze dwell period of the gaze area of the wearer are determined according to the eye movement data and the environmental image.
[0059] In the embodiment, on the one hand, the gaze dwell period of the gaze area of the wearer is determined according to the eye movement data and the environmental data, and on the other hand, the gaze content type of the gaze area of the wearer is determined according to the eye movement data and the environmental image. For the latter, after the gaze area of the wearer is determined according to the eye movement data and the environmental image, the type of the gaze content in the gaze area in the environmental image is identified. The type is exemplarily PPT, face, and whiteboard.
[0060] On the basis of the step S1220, the voice recording method provided by the application further includes the following step S1310 for determining the target gaze dwell period before the step S1300.
[0061] In the step S1310, a gaze dwell period with a length greater than a preset length and a gaze content type corresponding to the gaze area being a preset type is determined as the target gaze dwell period.
[0062] The preset type can be specified in advance by the wearer according to a content type that needs to be focused on by the wearer.
[0063] Through the step S1310, the target gaze dwell period reflecting the gaze state of the wearer can be more accurately determined in combination with the length and the gaze content type.
[0064] On the basis of the previous embodiment, in an embodiment of the present application, the voice recording method further comprises the steps of adding a label to the independent voice segment as shown in steps S1500 to S1700.
[0065] In step S1500, a keyword in the independent voice segment is determined for any independent voice segment.
[0066] For the above step S1500, the keyword in the independent voice segment can be determined by a keyword extraction algorithm.
[0067] In step S1600, a target label of the independent voice segment is determined according to the gaze content type corresponding to the independent voice segment, the keyword, and a preset mapping relationship.
[0068] The preset mapping relationship is used to reflect the corresponding relationship among the gaze content type, the keyword, and the label, and is stored in the intelligent head-mounted device in advance. In an example, a set of corresponding relationships in the preset mapping relationship is: when the keyword is "growth rate" and "same period", and the gaze content type is "PPT", the label is "core data".
[0069] In the present embodiment, the gaze content type corresponding to the target gaze dwell time period corresponding to the independent voice segment is the gaze content type corresponding to the independent voice segment, that is, the gaze content type in the above step S1310.
[0070] For the above step S1600, the label matched with the gaze content type and the keyword in the preset mapping relationship is determined as the target label.
[0071] In an embodiment of the present application, the keyword in a set of relationships in the preset mapping relationship is usually multiple. For example, a set of relationships in the preset mapping relationship is that the gaze content type is A, the keyword is a and b, and the label is M. Another set of relationships is that the gaze content type is A, the keyword is a and d, and the label is N. On this basis, when the gaze content type determined based on step S1220 is A, and the keyword determined based on the above step S1500 is a, it is not possible to determine whether the target label is N or M. In order to solve this problem, the above step S1600 is specifically implemented by the following steps S1610 and S1611.
[0072] In step S1610, at least one candidate label of the independent voice segment is determined according to the gaze content type corresponding to the independent voice segment, the keyword, and the preset mapping relationship.
[0073] Specifically, the labels matched with the gaze content type corresponding to the independent voice segment and the keyword in the preset mapping relationship are all determined as candidate labels.
[0074] Step S1611, determining the candidate label with the highest priority as the target label.
[0075] In this embodiment, different priorities are set for different labels in the preset mapping relationship, and the label with the highest priority in the candidate labels is determined as the target label. In this regard, for the above example, if the priority of label M is higher than that of label N in the preset mapping relationship, then the target label is determined as M.
[0076] In one example, the preset mapping relationship includes expert opinions, core data, and routine discussions, wherein the priority of the expert opinions is higher than that of the core data, and the priority of the core data is higher than that of the routine discussions.
[0077] Step S1700, adding the target label to the independent voice segment.
[0078] Through the above steps S1500 to S1700, a scenario-based label can also be generated for each independent voice segment. In addition, since each independent voice segment reflects the gaze content that the wearer focuses on, the target label added to the independent voice segment by the above steps S1500 to S1700 matches the real focus of the wearer.
[0079] In one embodiment of the present application, the above step S1100 includes a step of obtaining an environmental audio signal, which is specifically implemented through the following steps S1110 and S1111.
[0080] Step S1110, obtaining an environmental sub-audio signal collected by each microphone in the microphone array.
[0081] In this embodiment, a microphone array is provided in the smart head-mounted device, and the microphone array includes a plurality of microphones. During audio collection, the audio signal collected by each microphone is recorded as an environmental sub-audio signal.
[0082] Step S1111, synthesizing the environmental sub-audio signal collected by each microphone according to the weight of each microphone in the microphone array to obtain the environmental audio signal.
[0083] In this embodiment, the weight of each microphone is the same at the initial time. And the weight of the microphone is reset through the following step S1112.
[0084] Step S1112, in the case where the gaze point target coordinates of the wearer are determined, setting the weight of the environmental sub-audio signal collected by the microphone in the gaze direction corresponding to the gaze point target coordinates to be greater than the weight of the environmental sub-audio signal collected by the microphone outside the gaze direction.
[0085] In an embodiment of the present application, the weight of the environmental sub-audio signal collected by the microphone in the gaze direction corresponding to the gaze point target coordinate is set to a fixed weight value with the maximum weight, and the weight of the environmental sub-audio signal collected by the microphone outside the gaze direction corresponding to the gaze point target coordinate is set to a weight value smaller than the maximum fixed weight value. Alternatively, the above step S1112 is implemented according to the principle that the farther the microphone is from the gaze direction corresponding to the gaze point target coordinate, the smaller the weight of the environmental sub-audio signal collected by the microphone.
[0086] Through the above step S1112, the sound collection focus of the microphone array can be dynamically adjusted according to the gaze direction of the wearer, and the environmental sub-audio signal component in the gaze direction of the wearer in the collected environmental audio signal is enhanced.
[0087] As shown in Figure 2 The present application also provides a voice recording device 200 applied to a smart head-mounted device, which comprises: An acquisition module 210 is configured to acquire eye movement data of a wearer, an environmental image and an environmental audio signal in a visual angle of the wearer during audio signal collection, the environmental audio signal being an audio signal to be recorded, and the eye movement data, the environmental image and the environmental audio signal being associated with a time stamp representing a corresponding collection time point; A first determination module 220 is configured to determine a gaze staying time period of a gaze area of the wearer according to the eye movement data and the environmental image; A second determination module 230 is configured to determine an environmental audio signal acquired during a target gaze staying time period as an independent voice segment, the target gaze staying time period being a gaze staying time period with a time length greater than a preset time length; A marking module 240 is configured to determine a focus weight of an independent voice segment according to a time length of a gaze staying time period of the independent voice segment, and mark the independent voice segment with the corresponding focus weight.
[0088] In an embodiment of the present application, the first determination module 220 is specifically configured to determine a gaze content type and a gaze staying time period of a gaze area of the wearer according to the eye movement data and the environmental image; In the present embodiment, the voice recording device 200 provided by the present application further comprises: A third determination module is configured to determine a gaze staying time period with a time length greater than the preset time length and a gaze content type of a corresponding gaze area being a preset type as a target gaze staying time period.
[0089] In an embodiment of the present application, the marking module 240 is specifically configured to determine a confidence degree of an independent voice segment according to the independent voice segment. determine a focus weight of the independent speech segment according to the gaze dwell time corresponding to the independent speech segment, the confidence of the independent speech segment, and a sum of gaze dwell time corresponding to all the independent speech segments.
[0090] In an embodiment of the present application, the marking module 240 is further configured to: determine a keyword in the independent speech segment for any of the independent speech segments; determine a target label of the independent speech segment according to the gaze content type corresponding to the independent speech segment, the keyword, and a preset mapping relationship, wherein the preset mapping relationship is used to reflect a corresponding relationship among the gaze content type, the keyword, and the label; add the target label to the independent speech segment.
[0091] In an embodiment of the present application, the marking module 240 is specifically configured to determine at least one candidate label of the independent speech segment according to the gaze content type corresponding to the independent speech segment, the keyword, and the preset mapping relationship. determine a target label of the independent speech segment according to the gaze content type corresponding to the independent speech segment, the keyword, and a preset mapping relationship, wherein the preset mapping relationship is used to reflect a corresponding relationship among the gaze content type, the keyword, and the label;
[0092] In an embodiment of the present application, the first determining module 220 is specifically configured to determine an initial gaze point coordinate of the wearer according to the eye movement data. obtain head posture information of the wearer during the collection of the eye movement data; calibrate the initial gaze point coordinate according to the head posture information to obtain a target gaze point coordinate; determine a gaze area of the wearer according to the target gaze point coordinate and the environmental image.
[0093] In an embodiment of the present application, the obtaining module 210 is specifically configured to obtain an environmental sub-audio signal collected by each microphone in the microphone array. combine the environmental sub-audio signal collected by each microphone in the microphone array according to a weight of each microphone to obtain the environmental audio signal. In this embodiment, the voice recording device 200 provided by the present application further comprises: In a case where the target gaze point coordinate of the wearer is determined, the setting module is configured to set a weight of the environmental sub-audio signal collected by the microphone in a direction corresponding to the target gaze point coordinate to be greater than a weight of the environmental sub-audio signal collected by the microphone outside the direction.
[0094] The application also provides a smart head-mounted device, the smart head-mounted device comprising any one of the voice recording devices 200 provided in the device embodiments described above; Alternatively, as shown in FIG. 3, the smart head-mounted device 300 comprises a memory 310 and a processor 320, the memory 310 is configured to store computer instructions, and the processor 320 is configured to call the computer instructions from the memory 310 to execute the voice recording method according to any one of the method embodiments described above. Figure 3
[0095] The application also provides a computer readable storage medium, which stores a computer program, and the computer program, when executed by a processor, implements the voice recording method according to any one of the method embodiments described above.
[0096] The application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present application.
[0097] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0098] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0099] Computer readable program instructions for carrying out operations of the present application can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.
[0100] The computer readable program instructions can also be loaded onto a computing / processing device, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computing / processing device, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computing / processing device, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0101] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data, programs, program modules, e.g., instructions for operation, or digital content stored thereon or therein for a short time or not at all. The computer readable storage medium can also have instructions stored thereon or therein which may
[0102] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0103] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0104] Having described various embodiments of the application, it is to be understood that the above description is meant not to be exhaustive or limited to the various embodiments disclosed. Many modifications and variations are possible in light of the above teachings without departing from the scope and spirit of the described embodiments. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the various embodiments described herein, which will be limited only by the appended claims.
Claims
1. A voice recording method, characterized by, The application is applied to an intelligent head-mounted device, comprising: obtaining eye movement data of a wearer during audio acquisition, an environment image under a visual angle of the wearer and an environment audio signal, the environment audio signal being an audio signal to be recorded, and the eye movement data, the environment image and the environment audio signal being associated with a time stamp representing a corresponding acquisition time; determining a gaze dwell period of a gaze area of the wearer according to the eye movement data and the environment image; determining an environment audio signal acquired during a target gaze dwell period as an independent voice segment, the target gaze dwell period being a gaze dwell period with a length greater than a preset length; for any independent voice segment, determining an attention weight of the independent voice segment according to a corresponding length of the gaze dwell period of the independent voice segment, and marking the independent voice segment with the corresponding attention weight.
2. The method of claim 1, wherein, The method further comprises: determining a gaze content type and a gaze dwell period of the gaze area of the wearer according to the eye movement data and the environment image; before the environment audio signal acquired during the target gaze dwell period is determined as an independent voice segment, the method further comprises: determining a gaze dwell period with a length greater than the preset length and a gaze content type corresponding to the gaze area as a target gaze dwell period.
3. The method of claim 1, wherein, The method further comprises: determining a confidence degree of the independent voice segment according to the independent voice segment; determining the attention weight of the independent voice segment according to a corresponding length of the gaze dwell period of the independent voice segment, the confidence degree of the independent voice segment and a sum of corresponding lengths of gaze dwell periods of all the independent voice segments.
4. The method of claim 2, wherein, After the environment audio signal acquired during the target gaze dwell period is determined as an independent voice segment, the method further comprises: for any independent voice segment, determining a keyword in the independent voice segment; determining a target label of the independent voice segment according to a gaze content type corresponding to the independent voice segment, the keyword and a preset mapping relationship, wherein the preset mapping relationship reflects a corresponding relationship among gaze content types, keywords and labels; adding the target label to the independent voice segment.
5. The method of claim 4, wherein, The method further comprises: determining at least one candidate label of the independent voice segment according to the gaze content type corresponding to the independent voice segment, the keyword and the preset mapping relationship; determining a target label from the at least one candidate label with the highest priority.
6. The method of claim 1, wherein, The method further comprises: determining an initial coordinate of a gaze point of the wearer according to the eye movement data; obtaining head posture information of the wearer during acquisition of the eye movement data; According to the head posture information, the gaze point initial coordinates are calibrated to obtain gaze point target coordinates; According to the gaze point target coordinates and the environment image, a gaze area of the wearer is determined.
7. The method of claim 1, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises:
8. A voice recording device, characterized by The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises:
9. A smart headset, characterized by The method further comprises: The method further comprises:
10. A computer-readable storage medium, characterized in that, The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The
Citation Information
Patent Citations
Eye movement tracking method and device based on depth camera
CN116704572A
Sound source positioning method and system, electronic equipment and storage medium
CN120143052A
Information management apparatus, information management system and program
JP2004259198A
Information processing apparatus, information processing method, and non-transitory computer readable recording medium
US20190355382A1