Data processing method and device and glasses type wearable equipment
By using multiple interval sound acquisition components on glasses wearable devices, the voice data arrival time and pronunciation point distance are detected, and the problem of difficulty in detecting voiceprint attacks is solved, and high-security live voiceprint verification is achieved.
Patent Information
- Application Number
- CN202510474758.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art is difficult to effectively detect and defend against voiceprint attacks, especially in glasses wearable devices such as smart glasses. The security of voiceprint verification is affected by device diversity, instability in features and deep forgery technologies.
By setting a plurality of sound acquisition components with each other spaced apart greater than a preset distance on the glasses wearable device, receiving voice data for user verification, determining the time when the voice data of each character in the preset text content in the voice data reaches each sound acquisition component, and determining whether the user has a preset risk in voiceprint verification based on these times and the distance between the sound acquisition component and the pronunciation point.
It realizes convenient live voiceprint detection, improves the security of voiceprint verification, can effectively resist recording and playback attacks and voice synthesis attacks, and enhances the reliability of user identity verification.
Smart Images

Figure CN119989323A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method, device and eyewear-type wearable device. Background Art
[0002] The payment application scenarios of wearable glasses (such as smart glasses) are increasingly integrated into people's daily lives. Wearable glasses can be equipped with components such as cameras and microphones. Therefore, scanning codes and completing payments in offline scenarios are gradually entering people's lives. Taking smart glasses as an example, there is usually a microphone in smart glasses. Therefore, when executing payment and other services involving user privacy data, identity verification usually uses voiceprints, which has been reached by various smart glasses manufacturers. In this context, the security of voiceprint verification has become particularly important. Common voiceprint attacks include recording playback attacks and speech synthesis attacks. Common detection methods are limited by the diversity of devices, resulting in unstable features, and the increasing depth of fake technology. The difficulty of identification always exists. To this end, it is necessary to provide a better way to verify whether the user has preset risks in voiceprint verification. Summary of the invention
[0003] The purpose of the embodiments of this specification is to provide a better verification method for determining whether a user has a preset risk in voiceprint verification.
[0004] In order to implement the above technical solution, the embodiments of this specification are implemented as follows: A data processing method provided in an embodiment of the present specification is applied to a glasses-type wearable device, wherein the glasses-type wearable device is provided with multiple sound collection components that are mutually spaced greater than a preset distance, and the method comprises: receiving voice data for user verification through each of the multiple sound collection components, wherein the voice data comprises voice data of preset text content; determining, based on the received voice data, the time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component; determining whether the user has a preset risk in voiceprint verification based on the determined time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component and the reference time when the voice data of each character arrives at the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each of the multiple sound collection components.
[0005] A data processing device is provided in an embodiment of the present specification, wherein the data processing device is arranged in a glasses-type wearable device, wherein the glasses-type wearable device is provided with a plurality of sound collection components which are mutually spaced greater than a preset distance, and the device comprises: a data receiving module, receiving voice data for verifying a user through each of the plurality of sound collection components, wherein the voice data comprises voice data of preset text content; a data processing module, determining, based on the received voice data, a time when the voice data of each character in the preset text content in the voice data reaches each sound collection component; a data verification module, determining, based on the determined time when the voice data of each character in the preset text content in the voice data reaches each sound collection component, and a reference time when the voice data of each character reaches the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each of the plurality of sound collection components, whether the user has a preset risk in voiceprint verification.
[0006] An embodiment of the present specification provides a glasses-type wearable device, comprising multiple sound collection components, a glasses-type wearable device body and a control and power supply component, wherein: the control and power supply component is arranged in the glasses-type wearable device body; the multiple sound collection components are respectively arranged in the glasses-type wearable device body and are respectively connected to the control and power supply component; the control and power supply component is configured to control at least two of the multiple sound collection components to receive voice data including preset text content, so as to determine whether the user has a preset risk in voiceprint verification based on the voice data of each character in the preset text content in the voice data and the time when the pronunciation point of each character reaches each of the at least two sound collection components respectively; the mutual interval between any two of the at least two sound collection components is greater than a preset distance, and the preset distance is determined based on the sampling frequency of the sound collection components of the at least two sound collection components.
[0007] The embodiments of the present specification also provide a storage medium, which is used to store computer-executable instructions. When the executable instructions are executed by a processor, the following process is implemented: receiving voice data for user verification through each sound collection component of multiple sound collection components, wherein the voice data includes voice data of preset text content; based on the received voice data, determining the time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component; based on the determined time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component and the reference time when the voice data of each character arrives at the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each sound collection component of the multiple sound collection components, determining whether the user has a preset risk in voiceprint verification.
[0008] The embodiments of the present specification also provide a computer program product, including a computer program, which implements the following process when executed by a processor: receiving voice data for user verification through each of a plurality of voice collection components, the voice data including voice data of preset text content; determining the time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component based on the received voice data; determining whether the user has a preset risk in voiceprint verification based on the determined time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component and the reference time when the voice data of each character arrives at the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each of the plurality of sound collection components. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor. Figure 1 This is a schematic diagram of the structure of a glasses-type wearable device in this manual; Figure 2 This is a schematic diagram of the structure of a glasses-type wearable device with a band-shaped retractable binding unit in this specification; Figure 3 This is a schematic diagram of the structure of a temple-type eyeglasses wearable device in this specification; Figure 4A schematic diagram of a control and power supply component of this specification; Figure 5 This is a schematic diagram of the structure of a glasses-type wearable device in this manual, with the locations of control and power components marked; Figure 6 A schematic diagram of a data processing process of this specification; Figure 7 This is a structural diagram of a data processing system consisting of a glasses-type wearable device and a speaker in this specification; Figure 8 A schematic diagram of a data processing process for determining user risk in this specification; Fig. 9 This is a schematic diagram of the structure of a data processing system composed of a user and a speaker in this specification; Fig.10 This is a schematic diagram of the oral pronunciation position of this manual; Fig.11 A schematic diagram of a data processing process for determining two user risks in this specification; Fig.12 This is a schematic diagram of different pronunciation points and different microphones in this manual; Fig.13 This is a schematic diagram of a data processing process for voice data collection in this specification; Fig.14 This is a schematic diagram of the structure of a data processing device in this specification. DETAILED DESCRIPTION
[0010] The embodiments of this specification provide a data processing method, device and eyewear-type wearable device.
[0011] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0012] The embodiments of this specification provide a mechanism for resisting voiceprint attacks based on letter-level sound source positioning on glasses-type wearable devices. Since voiceprints are often attacked, common voiceprint attacks include recording playback attacks, speech synthesis attacks, etc., and the detection methods of common voiceprint attacks are limited by the diversity of devices, resulting in unstable features, and the increasingly advanced deep fake technology, so the difficulty of identification always exists. For this reason, it is necessary to provide a better verification method for whether a user has a preset risk in voiceprint verification. Based on this, the embodiments of this specification combine the characteristics of glasses-type wearable devices and propose a method for resisting voiceprint attacks based on letter-level sound source positioning on glasses-type wearable devices. Specifically, the microphone array on the glasses-type wearable device is used, combined with the different pronunciation positions (or pronunciation points) of each character in the mouth, and the time when the pronunciation point of each character reaches each microphone is compared to determine whether the user has a preset risk in voiceprint verification, thereby realizing convenient live voiceprint detection. For specific processing, please refer to the specific content in the following embodiments.
[0013] like Figure 1 As shown, the embodiment of this specification provides a glasses-type wearable device, and the glasses-type wearable device is specifically such as smart glasses, virtual reality VR, augmented reality glasses AR, cloud glasses, etc. The glasses-type wearable device includes multiple sound collection components 100, a glasses-type wearable device body 200 and a control and power supply component 300, wherein: The sound collection component 100 can be a microphone, a microphone, etc. The sound collection component 100 can be divided into a variety of different types according to different roles, functions, sound collection principles, etc., such as dynamic microphones, condenser microphones, piezoelectric microphones, air conduction microphones, bone conduction microphones, silicon micro-microphones, etc., which can be set according to actual conditions. The number of multiple sound collection components 100 can be set according to actual conditions, such as 3 sound collection components, 5 sound collection components (such as Figure 1 In addition, the distribution of multiple sound collection components 100 can be set according to actual conditions. For example, multiple sound collection components 100 are set in a specified area, or multiple sound collection components 100 are set in multiple different specified areas.
[0014] The eyeglasses wearable device body 200 may be a device entity composed of materials that encapsulate various components in the eyeglasses wearable device body 200. The above-mentioned encapsulation material may be plastic, ceramic, glass or metal, etc. In this embodiment, the encapsulation material may be plastic. The eyeglasses wearable device body 200 may include units or components that constitute the basic framework of the eyeglasses wearable device. For example, the eyeglasses wearable device body 200 may include a frame 210, a nose pad 220, and a binding component 230 that constrains the frame 210 in front of the user's eyes. The binding component 230 may be a temple 231 in common glasses (such as myopia glasses or sunglasses, etc.), such as Figure 3 As shown, the binding assembly 230 can also be a belt-shaped retractable binding unit 232, such as Figure 2 As shown in FIG. 1 , the frame 210 and the restraint assembly 230 may be connected by a hinge, such as Figure 1 As shown, it can also be connected by bundling, such as Figure 2 The nose pads 220 include two nose pads, which are respectively arranged on both sides of the frame 210. Figure 3 As shown, the eyeglasses-type wearable device body 200 is supported to stabilize and the frame 210 is located in front of the user's eyes.
[0015] The control and power supply component 300 may include a control component 310, a power supply 320 and an integrated circuit component 330. The control component 310 may include a controller, and may also include a cache unit, a storage unit, etc. The power supply 320 may be a battery, or an external power supply connected through a power interface, etc. The integrated circuit component 330 may also include a variety of different functional circuits, such as an analog-to-digital conversion circuit, an amplifier circuit, a short-circuit protection circuit, a voltage stabilization circuit, a current stabilization circuit, etc., and the corresponding functional circuits may be set in the integrated circuit component 300 according to actual needs. The control component 310 and the power supply 320 are interconnected through the integrated circuit component 330, that is, the power supply 320 provides power to each component connected thereto through the integrated circuit component 330, and the control component 310 sends control instructions to the designated component and receives request messages or instruction messages sent by each component through the integrated circuit component 330, which may be set according to actual conditions.
[0016] It should be noted that the control and power supply assembly 300 may have relatively large differences due to different configurations or performances, such as Figure 4As shown, the control component 310 may include one or more processors 311 and memory 312, etc., and the memory 312 may store one or more storage applications or data. Among them, the memory 312 may be a short-term storage or a permanent storage. The application stored in the memory 312 may include one or more modules (not shown in the figure), and each module may include a series of executable instructions for the eyeglasses wearable device. Furthermore, the processor 311 may be configured to communicate with the memory 312 to execute a series of executable instructions in the memory 312 on the eyeglasses wearable device. The control and power supply component 300 may also include one or more power supplies 320, an integrated circuit component 330, one or more wireless network interfaces 340, and one or more input and output interfaces 350.
[0017] In the embodiments of this specification, Figure 1-Figure 4 As shown, the control and power supply component 300 usually includes one or more functional circuits and electronic components. Since these functional circuits contain a large number of small electronic components, such as diodes and / or resistors and / or capacitors, and the number of these components may be large, these electronic components usually need to work in a dry environment. If these electronic components are exposed to the current environment, they are easily damp or wet, which can easily cause damage to the control and power supply component 300. To this end, the control and power supply component 300 can be set in the eyeglass wearable device body 200. Among them, the control and power supply component 300 can be set in any area or position in the eyeglass wearable device body 200, such as the control and power supply component 300 can be encapsulated in the temple 231 in the eyeglass wearable device body 200 (it can be one of the temples 231, or it can be separately installed in two temples 231, etc.), and can also be encapsulated in a designated position of the frame 210 in the eyeglass wearable device body 200, for example, Figure 5 As shown, the control and power supply component 300 is encapsulated in the middle connecting part of the frame 210 in the eyeglass-type wearable device body 200, so that the user does not need to worry about the waterproof and moisture-proof problems of the control and power supply component 300, and it is convenient for the user to use.
[0018] Multiple sound collection components 100 are respectively arranged in the eyeglass wearable device body 200 and are respectively connected to the control and power supply component 300. Through the above-mentioned settings and connection methods, the control and power supply component 300 can send specified control instructions (such as control instructions for turning off sound collection, control instructions for turning on sound collection, etc.) to the multiple sound collection components 100 respectively, and can provide power to each sound collection component 100, so that some or all of the multiple sound collection components 100 can work normally.
[0019] The control and power supply component 300 can control at least two sound collection components 100 among the multiple sound collection components 100 to receive voice data including preset text content for user verification, wherein the mutual interval between any two sound collection components 100 among the at least two sound collection components 100 is greater than a preset distance, wherein the preset distance can be determined based on the sampling frequency of the sound collection components among the at least two sound collection components 100.
[0020] In practical applications, in order to ensure the accuracy of live pronunciation detection, the preset distance corresponding to the interval between any two sound collection components 100 of at least two sound collection components 100 can be determined in the following manner: the setting target of the preset distance can be the sound sampling accuracy of the sound collection component 100 corresponding to the time difference between the pronunciation point of the character reaching one sound collection component 100 and reaching another sound collection component 100 is greater than or equal to a preset multiple (such as 1.2 times or 2 times, etc.). For example, if the sampling frequency of the sound collection component 100 is 48kHz, then the sound sampling accuracy of the sound collection component 100 is about 20 microseconds. The setting target of the preset distance can be the time difference between the pronunciation point of any character reaching at least two sound collection components 100. If the time difference between any two sound collection components 100 in the wearable device 100 is greater than or equal to 1.2 times or 2 times the sound sampling accuracy corresponding to the sound collection component 100 (that is, greater than or equal to 24 microseconds or 40 microseconds), the preset distance can be determined based on the above-set target, and then according to the actual interval between any two sound collection components 100 in the eyeglass wearable device, at least two suitable sound collection components 100 can be selected from the multiple sound collection components 100. For example, if the preset distance is 2 cm, the mutual interval between any two sound collection components 100 in the at least two sound collection components 100 can be 3 cm, 4 cm or 6 cm, etc., so as to perform related processing such as receiving voice data including preset text content.
[0021] Through the above structure of the glasses-type wearable device, the following data processing method can be implemented, that is, the control and power supply component 300 controls at least two of the multiple sound collection components 100 to receive voice data including preset text content for user verification, and can determine whether the user has a preset risk in voiceprint verification based on the voice data of each character in the preset text content in the voice data and the time when the pronunciation point of each character reaches each of the at least two sound collection components 100. Specifically, the glasses-type wearable device controls at least two of the multiple sound collection components 100 to receive voice data including preset text content for user verification through the control and power supply component 300. Then, based on the received voice data, the control and power supply component 300 determines the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component 100. Finally, the control and power supply component 300 determines the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component 100, and the distance between the pronunciation point of each character in the preset text content and each sound collection component 100 in the multiple sound collection components 100 determines the reference time when the voice data of each character reaches the sound collection component 100, so as to determine whether the user has a preset risk in voiceprint verification. For the specific processing process, please refer to the relevant content of the following data processing method.
[0022] The embodiment of the present specification provides a glasses-type wearable device, which is provided with a plurality of sound collection components that are spaced apart from each other by a distance greater than a preset distance, and receives voice data for verifying a user through each of the plurality of sound collection components, wherein the voice data includes voice data of preset text content, and then the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component can be determined, and finally, the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component and the distance between the pronunciation point of each character in the preset text content and each of the plurality of sound collection components can be determined. The base time when the voice data of each character arrives at the sound collection component is used to determine whether the user has a preset risk in voiceprint verification. In this way, by utilizing multiple sound collection components on the eyeglass-type wearable device and combining the different pronunciation positions (or pronunciation points) of each character in the preset text content in the mouth, the base time when the pronunciation point of each character arrives at each sound collection component is compared with the time when each sound collection component receives the voice data of each character in the preset text content, to determine whether the user has a preset risk in voiceprint verification. This realizes convenient live voiceprint detection and achieves the purpose of resisting voiceprint attacks based on letter-level sound source positioning on eyeglass-type wearable devices.
[0023] In actual applications, the eyeglass wearable device body 200 includes two temples 231, a frame 210 and two nose pads 220, wherein at least two sound collection components 100 are two sound collection components 100, and the two sound collection components 100 are evenly distributed on one of the two temples 231, and the second sound collection component of the two sound collection components 100 is adjacent to the side of the frame 210 connected to the temple 231, and the first sound collection component of the two sound collection components 100 is away from the side of the frame 210 connected to the temple 231.
[0024] In addition to the above-mentioned distribution of the two sound collecting components 100, the two sound collecting components 100 can also be respectively arranged on the two temples 231, that is, one of the two sound collecting components 100 is arranged at a certain position on the left temple 231 (which can be any position, such as the middle position of the temple 231 or a position close to one side of the frame 210, etc.), and the other sound collecting component 100 is arranged at a certain position on the right temple 231 (which can be any position, such as the middle position of the temple 231 or a position close to one side of the frame 210, etc.).
[0025] In addition to the above distribution of the two sound collecting components 100, one of the two sound collecting components 100 is set on the frame 210, and the other sound collecting component 100 is set on the frame 210 or the temple 231. Specifically, one of the two sound collecting components 100 is set on the frame 210, and the other sound collecting component 100 is also set on the frame 210, and the two frames 210 can be frames for the same eye of the user or the two frames 210 can be frames for different eyes of the user. In addition, one of the two sound collecting components 100 is set on the frame 210, and the other sound collecting component 100 is also set on the temple 231, and the side of the user's eye corresponding to the frame 210 is the same as the side where the temple 231 is located, such as both are on the left side or both are on the right side, etc., or, the side of the user's eye corresponding to the frame 210 is different from the side where the temple 231 is located, such as one is on the left side and the other is on the right side, etc. Specifically, the user's eye corresponding to the frame 210 is the left eye, and the temple 231 is located on the left side (that is, the left temple 231).
[0026] One of the two sound collecting components 100 is arranged on one of the two nose pads 220, and the other sound collecting component 100 is arranged on the other nose pad 220, the temple 231 or the frame 210, that is, one of the two sound collecting components 100 is arranged on one of the two nose pads 220, and the other sound collecting component 100 is arranged on the other nose pad 220, or one of the two sound collecting components 100 is arranged on one of the two nose pads 220, and the other sound collecting component 100 is arranged on the temple 231, which can be any temple 231 of the two temples 231, or one of the two sound collecting components 100 is arranged on one of the two nose pads 220, and the other sound collecting component 100 is arranged on the frame 210, and the user's eye corresponding to the frame 210 can be the eye on either side.
[0027] In practical applications, at least two sound collection components 100 include a bone conduction sound collection component, which is arranged at a position in contact with the user's body when worn. The bone conduction sound collection component can be arranged inside the nose pad 220 in the eyeglass wearable device body 200 and in contact with the user's nose bridge, and / or the bone conduction sound collection component is arranged on the temple 231 and in contact with the user's cheek and / or ear.
[0028] In actual applications, the bone conduction sound collection component can also be set on the belt-shaped retractable binding unit 232 and contact the head part corresponding to the user's skull, wherein the head part corresponding to the user's skull that contacts the bone conduction sound collection component is covered by the belt-shaped retractable binding unit 232.
[0029] In actual application, the sound collection components 100 other than at least two sound collection components 100 in the multiple sound collection components 100 include a bone conduction sound collection component, and the bone conduction sound collection component is arranged at a position in contact with the user's body in the wearing state. Based on the above structure, the bone conduction sound collection component can receive the bone conduction sound signal corresponding to the voice data, and when the bone conduction sound signal is not received, the control and power supply component 300 is triggered to generate a control instruction related to the preset risk of the user in voiceprint verification.
[0030] In practical applications, as described above, the bone conduction sound collection component (i.e., the sound collection components other than at least two sound collection components in the plurality of sound collection components include the bone conduction sound collection component) is disposed on the inner side of the nose pad 220 in the eyeglass wearable device body 200 and contacts the user's nose bridge, and / or the bone conduction sound collection component is disposed on the temple 231 and contacts the user's cheek and / or ear. Alternatively, the bone conduction sound collection component may also be disposed on the band-shaped retractable binding unit 232 and contact the head portion corresponding to the user's skull, wherein the head portion corresponding to the user's skull contacting the bone conduction sound collection component is covered by the band-shaped retractable binding unit 232.
[0031] The embodiment of the present specification provides a glasses-type wearable device, which is provided with a plurality of sound collection components that are spaced apart from each other by a distance greater than a preset distance, and receives voice data for verifying a user through each of the plurality of sound collection components, wherein the voice data includes voice data of preset text content, and then the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component can be determined, and finally, the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component and the distance between the pronunciation point of each character in the preset text content and each of the plurality of sound collection components can be determined. The base time when the voice data of each character arrives at the sound collection component is used to determine whether the user has a preset risk in voiceprint verification. In this way, by utilizing multiple sound collection components on the eyeglass-type wearable device and combining the different pronunciation positions (or pronunciation points) of each character in the preset text content in the mouth, the base time when the pronunciation point of each character arrives at each sound collection component is compared with the time when each sound collection component receives the voice data of each character in the preset text content, to determine whether the user has a preset risk in voiceprint verification. This realizes convenient live voiceprint detection and achieves the purpose of resisting voiceprint attacks based on letter-level sound source positioning on eyeglass-type wearable devices.
[0032] like Figure 6 As shown, the embodiment of this specification provides a data processing method, and the execution subject of the method can be as described above Figure 1-Figure 5 The eyeglasses-type wearable device shown is specifically a smart glasses, virtual reality VR, augmented reality glasses AR, cloud glasses, etc. The eyeglasses-type wearable device is provided with multiple sound collection components that are spaced apart by a distance greater than a preset distance, and the distance between any two of the multiple sound collection components is greater than the preset distance. For example, the preset distance is 2 cm, and the distance between any two sound collection components can be 3 cm, 4 cm, or 6 cm, etc. It should be noted that the multiple sound collection components that are spaced apart by a distance greater than a preset distance in this embodiment are at least two sound collection components in the above-mentioned embodiment of the eyeglasses-type wearable device. The method may specifically include the following steps: In step S102, voice data for user verification is received by each of the plurality of voice collecting components, where the voice data includes voice data of preset text content.
[0033] The sound collection component can be used to collect sound signals. The sound collection component can include multiple types, such as microphones, pickups with specified functions, etc., which can be set according to actual conditions. Multiple sound collection components can be distributed in different positions of the eyeglass wearable device. For example, three sound collection components can be included, which are distributed in: one sound collection component is set on each of the two temples, and one sound collection component is set on the lower frame; for example, Figure 1 As shown, five sound collection components may be included, which are distributed in: two sound collection components are set on each of the two temples, and one sound collection component is set on the lower frame, etc., which can be set according to actual conditions. The user can be any user who needs to verify whether he or she has a preset risk in voiceprint verification. The preset text content can include multiple types. Specifically, the preset text content can be a word, a sentence, etc. For example, the preset text content can include payment, payment or start verification, etc., which can be set according to actual conditions.
[0034] In implementation, the collection of current voice data through the sound collection component can be triggered in a variety of scenarios or in a variety of different ways. For example, when it is necessary to verify whether a user has a preset risk in biometric verification, a designated button can be clicked (it can be a physical button on the eyeglasses-type wearable device, or a virtual button presented by the eyeglasses-type wearable device, etc.). At this time, the eyeglasses-type wearable device starts the designated sound collection component set by itself, wherein the started sound collection component can be all sound collection components in the eyeglasses-type wearable device, or a specified number of sound collection components (such as selecting 3 or 2 sound collection components from 5 sound collection components, etc.), or a sound collection component at a specified position (such as selecting 2 sound collection components located on the temples from 5 sound collection components, or selecting 1 sound collection component located on the frame and 1 sound collection component on the temples from 5 sound collection components, etc.), etc. Then, the user can read out the preset text content by voice. For example, the user can read out "payment" by voice. At this time, the user's voice data will be captured by each of the multiple sound collection components, so that the eyeglass-type wearable device can receive the voice data for user verification through each of the multiple sound collection components.
[0035] For another example, the user is currently executing a certain service (such as resource transfer service, privacy data reading service, etc., which require user identity verification) through a glasses-type wearable device. During this process, when it is necessary to determine whether the current user is a real user, the glasses-type wearable device starts the specified sound collection component set by itself, and then the user can read the preset text content by voice. At this time, the user's voice data will be captured by each of the multiple sound collection components, so that the glasses-type wearable device can receive the voice data used to verify the user through each of the multiple sound collection components. The above methods are only two optional methods. In actual applications, it can also include a variety of different implementation methods, which can be set according to actual conditions. The embodiments of this specification do not limit this.
[0036] In step S104, based on the received voice data, the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component is determined.
[0037] In implementation, in the process of receiving voice data through each of the multiple sound collection components, the eyeglass wearable device can start timing when the multiple sound collection components are started, or it can also start timing at a preset timing time point, and after each sound collection component receives the voice of each character in the voice data, the reception time of the voice of the character is recorded. For example, the user can read "payment" by voice, and each sound collection component can record the reception time of the voice data of "payment", and then each sound collection component records the reception time of the voice data of "money". In the above manner, the reception time of the voice of each character in the voice data received by each sound collection component can be recorded. After the voice data is obtained through the processing of the above step S102, the reception time of the voice of each character in the voice data received by each sound collection component can be extracted from the voice data, thereby obtaining the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component.
[0038] In step S106, based on the time when the voice data of each character in the preset text content in the determined voice data reaches each sound collection component, and the reference time when the voice data of each character reaches the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components, it is determined whether the user has a preset risk in voiceprint verification.
[0039] Among them, the pronunciation point of a character can be the position of the pronunciation of the character in the oral cavity when the character is read. The positions of the pronunciations of the main pronunciation letters of different characters in the oral cavity may be different. For example, the pronunciation position of the main pronunciation letter G of the character "攻 (pronounced as gong)" in the oral cavity is the soft palate part, and the pronunciation position of the main pronunciation letter J of the character "击 (pronounced as ji)" in the oral cavity is the hard palate part. Another example is that the pronunciation position of the main pronunciation letter F of the character "防 (pronounced as fang)" in the oral cavity is the labiodental part, and the pronunciation position of the main pronunciation letter K of the character "控 (pronounced as kong)" in the oral cavity is the soft palate part.
[0040] In implementation, since the preset text content is preset, therefore, the time (i.e., the reference time) for the voice data of each character to reach the voice acquisition component can be calculated in advance according to the distance between the pronunciation point of each character in the preset text content and each voice acquisition component among multiple voice acquisition components, and this reference time can be stored in the glasses-type wearable device. Since in the process of voiceprint verification, common voiceprint attacks can be achieved through methods such as recording and playback attacks and voice synthesis attacks. In the above methods, usually, the speaker is placed at a certain distance from the user to play the voice data of the preset text content. As Figure 7 shown, if the speaker is far from the user, there is almost no difference in the time for the voice of each character in the voice data of the preset text content to reach each voice acquisition component. If the speaker is close to the user, the time sequence for the voice of each character in the voice data of the preset text content to reach multiple voice acquisition components in sequence is the same. Based on this, after determining the time for the voice data of each character in the preset text content in the voice data to reach each voice acquisition component through the above method, the determined time can be matched with the corresponding reference time. If the determined time does not match the corresponding reference time, it indicates that the user has a preset risk in voiceprint verification. At this time, a prompt message of verification failure can be output, and the user can be refused to perform voiceprint verification based on the current voice data. If the determined time matches the corresponding reference time, it indicates that the user does not have a preset risk in voiceprint verification. At this time, the user can continue to execute the corresponding service or process.
[0041] Or, further calculations can also be made based on the time for the voice data of each character in the preset text content in the determined voice data to reach each voice acquisition component. For example, the difference in the time when different voice acquisition components receive the voice data of the same character can be calculated. At the same time, the difference in the distance from different voice acquisition components to the pronunciation point of the same character can be calculated, and the difference in the corresponding reference time can be calculated through this difference in distance. As described above, as Figure 7As shown, if the speaker is far away from the user, the time for the voice of each character in the voice data of the preset text content to reach each sound collection component is almost the same. If the speaker is close to the user, the time sequence for the voice of each character in the voice data of the preset text content to reach multiple sound collection components in sequence is the same. Based on this, the difference in time obtained above can be matched with the difference in the corresponding reference time. If the determined time does not match the corresponding reference time, it indicates that the user has a preset risk in voiceprint verification. If the determined time matches the corresponding reference time, it indicates that the user does not have a preset risk in voiceprint verification. In addition, in addition to determining whether the user has a preset risk in voiceprint verification in the above manner, it is also possible to determine whether the user has a preset risk in voiceprint verification in a variety of other different ways, which can be set according to actual conditions, and the embodiments of this specification do not limit this.
[0042] The embodiment of the present specification provides a data processing method, which is applied to a glasses-type wearable device, wherein a plurality of sound collection components spaced apart from each other by a distance greater than a preset distance are provided on the glasses-type wearable device, and voice data for verifying a user is received by each of the plurality of sound collection components, wherein the voice data includes voice data of preset text content, and then the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component can be determined, and finally, based on the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component, and the correlation between the pronunciation point of each character in the preset text content and each sound collection component of the plurality of sound collection components, the voice data of the glasses-type wearable device is received by each of the plurality of sound collection components. The distance between the sound collection components is used to determine the benchmark time for the voice data of each character to arrive at the sound collection component, so as to determine whether the user has a preset risk in voiceprint verification. In this way, by using multiple sound collection components on the eyeglass-type wearable device and combining the different pronunciation positions (or pronunciation points) of each character in the preset text content in the mouth, the benchmark time for the pronunciation point of each character to arrive at each sound collection component is compared with the time when each sound collection component receives the voice data of each character in the preset text content, so as to judge whether the user has a preset risk in voiceprint verification, thereby realizing convenient live voiceprint detection and realizing the purpose of resisting voiceprint attacks based on letter-level sound source positioning on eyeglass-type wearable devices.
[0043] In practical applications, the specific processing methods of the above step S106 can be various. The following is an optional processing method, which can specifically include the processing of the following steps S10602 to S10608. Based on this, in the above Figure 6 Based on this, the method specifically includes the following steps: Figure 8 shown.
[0044] In step S10602, based on the time when the voice data of each character in the preset text content in the determined voice data reaches each sound collection component, the time difference between the voice data of each character reaching any two sound collection components among the multiple sound collection components is determined.
[0045] In implementation, the above-mentioned eyeglass-type wearable device starts timing when multiple sound collection components are started or starts timing at a preset timing time point, and then calculates the time when the voice data of each character reaches each sound collection component. Since the time point for starting timing may be inaccurate for the obtained time, the calculation can be based on the time difference between the voice data of each character reaching any two of the multiple sound collection components, thereby eliminating the impact of the inaccuracy of the time point for starting timing on the obtained time. Based on this, the time difference between the voice data of each character reaching any two of the multiple sound collection components can be calculated based on the time when the voice data of each character in the preset text content in the determined voice data reaches each sound collection component.
[0046] In step S10604, based on the time difference between the voice data of each character reaching any two of the multiple sound collection components, a time difference sequence corresponding to the characters in the preset text content is constructed.
[0047] In implementation, for example, the preset text content includes two characters, and the multiple sound collection components include two sound collection components. The time difference between the voice data of the first character reaching the two sound collection components is A, and the time difference between the voice data of the second character reaching the two sound collection components is B. Then the time difference sequence corresponding to the characters in the preset text content is (A, B).
[0048] In step S10606, a reference time difference sequence corresponding to characters in the preset text content is obtained, and the reference time difference sequence corresponding to characters in the preset text content is constructed based on the reference time of voice data of each character arriving at the sound collection component determined based on the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components.
[0049] In implementation, the reference time difference sequence corresponding to the characters in the preset text content can be calculated by the same processing method as above. Specifically, based on the distance between the pronunciation point of each character in the preset text content and each of the multiple sound collection components, the reference time for the voice data of each character to reach each of the multiple sound collection components can be calculated, and then the reference time difference for the voice data of each character to reach any two of the multiple sound collection components can be calculated, so as to obtain the reference time difference sequence corresponding to the characters in the preset text content.
[0050] In step S10608, if the time difference sequence corresponding to the characters in the preset text content matches the reference time difference sequence corresponding to the characters in the preset text content, it is determined that the user does not have a preset risk in voiceprint verification.
[0051] In implementation, during the process of voiceprint verification, common voiceprint attacks can be achieved through methods such as recording and playback attacks and voice synthesis attacks. In the above methods, usually the speaker is placed at a certain distance from the user to play the voice data of the preset text content, such as Fig. 9 As shown, the multiple sound collection components are two microphones (i.e., microphone 1 and microphone 2), and the preset text content is "payment". If the speaker is far from the user, the time for the voice of each character in the voice data of "payment" to reach each microphone is almost the same, and the time differences obtained at this time are all 0. If the speaker is close to the user, the time sequence for the voice of each character in the voice data of "payment" to reach the multiple sound collection components in turn is the same, that is, the time sequence for the voice data of "fu" to reach microphone 1 and microphone 2 in turn is the same as the time sequence for the voice data of "kuan" to reach microphone 1 and microphone 2 in turn. Based on this, it can be determined whether the user has a preset risk in voiceprint verification by matching the time difference sequence corresponding to the characters in the preset text content with the reference time difference sequence corresponding to the characters in the preset text content.
[0052] In practical applications, the specific construction method of the reference time difference sequence corresponding to the characters in the above preset text content can be various. Hereinafter, an optional processing method is provided, which specifically may include the processing from step A02 to step A10.
[0053] In step A02, determine the pronunciation point of each character in the oral cavity for the preset text content.
[0054] In practical applications, the pronunciation points of the main pronunciation letters of different characters in the oral cavity may be different. For example, such as Fig.10As shown, [b], [p], [m], etc. are bilabial consonants, and the pronunciation points in the oral cavity are located at the bilabial part. [f], [v], etc. are labiodental consonants, and the pronunciation points in the oral cavity are located at the labiodental part. [g], [k], , etc. are velar consonants, and the pronunciation points in the oral cavity are located at the velar part. , [tʃ], [j], etc. are palatal consonants, and the pronunciation points in the oral cavity are located at the palatal part. , [ʃ], etc. are post-alveolar consonants, and the pronunciation points in the oral cavity are located at the post-alveolar part. [d], [t], [n], [l], [z], [s], [r], etc. are alveolar consonants, and the pronunciation points in the oral cavity are located at the alveolar part. [θ], [ð], etc. are interdental consonants, and the pronunciation points in the oral cavity are located between the teeth. [h], etc. are glottal consonants, and the pronunciation points in the oral cavity are located at the throat part. For example, if the preset text content is "付款 (pronounced as Fu Kuan)", then based on Fig.10 , the main pronunciation letter "F" of the character "付" is a labiodental consonant, and the pronunciation point in the oral cavity is located at the labiodental part. The main pronunciation letter "K" of the character "款" is a velar consonant, and the pronunciation point in the oral cavity is located at the velar part. Another example, if the preset text content is "支付 (pronounced as Zhi Fu)", then based on Fig.10 , the main pronunciation letter "ZH" (same as ) of the character "支" is a palatal consonant, and the pronunciation point in the oral cavity is located at the palatal part. The main pronunciation letter "F" of the character "付" is a labiodental consonant, and the pronunciation point in the oral cavity is located at the labiodental part, etc. It can be set according to the actual situation.
[0055] In step A04, obtain the distance between the pronunciation point of each character in the preset text content and each of the multiple sound collection components.
[0056] In implementation, for example, if the preset text content is "付款", then the distance between the pronunciation point (i.e., the labiodental part) of the main pronunciation letter "F" of the character "付" in the oral cavity and each of the multiple sound collection components can be determined. Correspondingly, the distance between the pronunciation point (i.e., the velar part) of the main pronunciation letter "K" of the character "款" in the oral cavity and each of the multiple sound collection components can be determined. Based on the above method, the distance between the pronunciation point of each character in the preset text content and each of the multiple sound collection components can be determined.
[0057] In step A06, based on the distance between the pronunciation point of each character in the preset text content and each of the multiple sound collection components, determine the distance difference between the pronunciation point of each character in the preset text content and any two of the multiple sound collection components.
[0058] In implementation, for example, if the preset text content is "payment", the difference can be calculated between the distance from the pronunciation point of the main pronunciation letter "F" of the character "pay" (i.e., the lip - tooth position) in the oral cavity to the first sound collection component among multiple sound collection components and the distance from the pronunciation point of the main pronunciation letter of this character in the oral cavity to the second sound collection component among multiple sound collection components. Through the above - mentioned method, the distance differences between the pronunciation point of the main pronunciation letter of this character in the oral cavity and any two sound collection components among multiple sound collection components can be calculated. Then, in the same way, the distance differences between the pronunciation point of the main pronunciation letter of the character "ment" in the oral cavity and any two sound collection components among multiple sound collection components can be calculated, so that the distance differences between the pronunciation points of each character in the preset text content and any two sound collection components among multiple sound collection components can be obtained.
[0059] In step A08, based on the distance differences between the pronunciation points of each character in the preset text content and any two sound collection components among multiple sound collection components, and the speed of sound, determine the reference time differences for the voice data of each character in the preset text content to reach any two sound collection components among multiple sound collection components.
[0060] In implementation, the distances between the pronunciation points of each character in the preset text content and each sound collection component among multiple sound collection components can be measured in advance. Then, based on the measured distances, the distance differences between the pronunciation points of each character and any two sound collection components among multiple sound collection components can be calculated. Then, the calculated distance differences can be divided by the speed of sound (i.e., the propagation speed of sound, 340 m / s) to obtain the corresponding reference time differences, so that the reference time differences for the voice data of each character in the preset text content to reach any two sound collection components among multiple sound collection components can be obtained.
[0061] In step A10, based on the reference time differences for the voice data of each character in the preset text content to reach any two sound collection components among multiple sound collection components, construct a reference time - difference sequence corresponding to the characters in the preset text content.
[0062] In implementation, the determined reference time differences can be sorted using the same sorting rule as the arrangement rule of the above - mentioned time - difference sequence, and the sorted reference time differences can be used as the reference time - difference sequence corresponding to the characters in the preset text content.
[0063] In practical applications, the specific processing method of the above - mentioned step S106 can be implemented not only through the above - mentioned method, but also through various methods. Hereinafter, an optional processing method is provided. Specifically, it can include the processing from step S10610 to step S10616. Based on this, in the above Figure 8Based on this, the method specifically includes the following steps: Fig.11 shown.
[0064] In step S10610, based on the time at which the voice data of each character in the preset text content in the determined voice data arrives at each sound collection component, a sub-sequence of the order in which the voice data of each character arrives at the multiple sound collection components is determined.
[0065] In implementation, in order to simplify the processing procedure or further verify the processing results of step S10602 to step S10608, the different orders in which the sound arrives at the sound collection component can be used to determine whether the user has a preset risk in voiceprint verification. Specifically, the time when the voice data of each character in the preset text content in the determined voice data arrives at each sound collection component can be compared, and based on the comparison result, the order in which the voice data of each character arrives at multiple sound collection components can be determined, and a corresponding sub-sequence sequence can be constructed based on multiple sound collection components with a sequence. For example, the preset text content includes two characters, and the multiple sound collection components are two sound collection components, which are respectively recorded as S1 and S2. The first character reaches S1 first and then reaches S2. For the first character, the sub-sequence sequence is (S1, S2). The second character reaches S2 first and then reaches S1. For the second character, the sub-sequence sequence is (S2, S1). Alternatively, a certain sequence sequence can be used as a benchmark, such as taking the sequence sequence (S1, S2) (the sequence sequence indicates that it reaches S1 first and then reaches S2) as a benchmark, and comparing other results with the benchmark to determine the corresponding The corresponding sub-sequence sequence, based on the above example, the first character reaches S1 first and then reaches S2, then for the first character, the result is the same as the above benchmark, at this time, for the first character, its sub-sequence sequence can be recorded as 1, the second character reaches S2 first and then reaches S1, then for the second character, its result is opposite to the above benchmark, at this time, for the second character, its sub-sequence sequence can be recorded as -1, in addition, the sub-sequence sequence of the order in which the voice data of each character arrives at multiple sound collection components can be determined in a variety of different ways, which can be set according to actual conditions.
[0066] In step S10612, a sequence corresponding to the characters in the preset text content is constructed based on the sub-sequence of the order in which the voice data of each character arrives at the multiple sound collection components.
[0067] In implementation, based on the above example, for the first character, the sub-sequence sequence is (S1, S2), and for the second character, the sub-sequence sequence is (S2, S1), then the sequence sequence corresponding to the characters in the preset text content can be (S1, S2; S2, S1), or, for the first character, its sub-sequence sequence is recorded as 1, and for the second character, its sub-sequence sequence is recorded as -1, then the sequence sequence corresponding to the characters in the preset text content can be (1, -1), etc.
[0068] In step S10614, a reference order sequence corresponding to characters in the preset text content is obtained, and the reference order sequence corresponding to characters in the preset text content is constructed based on a reference time for voice data of each character to arrive at a sound collection component determined by a distance between a pronunciation point of each character in the preset text content and each sound collection component in a plurality of sound collection components.
[0069] The method of constructing the reference order sequence corresponding to the characters in the above-mentioned preset text content can be similar to the processing of the above-mentioned step S10610 and step S10612, which will not be repeated here.
[0070] In step S10616, if the order sequence corresponding to the characters in the preset text content matches the reference order sequence corresponding to the characters in the preset text content, it is determined that the user does not have a preset risk in voiceprint verification.
[0071] In practice, during the voiceprint verification process, common voiceprint attacks can be achieved through recording playback attacks, speech synthesis attacks, etc. In the above methods, the speaker is usually placed at a certain distance from the user to play voice data with preset text content, such as Fig. 9As shown, the multiple voice acquisition components are two microphones (i.e., microphone 1 and microphone 2), and the preset text content is "payment". If the speaker is far away from the user, there is little difference in the time for the voice of each character in the voice data of "payment" to reach each microphone. At this time, the order sequence corresponding to the characters in the preset text content cannot be determined. However, if the speaker is close to the user, the order sequence in which the voice of each character in the voice data of "payment" reaches the multiple voice acquisition components is the same. That is, the order sequence in which the voice of each character in the voice data of "fu" reaches the multiple voice acquisition components is microphone 2 - microphone 1, and the order sequence in which the voice of each character in the voice data of "kuan" reaches the multiple voice acquisition components is also microphone 2 - microphone 1. Based on this, the elements in the order sequence corresponding to the characters in the preset text content can be compared with the elements in the reference order sequence corresponding to the characters in the preset text content. If it is determined after comparison that the order sequence corresponding to the characters in the preset text content is the same as the reference order sequence corresponding to the characters in the preset text content, it is determined that the order sequence corresponding to the characters in the preset text content matches the reference order sequence corresponding to the characters in the preset text content. Or, if the error between the elements of the two is within the preset error range, it can be determined that the order sequence corresponding to the characters in the preset text content matches the reference order sequence corresponding to the characters in the preset text content. If the error between the elements of the two is not within the preset error range, it can be determined that the user has a preset risk in voiceprint verification.
[0072] It should be noted that the processing of steps S10610 to S10616 can be implemented in parallel with the processing of steps S10602 to S10608. In practical applications, the processing of steps S10610 to S10616 and the processing of steps S10602 to S10608 can also be used as a further verification process for each other. That is, the processing of steps S10610 to S10616 and the processing of steps S10602 to S10608 can be connected in series as an overall verification method, which can be specifically set according to the actual situation.
[0073] In practical applications, the multiple voice acquisition components are two voice acquisition components, which are evenly distributed on the temple. The second voice acquisition component among the two voice acquisition components is adjacent to the side of the frame connected to the temple, and the first voice acquisition component among the two voice acquisition components is far from the side of the frame connected to the temple.
[0074] In implementation, such as Fig.12As shown, taking the preset text content "payment" as an example, the distance between the labiodental sound F (the pronunciation point is located in the labiodental part of the mouth) and the velar sound K (the pronunciation point is located in the soft palate part of the mouth) is generally 10-14cm (centimeter), with an average distance of 12cm. The distance between the labiodental sound F (the pronunciation point is located in the labiodental part of the mouth) and the center of the eyeglass wearable device is generally 6-8cm, with an average distance of 7cm. Based on the above average distance, the temples can be divided into three parts, and two sound collection groups are placed at the two middle positions (the two positions are about 4cm apart, and one position is about 4cm away from the frame) respectively. The sound collection components are microphones, and from left to right they are microphone 1 and microphone 2. Based on the above-mentioned measured distance, the distances between the pronunciation points of each character in the preset text content and microphone 1 and microphone 2 can be measured respectively, and then the distance difference between the pronunciation points of each character in the preset text content and microphone 1 and microphone 2 can be calculated. The distance difference and the speed of sound can be used to calculate the reference time difference for the voice data of each character in the preset text content to reach microphone 1 and microphone 2, so that a reference time difference sequence corresponding to the characters in the preset text content can be constructed.
[0075] In addition, based on the measured distances between the pronunciation points of each character in the preset text content and microphone 1 and microphone 2, respectively, the reference time for the voice data of each character to arrive at microphone 1 and microphone 2 can be calculated, and then the order in which the voice data of each character arrives at microphone 1 and microphone 2 can be determined, thereby obtaining a reference order sequence corresponding to the characters in the preset text content, such as (microphone 2, microphone 1; microphone 1, microphone 2) or taking the order sequence of (microphone 2, microphone 1) (this order sequence indicates that it arrives at microphone 2 first and then arrives at microphone 1) as a reference, the reference order sequence corresponding to the characters in the preset text content can be (1, -1).
[0076] In practical applications, the specific processing methods of the above step S102 can be various. The following is an optional processing method, which can specifically include the processing of the following steps S1022 to S1026. Based on this, in the above Fig.11 Based on this, the method specifically includes the following steps: Fig.13 shown.
[0077] In step S1022, a resource transfer instruction based on voiceprint triggered by a user is received.
[0078] In implementation, the triggering execution methods of the processing of the above steps S102 to S106 may include multiple methods. For example, when performing voiceprint-based resource transfer services (such as voiceprint-based payment services, voiceprint-based transfer services, and other voiceprint-based online transaction services) through eyeglasses-type wearable devices, the user can click on a pre-set voiceprint-based resource transfer button or hyperlink, etc. At this time, the eyeglasses-type wearable device can generate a voiceprint-based resource transfer instruction, and the eyeglasses-type wearable device can receive the voiceprint-based resource transfer instruction.
[0079] In step S1024, a resource information sheet corresponding to the resource transfer instruction is obtained and displayed, and confirmation prompt information corresponding to the resource information sheet is displayed. The confirmation prompt information is used to prompt the user to read the preset text content by voice.
[0080] The resource information sheet may be a list of information related to resource transfer, such as orders, contracts, etc.
[0081] In step S1026, voice data for verifying the user is received by each of the multiple sound collecting components.
[0082] Based on the above processing, if it is determined that the user does not have the preset risk in voiceprint verification, the user can be allowed to perform resource transfer processing through voiceprint. Specifically, the user can input the corresponding voiceprint data, and the glasses-type wearable device can match the voiceprint data input by the user with the pre-stored reference voiceprint data of the user. If it matches, the resource transfer processing can be performed. If it does not match, the user can be reminded that the resource transfer failed. At this time, the user can re-initiate the resource transfer processing. If it is determined that the user has the preset risk in voiceprint verification, the user can be denied to perform subsequent resource transfer processing based on voiceprint, that is, the user is not allowed to enter the corresponding voiceprint data to perform resource transfer processing.
[0083] In actual applications, after determining whether the user has a preset risk in voiceprint verification, corresponding business processing can also be performed through the processing methods of the following steps B2 and B4.
[0084] In step B2, if it is determined that the user does not have a preset risk in voiceprint verification, a resource transfer process based on the voiceprint is performed based on the voice data.
[0085] In implementation, in order to simplify the processing and improve the security of resource transfer, the above voice data can be directly used as voiceprint data for resource transfer. At this time, the preset text content can be the content of prompt text information to guide the user to transfer resources, such as "Please say "payment" to make a payment". If it is determined that the user does not have the preset risk in voiceprint verification, the above voice data can be directly matched with the pre-stored baseline voiceprint data of the user. If it matches, the resource transfer process can be executed. If it does not match, the user can be reminded that the resource transfer failed. At this time, the user can re-initiate the resource transfer process.
[0086] In step B4, if it is determined that the user has a preset risk in voiceprint verification, the resource transfer process based on the voiceprint is rejected based on the voice data.
[0087] In implementation, if it is determined that the user has a preset risk in voiceprint verification, the user is denied from directly using the above voice data for resource transfer processing, that is, the user is denied from directly using the above voice data to match the pre-stored baseline voiceprint data of the user.
[0088] In practical applications, the preset distance between any two sound collection components among the multiple sound collection components is determined based on the sampling frequency of the sound collection components.
[0089] In implementation, in order to ensure the accuracy of live pronunciation detection, the preset distance corresponding to the interval between any two sound collection components among the multiple sound collection components can be determined in the following manner: the setting target of the preset distance can be that the time difference between the voice of the character reaching one sound collection component and reaching another sound collection component is greater than or equal to 1.5 times or 2 times the sound sampling accuracy corresponding to the sound collection component, for example, based on Fig.12In the example shown, if the sampling frequency of the microphone is 48kHz, the sound sampling accuracy of the microphone is about 20 microseconds. The setting target of the preset distance can be that the time difference between the pronunciation point of the labiodental sound F arriving at microphone 2 and that arriving at microphone 1 is greater than or equal to 1.5 times or 2 times the sound sampling accuracy of the sound collection component (that is, greater than or equal to 30 microseconds or 40 microseconds), and the time difference between the pronunciation point of the velar sound K arriving at microphone 1 and that arriving at microphone 2 is greater than or equal to 30 microseconds or 40 microseconds; if the sampling frequency of the microphone is 128kHz, the sound sampling accuracy of the microphone is about 20 microseconds. The sampling accuracy is about 7.8 microseconds. The setting target of the preset distance can be that the time difference between the pronunciation point of the labiodental sound F arriving at microphone 2 and that arriving at microphone 1 is greater than or equal to 11.7 microseconds or 15.6 microseconds, the time difference between the pronunciation point of the velar sound K arriving at microphone 1 and that arriving at microphone 2 is greater than or equal to 11.7 microseconds or 15.6 microseconds, etc. The preset distance can be determined based on the above-set target, and then according to the actual interval between any two sound collection components in the eyeglass-type wearable device, appropriate multiple sound collections can be selected to execute the above-mentioned step S102 and other related processing.
[0090] The embodiment of the present specification provides a data processing method, which is applied to a glasses-type wearable device, wherein a plurality of sound collection components spaced apart from each other by a distance greater than a preset distance are provided on the glasses-type wearable device, and voice data for verifying a user is received by each of the plurality of sound collection components, wherein the voice data includes voice data of preset text content, and then the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component can be determined, and finally, based on the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component, and the correlation between the pronunciation point of each character in the preset text content and each sound collection component of the plurality of sound collection components, the voice data of the glasses-type wearable device is received by each of the plurality of sound collection components. The distance between the sound collection components is used to determine the benchmark time for the voice data of each character to arrive at the sound collection component, so as to determine whether the user has a preset risk in voiceprint verification. In this way, by using multiple sound collection components on the eyeglass-type wearable device and combining the different pronunciation positions (or pronunciation points) of each character in the preset text content in the mouth, the benchmark time for the pronunciation point of each character to arrive at each sound collection component is compared with the time when each sound collection component receives the voice data of each character in the preset text content, so as to judge whether the user has a preset risk in voiceprint verification, thereby realizing convenient live voiceprint detection and realizing the purpose of resisting voiceprint attacks based on letter-level sound source positioning on eyeglass-type wearable devices.
[0091] The above is a data processing method provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, which is arranged in a glasses-type wearable device. The glasses-type wearable device is provided with a plurality of sound collection components which are spaced apart from each other by a distance greater than a preset distance, such as Fig.14 shown.
[0092] The data processing device comprises: a data receiving module 1401, a data processing module 1402 and a data verification module 1403, wherein: The data receiving module 1401 receives voice data for user verification through each of the multiple voice collecting components, wherein the voice data includes voice data of preset text content; The data processing module 1402 determines, based on the received voice data, the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component; The data verification module 1403 determines whether the user has a preset risk in voiceprint verification based on the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component, and the benchmark time when the voice data of each character reaches the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components.
[0093] In the embodiment of this specification, the data verification module 1403 includes: a time difference determining unit, which determines a time difference between the voice data of each character arriving at any two of the plurality of sound collecting components based on the determined time when the voice data of each character in the preset text content in the voice data arrives at each sound collecting component; A time difference sequence construction unit, which constructs a time difference sequence corresponding to a character in a preset text content based on a time difference between voice data of each character reaching any two of the plurality of sound collection components; A reference time difference sequence acquisition unit is used to acquire a reference time difference sequence corresponding to characters in a preset text content, wherein the reference time difference sequence corresponding to characters in the preset text content is constructed based on a reference time at which voice data of each character reaches a sound collection component, determined based on a distance between a pronunciation point of each character in the preset text content and each sound collection component in the plurality of sound collection components; The first data verification unit determines that the user does not have a preset risk in voiceprint verification if the time difference sequence corresponding to the characters in the preset text content matches the reference time difference sequence corresponding to the characters in the preset text content.
[0094] In the embodiment of this specification, the device further includes: A pronunciation point determination module, which determines the pronunciation point of each character in the preset text content in the oral cavity; A distance acquisition module, which acquires the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components; a distance difference determination module, which determines the distance difference between the pronunciation point of each character in the preset text content and any two sound collection components in the multiple sound collection components based on the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components; a reference time difference determination module, which determines a reference time difference between the voice data of each character in the preset text content and reaching any two of the multiple sound collection components based on the distance difference between the pronunciation point of each character in the preset text content and any two of the multiple sound collection components, and the sound speed; The reference time difference sequence determination module constructs a reference time difference sequence corresponding to the characters in the preset text content based on the reference time difference between the voice data of each character in the preset text content reaching any two of the multiple sound collection components.
[0095] In the embodiment of this specification, the data verification module 1403 includes: A sequence determination unit, based on the determined time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component, determines a sub-sequence of the order in which the voice data of each character arrives at the plurality of sound collection components; A sequence construction unit, which constructs a sequence corresponding to the characters in the preset text content based on a subsequence of the order in which the voice data of each character arrives at the plurality of sound collection components; A reference sequence acquisition unit is configured to acquire a reference sequence corresponding to characters in a preset text content, wherein the reference sequence corresponding to characters in the preset text content is constructed based on a reference time at which voice data of each character reaches a sound collection component, determined based on a distance between a pronunciation point of each character in the preset text content and each sound collection component in the plurality of sound collection components; The second data verification unit determines that the user does not have a preset risk in voiceprint verification if the sequence corresponding to the characters in the preset text content matches the reference sequence corresponding to the characters in the preset text content.
[0096] In the embodiment of the present specification, the multiple sound collecting components are two sound collecting components, and the two sound collecting components are evenly distributed on the temples. The second sound collecting component of the two sound collecting components is adjacent to the side of the frame connected to the temple, and the first sound collecting component of the two sound collecting components is away from the side of the frame connected to the temple.
[0097] In the embodiment of this specification, the data receiving module 1401 includes: An instruction receiving unit, receiving a voiceprint-based resource transfer instruction triggered by the user; An information display unit, which obtains and displays a resource information sheet corresponding to the resource transfer instruction, and displays confirmation prompt information corresponding to the resource information sheet, wherein the confirmation prompt information is used to prompt the user to read the preset text content by voice; The data receiving unit receives voice data for verifying the user through each of the multiple sound collecting components.
[0098] In the embodiment of this specification, the device further includes: A resource transfer module, which performs a resource transfer process based on the voiceprint based on the voice data if it is determined that the user does not have a preset risk in the voiceprint verification; The service rejection module refuses to perform voiceprint-based resource transfer processing based on the voice data if it is determined that the user has a preset risk in voiceprint verification.
[0099] In the embodiment of the present specification, the preset distance between any two sound collection components among the multiple sound collection components is determined based on the sampling frequency of the sound collection components.
[0100] The embodiment of the present specification provides a data processing device, which is arranged in a glasses-type wearable device, wherein a plurality of sound collection components which are spaced apart from each other by a distance greater than a preset distance are arranged on the glasses-type wearable device, and voice data for verifying a user is received by each of the plurality of sound collection components, wherein the voice data includes voice data of preset text content, and then the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component can be determined, and finally, based on the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component, and the correlation between the pronunciation point of each character in the preset text content and each sound collection component of the plurality of sound collection components, the voice data of the user can be processed. The distance between the sound collection components is used to determine the benchmark time for the voice data of each character to arrive at the sound collection component, so as to determine whether the user has a preset risk in voiceprint verification. In this way, by using multiple sound collection components on the eyeglass-type wearable device and combining the different pronunciation positions (or pronunciation points) of each character in the preset text content in the mouth, the benchmark time for the pronunciation point of each character to arrive at each sound collection component is compared with the time when each sound collection component receives the voice data of each character in the preset text content, so as to judge whether the user has a preset risk in voiceprint verification, thereby realizing convenient live voiceprint detection and realizing the purpose of resisting voiceprint attacks based on letter-level sound source positioning on eyeglass-type wearable devices.
[0101] Furthermore, based on the above Figures 6 to 13 One or more embodiments of this specification further provide a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc. When the computer executable instruction information stored in the storage medium is executed by the processor, the following process can be implemented: Voice data for user verification is received by each of a plurality of voice collection components, wherein the voice data includes voice data of preset text content; based on the received voice data, a time when the voice data of each character in the preset text content in the voice data reaches each sound collection component is determined; based on the determined time when the voice data of each character in the preset text content in the voice data reaches each sound collection component and a reference time when the voice data of each character reaches the sound collection component determined by a distance between a pronunciation point of each character in the preset text content and each of the plurality of sound collection components, it is determined whether the user has a preset risk in voiceprint verification.
[0102] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0103] The embodiment of the present specification provides a storage medium, which receives voice data for user verification through each of a plurality of voice collection components, wherein the voice data includes voice data of preset text content, and then determines the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component. Finally, it can be determined whether the user has a preset risk in voiceprint verification based on the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component and the reference time when the voice data of each character reaches the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each of the plurality of sound collection components. In this way, by using the plurality of sound collection components on the glasses-type wearable device, combined with the different pronunciation positions (or pronunciation points) of each character in the preset text content in the oral cavity, the reference time when the pronunciation point of each character reaches each sound collection component is compared with the time when the voice data of each character in the preset text content received by each sound collection component, to determine whether the user has a preset risk in voiceprint verification, thereby realizing convenient live voiceprint detection, and realizing the purpose of resisting voiceprint attacks based on letter-level sound source positioning on the glasses-type wearable device.
[0104] Furthermore, based on the above Figures 6 to 13 One or more embodiments of the present specification further provide a computer program product, including a computer program. When the computer program in the computer program product is executed by a processor, the following process can be implemented: Voice data for user verification is received by each of a plurality of voice collection components, wherein the voice data includes voice data of preset text content; based on the received voice data, a time when the voice data of each character in the preset text content in the voice data reaches each sound collection component is determined; based on the determined time when the voice data of each character in the preset text content in the voice data reaches each sound collection component and a reference time when the voice data of each character reaches the sound collection component determined by a distance between a pronunciation point of each character in the preset text content and each of the plurality of sound collection components, it is determined whether the user has a preset risk in voiceprint verification.
[0105] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0106] The embodiment of the present specification provides a computer program product, which receives voice data for user verification through each of a plurality of voice collection components, wherein the voice data includes voice data of preset text content, and then determines the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component. Finally, it can be determined whether the user has a preset risk in voiceprint verification based on the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component and the reference time when the voice data of each character reaches the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each of the plurality of sound collection components. In this way, by using the plurality of sound collection components on the eyeglass-type wearable device, combined with the different pronunciation positions (or pronunciation points) of each character in the preset text content in the oral cavity, the reference time when the pronunciation point of each character reaches each sound collection component is compared with the time when the voice data of each character in the preset text content received by each sound collection component, to determine whether the user has a preset risk in voiceprint verification, thereby realizing convenient live voiceprint detection, and realizing the purpose of resisting voiceprint attacks based on letter-level sound source positioning on the eyeglass-type wearable device.
[0107] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0108] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0109] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0110] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0111] For the convenience of description, the above devices are described in terms of functions and are described separately in various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0112] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0113] The embodiments of this specification are described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable fraud case serial and parallel device to produce a machine, so that the instructions executed by the processor of the computer or other programmable fraud case serial and parallel device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0114] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable fraud case serial and parallel device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0115] These computer program instructions may also be loaded onto a computer or other programmable device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0116] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0117] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0118] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0119] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0120] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0121] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0122] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0123] The above description is only an embodiment of this specification and is not intended to limit this document. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this specification.
Claims
1. A data processing method, applied to a glasses-type wearable device, wherein the glasses-type wearable device is provided with a plurality of sound collection components which are mutually spaced greater than a preset distance, the method comprising: Receiving voice data for verifying the user through each of the plurality of voice collecting components, wherein the voice data includes voice data of preset text content; Based on the received voice data, determining the time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component; Based on the determined time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component, and the reference time when the voice data of each character arrives at the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components, it is determined whether the user has a preset risk in voiceprint verification.
2. The method according to claim 1, wherein the determining of the time at which the voice data of each character in the preset text content in the voice data reaches each sound collection component and the reference time at which the voice data of each character reaches the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components, comprises: Based on the determined time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component, determine the time difference between the voice data of each character arriving at any two sound collection components among the multiple sound collection components; Based on the time difference between the voice data of each character reaching any two of the multiple sound collection components, construct a time difference sequence corresponding to the characters in the preset text content; Obtaining a reference time difference sequence corresponding to characters in the preset text content, wherein the reference time difference sequence corresponding to characters in the preset text content is constructed based on a reference time at which voice data of each character reaches a sound collection component determined by a distance between a pronunciation point of each character in the preset text content and each sound collection component in the plurality of sound collection components; If the time difference sequence corresponding to the characters in the preset text content matches the reference time difference sequence corresponding to the characters in the preset text content, it is determined that the user does not have a preset risk in voiceprint verification.
3. The method according to claim 2, further comprising: Determining the pronunciation point of each character in the preset text content in the oral cavity; Acquire the distance between the pronunciation point of each character in the preset text content and each sound collecting component in the multiple sound collecting components; Determine the distance difference between the pronunciation point of each character in the preset text content and any two sound collection components in the multiple sound collection components based on the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components; Determine a reference time difference between the voice data of each character in the preset text content and reaching any two of the multiple sound collection components based on the distance difference between the pronunciation point of each character in the preset text content and any two of the multiple sound collection components, and the sound speed; Based on the reference time difference between the voice data of each character in the preset text content and the arrival of the voice data of each character in the preset text content at any two of the plurality of sound collection components, a reference time difference sequence corresponding to the characters in the preset text content is constructed.
4. The method according to any one of claims 1 to 3, wherein determining whether the user has a preset risk in voiceprint verification based on the determined time when the voice data of each character in the preset text content in the voice data reaches each sound collection component and the reference time when the voice data of each character reaches the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components comprises: Based on the determined time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component, determine a sub-sequence of the order in which the voice data of each character arrives at the multiple sound collection components; Constructing a sequence corresponding to the characters in the preset text content based on the subsequence of the order in which the voice data of each character arrives at the plurality of sound collection components; Obtaining a reference sequence corresponding to characters in a preset text content, wherein the reference sequence corresponding to characters in the preset text content is constructed based on a reference time at which voice data of each character reaches a sound collection component determined by a distance between a pronunciation point of each character in the preset text content and each sound collection component in the plurality of sound collection components; If the order sequence corresponding to the characters in the preset text content matches the reference order sequence corresponding to the characters in the preset text content, it is determined that the user does not have a preset risk in voiceprint verification.
5. The method according to claim 4, wherein the multiple sound collecting components are two sound collecting components, the two sound collecting components are evenly distributed on the temples, the second sound collecting component of the two sound collecting components is adjacent to a side of the frame to which the temple is connected, and the first sound collecting component of the two sound collecting components is away from a side of the frame to which the temple is connected.
6. The method according to claim 5, wherein receiving voice data for verifying the user through each of the plurality of voice collecting components comprises: Receiving a voiceprint-based resource transfer instruction triggered by the user; Acquire and display the resource information sheet corresponding to the resource transfer instruction, and display confirmation prompt information corresponding to the resource information sheet, wherein the confirmation prompt information is used to prompt the user to read the preset text content by voice; Voice data for authenticating a user is received through each of the plurality of voice collecting components.
7. The method according to claim 6, further comprising: If it is determined that the user does not have a preset risk in voiceprint verification, performing voiceprint-based resource transfer processing based on the voice data; If it is determined that the user has a preset risk in voiceprint verification, the resource transfer process based on the voiceprint based on the voice data is rejected.
8. The method according to claim 7, wherein the preset distance between any two sound collecting components among the plurality of sound collecting components is determined based on the sampling frequency of the sound collecting components.
9. A data processing device, the data processing device being arranged in a glasses-type wearable device, the glasses-type wearable device being provided with a plurality of sound collecting components which are mutually spaced at a distance greater than a preset distance, the device comprising: A data receiving module receives voice data for user verification through each of the plurality of voice collecting components, wherein the voice data includes voice data of preset text content; A data processing module, based on the received voice data, determines the time when the voice data of each character in the preset text content in the voice data arrives at each sound collection component; The data verification module determines whether the user has a preset risk in voiceprint verification based on the time when the voice data of each character in the preset text content in the voice data reaches each sound collection component and the benchmark time when the voice data of each character reaches the sound collection component determined by the distance between the pronunciation point of each character in the preset text content and each sound collection component in the multiple sound collection components.
10. A glasses-type wearable device, comprising a plurality of sound collection components, a glasses-type wearable device body and a control and power supply component, wherein: The control and power supply components are arranged in the body of the eyeglass-type wearable device; The multiple sound collection components are respectively arranged in the body of the eyeglass-type wearable device and are respectively connected to the control and power supply components; The control and power supply component is configured to control at least two of the multiple sound collection components to receive voice data including preset text content, so as to determine whether the user has a preset risk in voiceprint verification based on the voice data of each character in the preset text content in the voice data and the time when the pronunciation point of each character reaches each of the at least two sound collection components respectively; The mutual interval between any two sound collecting components of the at least two sound collecting components is greater than a preset distance, and the preset distance is determined based on the sampling frequency of the sound collecting components of the at least two sound collecting components.
11. The eyeglass wearable device according to claim 10, wherein the eyeglass wearable device body comprises two temples, a frame and two nose pads, and the at least two sound collection components are two sound collection components, and the two sound collection components are evenly distributed on one of the two temples, and the second sound collection component of the two sound collection components is adjacent to the side of the frame connected to the temple, and the first sound collection component of the two sound collection components is away from the side of the frame connected to the temple; or, the two sound collection components are respectively arranged on the two temples, or, one of the two sound collection components is arranged on the frame, and the other sound collection component is arranged on the frame or the temple; or, one of the two sound collection components is arranged on one of the two nose pads, and the other sound collection component is arranged on the other nose pad, the temple or the frame.
12. The eyeglass-type wearable device according to claim 11, wherein the at least two sound collection components include a bone conduction sound collection component, and the bone conduction sound collection component is arranged at a position in contact with the user's body in a wearing state.
13. The eyeglasses wearable device according to claim 11, wherein the sound collecting components other than the at least two sound collecting components in the plurality of sound collecting components include a bone conduction sound collecting component, and the bone conduction sound collecting component is arranged at a position in contact with the user's body in a wearing state; The bone conduction sound collection component is configured to receive a bone conduction sound signal corresponding to the voice data, and when the bone conduction sound signal is not received, trigger the control and power supply component to generate a control instruction related to a preset risk of the user in voiceprint verification.
14. According to the eyeglasses-type wearable device according to claim 12 or 13, the bone conduction sound collection component is arranged on the inner side of the nose pad in the eyeglasses-type wearable device body and contacts the bridge of the nose of the user, and / or the bone conduction sound collection component is arranged on the temple and contacts the cheek and / or ear of the user.
Citation Information
Patent Citations
Information processing method and device and terminal
CN107767137A
Identity authentication method and device
CN110210196A
Identity verification method and device based on artificial intelligence, medium and electronic equipment
CN111949965A
Identity verification method and device, electronic equipment and computer readable medium
CN113051537A
AR glasses payment method and device, AR glasses, medium and program product
CN115808967A
Cited By
Intelligent control method and system of electric appliance
CN122395248A