Surgical safety verification method and system based on artificial intelligence voiceprint recognition

By segmenting the voice information of surgical participants and using voiceprint recognition that integrates multiple results, the problem of low reliability in AI-based surgical security verification has been solved, achieving higher reliability in identity verification.

CN115938371BActive Publication Date: 2026-01-30FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211220782.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2026-01-30
Estimated Expiration
2042-10-08

AI Technical Summary

Technical Problem

In existing technologies, the reliability of AI-based surgical security checks is not high, especially in the process of identity verification.

Method used

An AI-based voiceprint recognition method is used to collect and segment the voice information of the user to be verified. A pre-trained voiceprint recognition model is used to identify the segmented audio segments. Identity verification is performed through multiple audio recognition results, thereby improving the granularity and quantity of recognition.

Benefits of technology

By using fine-grained audio segment recognition and multi-result fusion, the reliability of surgical safety verification is improved, ensuring the accuracy of identity verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938371B_ABST
    Figure CN115938371B_ABST
Patent Text Reader

Abstract

This invention provides a surgical security verification method and system based on artificial intelligence voiceprint recognition, belonging to the field of artificial intelligence technology. In this invention, voice information is collected from the user to be verified to output corresponding audio of the user to be identified. Audio segmentation is performed on the audio of the user to be identified to form multiple audio segments corresponding to the audio of the user to be identified. Based on a pre-trained voiceprint recognition model, each audio segment of the multiple audio segments is recognized to output multiple audio recognition results corresponding to the multiple audio segments. Then, based on the multiple audio recognition results, an identity verification operation is performed on the user to be verified to output the user identity verification result. Based on the above method, the reliability of surgical security verification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a surgical safety verification method and system based on artificial intelligence voiceprint recognition. Background Technology

[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. With the research and advancement of AI technology, it is being researched and applied in multiple fields, such as smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, autonomous driving, robotics, and intelligent healthcare. As technology develops, AI will be applied in even more fields.

[0003] In smart healthcare applications, surgical safety can be verified. For example, voiceprint recognition can identify surgical participants, and the verification results can be compared with surgical case information. Furthermore, a verification form can be generated after verification. During the verification process, if omissions or errors are found during comparison with surgical case information, appropriate prompts can be provided to improve the reliability of the surgery.

[0004] However, in the existing technology, the reliability of relying on artificial intelligence to verify surgical safety, such as the identity verification of participants, may be low. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a surgical safety verification method and system based on artificial intelligence voiceprint recognition, so as to improve the reliability of surgical safety verification.

[0006] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:

[0007] A surgical security verification method based on artificial intelligence voiceprint recognition is applied to a security verification server. The surgical security verification method includes:

[0008] The voice information collection operation is performed on the user to be verified, so as to output the audio of the user to be identified corresponding to the user to be verified, the audio of the user to be identified includes multiple user audio frames;

[0009] The audio of the user to be identified is subjected to audio segmentation to form multiple audio segments of the user to be identified, and each of the multiple audio segments of the user to be identified includes multiple user audio frames.

[0010] Based on the pre-trained voiceprint recognition model, each of the multiple audio segments of the user to be identified is identified, and multiple audio recognition results corresponding to the multiple audio segments of the user to be identified are output. Then, based on the multiple audio recognition results, the identity verification operation of the user to be verified is performed, and the user identity verification result is output. The user identity verification result is used to reflect whether the user to be verified is a participant in the target surgery.

[0011] In some preferred embodiments, in the above-described surgical security verification method based on artificial intelligence voiceprint recognition, the step of collecting voice information from the user to be verified and outputting the audio of the user to be identified includes:

[0012] Upon receiving a verification instruction from a participant in the target surgery, a voice information collection instruction is generated and then sent to the target terminal device. After receiving the voice information collection instruction, the target terminal device performs a voice information collection operation on the user to be verified corresponding to the target terminal device to generate the audio of the user to be identified corresponding to the user to be verified, and then reports the audio of the user to be identified to the security verification server.

[0013] After sending the voice information collection command to the target terminal device, the system receives the user audio to be identified collected and reported by the target terminal device in accordance with the voice information collection command.

[0014] In some preferred embodiments, in the above-described surgical security verification method based on artificial intelligence voiceprint recognition, the step of generating a voice information collection instruction upon receiving a verification instruction from a participant in the target surgery, and then sending the voice information collection instruction to the target terminal device, includes:

[0015] Upon receiving a verification instruction from a participant in the target surgery, the verification instruction is parsed to output the surgical content information of the target surgery.

[0016] The surgical difficulty of the target surgery is determined based on the surgical content information, and the surgical difficulty information corresponding to the target surgery is output.

[0017] A voice information collection instruction is generated based on the surgical difficulty information, and then the voice information collection instruction is sent to the target terminal device. There is a positive correlation between the surgical difficulty information and the audio collection duration included in the voice information collection instruction.

[0018] In some preferred embodiments, in the above-described surgical security verification method based on artificial intelligence voiceprint recognition, the step of receiving the user audio to be identified collected and reported by the target terminal device according to the voice information collection instruction after sending the voice information collection instruction to the target terminal device includes:

[0019] After sending the voice information collection command to the target terminal device, the system receives the user audio to be identified collected and reported by the target terminal device according to the voice information collection command, and performs an audio validity identification operation on the user audio to be identified to output the audio validity identification result corresponding to the user audio to be identified. The audio validity identification result is used to reflect whether the user audio to be identified belongs to silent audio.

[0020] If the audio validity recognition result corresponding to the audio of the user to be identified indicates that the audio of the user to be identified is silent audio, the audio of the user to be identified is discarded, and the process is reversed to execute the steps of generating a voice information collection instruction and then sending the voice information collection instruction to the target terminal device when the verification instruction of the participant in the target surgery is received.

[0021] If the audio validity recognition result corresponding to the audio of the user to be identified reflects that the audio of the user to be identified is not silent audio, the audio of the user to be identified is retained.

[0022] In some preferred embodiments, in the above-described surgical security verification method based on artificial intelligence voiceprint recognition, the step of performing audio segmentation on the user audio to be identified to form multiple user audio segments corresponding to the user audio to be identified includes:

[0023] Perform an audio-to-text conversion operation on the audio of the user to be identified, and output the text data of the user to be identified corresponding to the audio of the user to be identified;

[0024] Based on the text data of the user to be identified, an audio segmentation operation is performed on the audio of the user to be identified to form multiple audio segments of the user to be identified.

[0025] In some preferred embodiments, in the above-described surgical security verification method based on artificial intelligence voiceprint recognition, the step of performing audio segmentation on the audio of the user to be identified based on the text data of the user to be identified to form multiple audio segments corresponding to the audio of the user to be identified includes:

[0026] The text data of the user to be identified is segmented to output multiple candidate text data fragments corresponding to the text data of the user to be identified;

[0027] For each pair of adjacent candidate text data segments among the plurality of candidate text data segments, a text similarity calculation operation is performed on the two candidate text data segments to output the text similarity between the two candidate text data segments. Then, the text similarity between the two candidate text data segments is compared with a pre-configured text similarity reference value. If the text similarity is less than or equal to the text similarity reference value, the two candidate text data segments are marked as corresponding non-related candidate text data segments.

[0028] If, among the plurality of candidate text data segments, there are at least two candidate text data segments that do not belong to each other's non-corresponding candidate text data segments, the step of segmenting the user text data to be identified and outputting the plurality of candidate text data segments corresponding to the user text data to be identified is reversed. If, among the plurality of candidate text data segments, every two adjacent candidate text data segments belong to each other's non-corresponding candidate text data segments, each candidate text data segment in the currently output plurality of candidate text data segments is marked as a target text data segment, so as to output a plurality of target text data segments.

[0029] Based on the multiple target text data segments, audio segmentation is performed on the audio of the user to be identified to form multiple candidate audio segments of the user to be identified.

[0030] For each pair of adjacent candidate audio segments among the plurality of candidate user audio segments to be identified, an audio similarity calculation operation is performed on the two candidate audio segments to be identified, so as to output the audio similarity between the two candidate audio segments to be identified.

[0031] If, in the plurality of candidate audio segments to be identified, there is at least one audio similarity greater than or equal to a preset audio similarity comparison value among the audio similarity between any two adjacent candidate audio segments to be identified, the step of segmenting the user text data to be identified to output a plurality of candidate text data segments corresponding to the user text data to be identified is reversed. If, in the plurality of candidate audio segments to be identified, there is no audio similarity greater than or equal to a preset audio similarity comparison value among the audio similarity between any two adjacent candidate audio segments to be identified, each candidate audio segment to be identified is marked as a user audio segment to be identified, thereby forming a plurality of user audio segments to be identified.

[0032] In some preferred embodiments, in the above-described surgical security verification method based on artificial intelligence voiceprint recognition, the step of calculating the audio similarity between any two adjacent candidate audio segments among the plurality of candidate candidate audio segments to output the audio similarity between the two candidate audio segments includes:

[0033] The two candidate user audio segments corresponding to the audio similarity calculation operation are respectively labeled as the first candidate user audio segment and the second candidate user audio segment. Then, a serialization operation is performed based on the first audio energy value corresponding to each first user audio frame included in the first candidate user audio segment to output the first audio energy sequence corresponding to the first candidate user audio segment. Then, a serialization operation is performed based on the second audio energy value corresponding to each second user audio frame included in the second candidate user audio segment to output the second audio energy sequence corresponding to the second candidate user audio segment.

[0034] Based on the number of target segments, the first audio energy sequence is segmented to output multiple corresponding first audio energy sequence segments. Then, based on the number of target segments, the second audio energy sequence is segmented to output multiple corresponding second audio energy sequence segments. The number of multiple first audio energy sequence segments is equal to the number of target segments, and the number of multiple second audio energy sequence segments is equal to the number of target segments.

[0035] For each first audio energy sequence segment, each first audio energy value included in the first audio energy sequence segment and the audio frame timing of the corresponding first user audio frame are subjected to coordinate transformation to output a first audio feature coordinate set corresponding to the first audio energy sequence segment. The first audio feature coordinate set includes multiple first audio feature coordinates. For each second audio energy sequence segment, each second audio energy value included in the second audio energy sequence segment and the audio frame timing of the corresponding second user audio frame are subjected to coordinate transformation to output a second audio feature coordinate set corresponding to the second audio energy sequence segment. The second audio feature coordinate set includes multiple second audio feature coordinates.

[0036] For each first audio energy sequence segment, the first audio feature coordinates included in the first audio feature coordinate set corresponding to the first audio energy sequence segment are connected in pairs to form a first connecting line between every two first audio feature coordinates. Then, based on the principle of minimizing the area, the first connecting line and the first audio feature coordinates are traversed to output the first region corresponding to the first audio energy sequence segment. The audio energy value corresponding to the center point of the first region is marked as the target first audio energy value corresponding to the first audio energy sequence segment. Each edge of the first region belongs to a first connecting line, and the vertices included in the first region coincide with the corresponding multiple first audio feature coordinates.

[0037] For each second audio energy sequence segment, the set of second audio feature coordinates corresponding to the second audio energy sequence segment is connected in pairs to form a second connecting line between each pair of second audio feature coordinates. Then, based on the principle of minimizing the area, the second connecting line and the second audio feature coordinates are traversed to output the second region corresponding to the second audio energy sequence segment. The audio energy value corresponding to the center point of the second region is marked as the target second audio energy value corresponding to the second audio energy sequence segment. Each edge of the second region is a second connecting line, and the vertices of the second region coincide with the corresponding set of second audio feature coordinates.

[0038] A set construction operation is performed based on the target first audio energy value corresponding to each first audio energy sequence segment to output an ordered set of first audio energy values. Then, a set construction operation is performed based on the target second audio energy value corresponding to each second audio energy sequence segment to output an ordered set of second audio energy values. Then, based on the difference between the target first audio energy value and the target second audio energy value at the corresponding set position, a similarity calculation operation is performed on the ordered set of first audio energy values ​​and the ordered set of second audio energy values ​​to output the audio similarity between the two candidate user audio segments to be identified.

[0039] In some preferred embodiments, in the above-described surgical security verification method based on artificial intelligence voiceprint recognition, the step of performing a recognition operation on each of the plurality of user audio segments to be identified according to a pre-trained voiceprint recognition model to output multiple audio recognition results corresponding to the plurality of user audio segments to be identified, and then performing an identity verification operation on the user to be verified based on the multiple audio recognition results to output the user identity verification result, includes:

[0040] Based on the pre-trained voiceprint recognition model, each of the multiple audio segments of the user to be identified is identified to perform recognition operations, so as to output multiple audio recognition results corresponding to the multiple audio segments of the user to be identified. Each audio recognition result is used to reflect the user identity information of the user to be verified.

[0041] The multiple audio recognition results are fused to output the target audio recognition result corresponding to the multiple audio recognition results. Then, the identity verification operation of the user to be verified is performed based on the target audio recognition result to output the user identity verification result.

[0042] In some preferred embodiments, in the above-described surgical security verification method based on artificial intelligence voiceprint recognition, the step of fusing the multiple audio recognition results to output a target audio recognition result corresponding to the multiple audio recognition results, and then performing an identity verification operation on the user to be verified based on the target audio recognition result to output the user identity verification result, includes:

[0043] Based on whether the reflected user identity information is the same, the multiple audio recognition results are classified to output at least one recognition result classification set corresponding to the multiple audio recognition results;

[0044] For each of the at least one recognition result classification sets, a statistical operation is performed on the number of audio recognition results included in the recognition result classification set to output the statistical value of the number of results corresponding to the recognition result classification set;

[0045] The audio recognition result corresponding to the recognition result category set with the largest corresponding result quantity statistical value is marked as the target audio recognition result corresponding to the multiple audio recognition results;

[0046] Based on the target audio recognition result, an identity verification operation is performed on the user to be verified to output the user identity verification result. If the user identity information reflected in the target audio recognition result does not match the identity information of the participants in the target surgery, the user identity verification result is used to indicate that the user to be verified is not a participant in the target surgery. If the user identity information reflected in the target audio recognition result matches the identity information of the participants in the target surgery, the user identity verification result is used to indicate that the user to be verified is a participant in the target surgery.

[0047] This invention also provides a surgical safety verification system based on artificial intelligence voiceprint recognition, applied to a safety verification server. The surgical safety verification system includes:

[0048] The voice information acquisition module is used to perform voice information acquisition operations on the user to be verified, so as to output the audio of the user to be identified corresponding to the user to be verified, and the audio of the user to be identified includes multiple user audio frames.

[0049] The audio segmentation module is used to perform audio segmentation on the user audio to be identified, so as to form multiple user audio segments corresponding to the user audio to be identified, and each user audio segment to be identified includes multiple user audio frames.

[0050] The identity verification module is used to perform recognition operations on each of the multiple audio segments of the user to be identified according to the pre-trained voiceprint recognition model, so as to output multiple audio recognition results corresponding to the multiple audio segments of the user to be identified, and then perform identity verification operations on the user to be verified according to the multiple audio recognition results, so as to output user identity verification results. The user identity verification results are used to reflect whether the user to be verified is a participant in the target surgery.

[0051] This invention provides a surgical safety verification method and system based on artificial intelligence voiceprint recognition. It can collect voice information from the user to be verified, outputting corresponding audio of the user to be identified. The audio is segmented to form multiple audio segments corresponding to the user's audio. Each audio segment is then recognized using a pre-trained voiceprint recognition model, outputting multiple audio recognition results. Finally, the user's identity is verified based on these multiple audio recognition results, outputting the user's identity verification result. By segmenting the user's audio into multiple audio segments, allowing for separate recognition operations, and then performing identity verification based on the multiple audio recognition results, the reliability of surgical safety verification can be improved to some extent by changing the granularity and number of recognition operations—that is, by making the recognition operations smaller and more numerous.

[0052] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0053] Figure 1 This is a structural block diagram of the security verification server provided in an embodiment of the present invention.

[0054] Figure 2 This is a flowchart illustrating the steps of the surgical safety verification method based on artificial intelligence voiceprint recognition provided in this embodiment of the invention.

[0055] Figure 3 This is a schematic diagram of the modules included in the surgical safety verification system based on artificial intelligence voiceprint recognition provided in an embodiment of the present invention. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0057] Reference Figure 1As shown, this embodiment of the invention provides a security verification server. The security verification server may include a memory and a processor.

[0058] For example, in one possible implementation, the memory and processor are electrically connected directly or indirectly to enable data transmission or interaction. For instance, they can be electrically connected via one or more communication buses or signal lines. The memory may store at least one software functional module, which can exist in the form of software or firmware. The processor can be used to execute the executable computer program stored in the memory, thereby implementing the surgical security verification method based on artificial intelligence voiceprint recognition provided in this embodiment of the invention.

[0059] For example, in one possible implementation, the memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The processor may be a general-purpose processor, including a Central Processing Unit (CPU), Network Processor (NP), System on Chip (SoC), etc.; it may also be a Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0060] For example, in one possible implementation, Figure 1 The structure shown is for illustrative purposes only; the security verification server may also include components such as... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown may include, for example, a communication unit for exchanging information with other devices (such as a terminal device for collecting voice information).

[0061] Reference Figure 2As shown, this embodiment of the invention also provides a surgical security verification method based on artificial intelligence voiceprint recognition, which can be applied to the aforementioned security verification server. The method steps defined in the relevant process of the surgical security verification method based on artificial intelligence voiceprint recognition can be implemented by the security verification server. The following will describe... Figure 2 The specific process shown will be explained in detail.

[0062] Step S110: Perform voice information collection operation on the user to be verified, so as to output the audio of the user to be identified corresponding to the user to be verified.

[0063] In this embodiment of the invention, the security verification server can perform voice information collection operations on the user to be verified, so as to output the audio of the user to be identified corresponding to the user to be verified. The audio of the user to be identified includes multiple user audio frames (which may be sequential in time).

[0064] Step S120: Perform audio segmentation on the audio of the user to be identified to form multiple audio segments of the user to be identified.

[0065] In this embodiment of the invention, the security verification server can perform audio segmentation on the audio of the user to be identified to form multiple audio segments corresponding to the audio of the user to be identified. Each of the multiple audio segments to be identified includes multiple audio frames (which may be sequentially continuous).

[0066] Step S130: Based on the pre-trained voiceprint recognition model, each of the multiple audio segments of the user to be identified is identified to perform a recognition operation, so as to output multiple audio recognition results corresponding to the multiple audio segments of the user to be identified. Then, based on the multiple audio recognition results, the identity verification operation of the user to be verified is performed to output the user identity verification result.

[0067] In this embodiment of the invention, the security verification server can perform recognition operations on each of the plurality of audio segments of the user to be identified according to a pre-trained voiceprint recognition model (the voiceprint recognition model can be a trained multi-class neural network model, the training process of which is not specifically limited here), so as to output multiple audio recognition results corresponding to the plurality of audio segments of the user to be identified, and then perform identity verification operations on the user to be verified according to the multiple audio recognition results, so as to output user identity verification results. The user identity verification results are used to reflect whether the user to be verified is a participant in the target surgery.

[0068] Based on the foregoing, since the audio of the user to be identified is first divided into multiple audio segments, which can be identified separately, and then the identity verification operation is performed based on the multiple audio recognition results, the reliability of surgical safety verification can be improved to a certain extent by changing the granularity and number of recognition operations, so that the granularity of the recognition operation is smaller and the number is greater.

[0069] For example, in one possible implementation, step S110 above may further include the following more detailed content:

[0070] Upon receiving a verification instruction from a participant in the target surgery, a voice information collection instruction is generated and then sent to the target terminal device. After receiving the voice information collection instruction, the target terminal device performs a voice information collection operation on the user to be verified corresponding to the target terminal device to generate the audio of the user to be identified corresponding to the user to be verified, and then reports the audio of the user to be identified to the security verification server.

[0071] After sending the voice information collection command to the target terminal device, the system receives the user audio to be identified collected and reported by the target terminal device in accordance with the voice information collection command.

[0072] For example, in one possible implementation, the step described above—generating a voice information collection instruction upon receiving a verification instruction from a participant in the target surgery, and then sending the voice information collection instruction to the target terminal device—can further include the following more detailed content:

[0073] Upon receiving a verification instruction from a participant in the target surgery, the verification instruction is parsed to output the surgical content information of the target surgery.

[0074] The surgical difficulty of the target surgery is determined based on the surgical content information (for example, the difficulty of various surgeries can be defined in advance, or the surgical content can be analyzed, such as the greater the number of steps, the greater the difficulty), so as to output the surgical difficulty information corresponding to the target surgery.

[0075] A voice information collection instruction is generated based on the surgical difficulty information, and then the voice information collection instruction is sent to the target terminal device. There is a positive correlation between the surgical difficulty information and the audio collection duration included in the voice information collection instruction.

[0076] For example, in one possible implementation, the step described above—receiving the user audio to be identified collected and reported by the target terminal device according to the voice information collection instruction after sending the voice information collection instruction to the target terminal device—can further include the following more detailed content:

[0077] After the voice information collection command is sent to the target terminal device, the system receives the user audio to be identified collected and reported by the target terminal device according to the voice information collection command, and performs an audio validity identification operation on the user audio to be identified to output the audio validity identification result corresponding to the user audio to be identified. The audio validity identification result is used to reflect whether the user audio to be identified belongs to silent audio (wherein, silent audio can mean that all audio frames are silent audio frames, or it can mean that a certain proportion of audio frames belong to silent audio frames).

[0078] If the audio validity recognition result corresponding to the audio of the user to be identified indicates that the audio of the user to be identified is silent audio, the audio of the user to be identified is discarded, and the process is reversed to execute the steps of generating a voice information collection instruction and then sending the voice information collection instruction to the target terminal device when the verification instruction of the participant in the target surgery is received.

[0079] If the audio validity recognition result corresponding to the audio of the user to be identified reflects that the audio of the user to be identified is not silent audio, the audio of the user to be identified is retained.

[0080] For example, in one possible implementation, step S120 above may further include the following more detailed content:

[0081] Perform an audio-to-text conversion operation on the audio of the user to be identified (refer to relevant existing technologies) to output the text data of the user to be identified corresponding to the audio of the user to be identified;

[0082] Based on the text data of the user to be identified, an audio segmentation operation is performed on the audio of the user to be identified to form multiple audio segments of the user to be identified.

[0083] For example, in one possible implementation, the step described above of performing audio segmentation on the audio of the user to be identified based on the text data of the user to be identified, to form multiple audio segments corresponding to the audio of the user to be identified, may further include the following more detailed content:

[0084] The user text data to be identified is segmented (arbitrary and random segmentation) to output multiple candidate text data fragments corresponding to the user text data to be identified;

[0085] For each pair of adjacent candidate text data segments among the plurality of candidate text data segments, a text similarity calculation operation is performed on the two candidate text data segments (the text similarity calculation method in the prior art can be referred to, and no specific limitation is made here) to output the text similarity between the two candidate text data segments. Then, the text similarity between the two candidate text data segments is compared with a pre-configured text similarity reference value. If the text similarity is less than or equal to the text similarity reference value, the two candidate text data segments are marked as corresponding non-related candidate text data segments.

[0086] If, among the plurality of candidate text data segments, there are at least two candidate text data segments that do not belong to each other's non-corresponding candidate text data segments, the step of segmenting the user text data to be identified and outputting the plurality of candidate text data segments corresponding to the user text data to be identified is reversed. If, among the plurality of candidate text data segments, every two adjacent candidate text data segments belong to each other's non-corresponding candidate text data segments, each candidate text data segment in the currently output plurality of candidate text data segments is marked as a target text data segment, so as to output a plurality of target text data segments.

[0087] Based on the multiple target text data segments, audio segmentation is performed on the audio of the user to be identified to form multiple candidate audio segments of the user to be identified (i.e., one target text data segment corresponds to one candidate audio segment of the user to be identified).

[0088] For each pair of adjacent candidate audio segments among the plurality of candidate user audio segments to be identified, an audio similarity calculation operation is performed on the two candidate audio segments to be identified, so as to output the audio similarity between the two candidate audio segments to be identified.

[0089] If, in the plurality of candidate audio segments to be identified, there is at least one audio similarity greater than or equal to a preset audio similarity comparison value among the audio similarity between any two adjacent candidate audio segments to be identified, the step of segmenting the user text data to be identified to output a plurality of candidate text data segments corresponding to the user text data to be identified is reversed. If, in the plurality of candidate audio segments to be identified, there is no audio similarity greater than or equal to a preset audio similarity comparison value among the audio similarity between any two adjacent candidate audio segments to be identified, each candidate audio segment to be identified is marked as a user audio segment to be identified, thereby forming a plurality of user audio segments to be identified.

[0090] For example, in one possible implementation, the step described above, which involves calculating the audio similarity between every two adjacent candidate audio segments among the plurality of candidate audio segments to be identified, and outputting the audio similarity between the two candidate audio segments, may further include the following more detailed content:

[0091] The two candidate user audio segments corresponding to the audio similarity calculation operation are respectively labeled as the first candidate user audio segment and the second candidate user audio segment. Then, a serialization operation is performed based on the first audio energy value corresponding to each first user audio frame included in the first candidate user audio segment (the audio energy values ​​can be sorted according to the temporal relationship of the corresponding audio frames) to output the first audio energy sequence corresponding to the first candidate user audio segment. Then, a serialization operation is performed based on the second audio energy value corresponding to each second user audio frame included in the second candidate user audio segment to output the second audio energy sequence corresponding to the second candidate user audio segment.

[0092] Based on the number of target segments, the first audio energy sequence is segmented to output multiple corresponding first audio energy sequence segments. Then, based on the number of target segments, the second audio energy sequence is segmented to output multiple corresponding second audio energy sequence segments. The number of multiple first audio energy sequence segments is equal to the number of target segments, and the number of multiple second audio energy sequence segments is equal to the number of target segments.

[0093] For each first audio energy sequence segment, the first audio energy value and the corresponding audio frame timing of the first user audio frame included in the first audio energy sequence segment are respectively subjected to coordinate operation (that is, the first audio energy value and the audio frame timing are used as corresponding two-dimensional coordinates) to output the first audio feature coordinate set corresponding to the first audio energy sequence segment. The first audio feature coordinate set includes multiple first audio feature coordinates. For each second audio energy sequence segment, the second audio energy value and the corresponding audio frame timing of the second user audio frame included in the second audio energy sequence segment are respectively subjected to coordinate operation to output the second audio feature coordinate set corresponding to the second audio energy sequence segment. The second audio feature coordinate set includes multiple second audio feature coordinates.

[0094] For each first audio energy sequence segment, the first audio feature coordinates included in the first audio feature coordinate set corresponding to the first audio energy sequence segment are connected in pairs to form a first connecting line between every two first audio feature coordinates. Then, based on the principle of minimizing the area, the first connecting line and the first audio feature coordinates are traversed to output the first region corresponding to the first audio energy sequence segment. The audio energy value corresponding to the center point of the first region is marked as the target first audio energy value corresponding to the first audio energy sequence segment. Each edge of the first region belongs to a first connecting line, and the vertices included in the first region coincide with the corresponding multiple first audio feature coordinates.

[0095] For each second audio energy sequence segment, the set of second audio feature coordinates corresponding to the second audio energy sequence segment is connected in pairs to form a second connecting line between each pair of second audio feature coordinates. Then, based on the principle of minimizing the area, the second connecting line and the second audio feature coordinates are traversed to output the second region corresponding to the second audio energy sequence segment. The audio energy value corresponding to the center point of the second region is marked as the target second audio energy value corresponding to the second audio energy sequence segment. Each edge of the second region is a second connecting line, and the vertices of the second region coincide with the corresponding set of second audio feature coordinates.

[0096] A set construction operation is performed based on the target first audio energy value corresponding to each first audio energy sequence segment to output an ordered set of first audio energy values. Then, a set construction operation is performed based on the target second audio energy value corresponding to each second audio energy sequence segment to output an ordered set of second audio energy values. Then, based on the difference between the target first audio energy value and the target second audio energy value at the corresponding set position, a similarity calculation operation is performed on the ordered set of first audio energy values ​​and the ordered set of second audio energy values ​​(for example, the difference between each set position can be fused first, and then the similarity can be determined based on the fused value, which can be negatively correlated with the fused value, and the fused value can be the sum of the differences) to output the audio similarity between the two candidate user audio segments to be identified.

[0097] For example, in another possible implementation, the step described above, which involves calculating the audio similarity between every two adjacent candidate audio segments among the plurality of candidate user audio segments to output the audio similarity between the two candidate user audio segments, may further include the following more detailed content:

[0098] The two candidate user audio segments corresponding to the audio similarity calculation operation are respectively labeled as the first candidate user audio segment and the second candidate user audio segment. Then, a serialization operation is performed based on the first audio energy value corresponding to each first user audio frame included in the first candidate user audio segment to output the first audio energy sequence corresponding to the first candidate user audio segment. Then, a serialization operation is performed based on the second audio energy value corresponding to each second user audio frame included in the second candidate user audio segment to output the second audio energy sequence corresponding to the second candidate user audio segment.

[0099] For each first audio energy value included in the first audio energy sequence, a coordinate transformation operation is performed on the first audio energy value and the audio frame timing of the first user audio frame corresponding to the first audio energy value to output the first audio feature coordinates corresponding to the first audio energy value. For each second audio energy value included in the second audio energy sequence, a coordinate transformation operation is performed on the second audio energy value and the audio frame timing of the second user audio frame corresponding to the second audio energy value to output the second audio feature coordinates corresponding to the second audio energy value.

[0100] Clustering operation is performed on the first audio feature coordinates (refer to relevant existing clustering techniques, such as the nearest neighbor algorithm, etc.) to output at least one first cluster set, and then the first audio feature coordinates corresponding to at least one first cluster center of the at least one first cluster set are marked as target first audio feature coordinates;

[0101] Clustering operation is performed on the second audio feature coordinates to output at least one second cluster set, and then the second audio feature coordinates corresponding to at least one second cluster center of the at least one second cluster set are marked as target second audio feature coordinates;

[0102] The system performs a statistical operation on the proportion of the number of the first audio feature coordinates of the target currently available to output a first proportion. Then, it performs a statistical operation on the proportion of the number of the second audio feature coordinates of the target currently available to output a second proportion. The system then compares the first proportion with a pre-configured proportion threshold and compares the second proportion with the proportion threshold.

[0103] When the first quantity proportion is less than the quantity proportion threshold, the following steps are performed: clustering the first audio feature coordinates based on other first audio feature coordinates besides the target first audio feature coordinates, to output at least one first cluster set, and then marking the first audio feature coordinates corresponding to at least one first cluster center of the at least one first cluster set as the target first audio feature coordinates, to form new target first audio feature coordinates, until the quantity proportion of the currently existing target first audio feature coordinates is greater than or equal to the quantity proportion threshold.

[0104] When the second quantity proportion is less than the quantity proportion threshold, the following steps are performed: clustering the second audio feature coordinates based on other second audio feature coordinates besides the target second audio feature coordinates, to output at least one second cluster set, and then marking the second audio feature coordinates corresponding to at least one second cluster center of the at least one second cluster set as the target second audio feature coordinates, to form new target second audio feature coordinates, until the quantity proportion of the currently existing target second audio feature coordinates is greater than or equal to the quantity proportion threshold.

[0105] If the proportion of the number of current target first audio feature coordinates is greater than or equal to the threshold, and the proportion of the number of current target second audio feature coordinates is greater than or equal to the threshold, the overlap of the current target first audio feature coordinates and the current target second audio feature coordinates is calculated (for example, a first set can be formed based on the current target first audio feature coordinates, a second set can be formed based on the current target second audio feature coordinates, and the overlap of the first set and the second set can be calculated) to output the target coordinate overlap. Then, the distribution state of the first audio energy value corresponding to the current target first audio feature coordinates is determined to output the first energy value distribution information (such as a distribution histogram). And, the distribution state of the second audio energy value corresponding to the current target second audio feature coordinates is determined to output the second energy value distribution information (similarly, it can also be a distribution histogram).

[0106] The first energy value distribution information and the second energy value distribution information are subjected to a similarity calculation operation of distribution state (for example, the similarity between two distribution histograms can be calculated, that is, the graphic similarity between two histograms can be calculated, and the similarity between two histograms can be calculated by referring to the existing graphic similarity or contour similarity calculation method), so as to output the target distribution similarity. Then, the target distribution similarity and the target coordinate overlap are fused (such as calculating the weighted mean of the target distribution similarity and the target coordinate overlap), so as to output the audio similarity between the two candidate user audio segments to be identified.

[0107] For example, in one possible implementation, step S130 above may further include the following more detailed content:

[0108] Based on the pre-trained voiceprint recognition model, each of the multiple audio segments of the user to be identified is identified to perform recognition operations, so as to output multiple audio recognition results corresponding to the multiple audio segments of the user to be identified. Each audio recognition result is used to reflect the user identity information of the user to be verified.

[0109] The multiple audio recognition results are fused to output the target audio recognition result corresponding to the multiple audio recognition results. Then, the identity verification operation of the user to be verified is performed based on the target audio recognition result to output the user identity verification result.

[0110] For example, in one possible implementation, the step of fusing the multiple audio recognition results to output a target audio recognition result corresponding to the multiple audio recognition results, and then performing an identity verification operation on the user to be verified based on the target audio recognition result to output the user identity verification result, may further include the following more detailed content:

[0111] Based on whether the reflected user identity information is the same, the multiple audio recognition results are classified to output at least one recognition result classification set corresponding to the multiple audio recognition results;

[0112] For each of the at least one recognition result classification sets, a statistical operation is performed on the number of audio recognition results included in the recognition result classification set to output the statistical value of the number of results corresponding to the recognition result classification set;

[0113] The audio recognition result corresponding to the recognition result category set with the largest corresponding result quantity statistical value is marked as the target audio recognition result corresponding to the multiple audio recognition results;

[0114] Based on the target audio recognition result, an identity verification operation is performed on the user to be verified to output the user identity verification result. If the user identity information reflected in the target audio recognition result does not match the identity information of the participants in the target surgery, the user identity verification result is used to indicate that the user to be verified is not a participant in the target surgery. If the user identity information reflected in the target audio recognition result matches the identity information of the participants in the target surgery, the user identity verification result is used to indicate that the user to be verified is a participant in the target surgery.

[0115] Reference Figure 3 As shown, this embodiment of the invention also provides a surgical security verification system based on artificial intelligence voiceprint recognition, which can be applied to the aforementioned security verification server. The surgical security verification system based on artificial intelligence voiceprint recognition may include the following software functional modules, specifically such as a voice information acquisition module, an audio segmentation module, and an identity verification module.

[0116] In one possible implementation, the voice information acquisition module is used to acquire voice information from the user to be verified, and output the audio of the user to be identified, which includes multiple audio frames. The audio segmentation module is used to segment the audio of the user to be identified, forming multiple audio segments corresponding to the audio of the user to be identified, each of the multiple audio segments including multiple audio frames. The identity verification module is used to perform recognition operations on each of the multiple audio segments according to a pre-trained voiceprint recognition model, and output multiple audio recognition results corresponding to the multiple audio segments. Then, based on the multiple audio recognition results, the module performs identity verification on the user to be verified, and outputs a user identity verification result, which reflects whether the user to be verified is a participant in the target surgery.

[0117] In summary, the surgical safety verification method and system based on artificial intelligence voiceprint recognition provided by this invention can collect voice information from the user to be verified, and output the corresponding audio of the user to be identified. The audio of the user to be identified is segmented to form multiple audio segments corresponding to the audio of the user to be identified. Each audio segment of the multiple audio segments is identified according to a pre-trained voiceprint recognition model, resulting in multiple audio recognition results. Then, based on the multiple audio recognition results, the identity verification operation of the user to be verified is performed, outputting the user identity verification result. Based on the foregoing, since the audio of the user to be identified is first segmented into multiple audio segments, identification operations can be performed separately. Then, the identity verification operation is performed based on the multiple audio recognition results. By changing the granularity and number of identification operations, the reliability of surgical safety verification can be improved to a certain extent.

[0118] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A surgical safety verification method based on artificial intelligence voiceprint recognition, characterized in that, The surgical safety verification method applied to a safety verification server comprises: performing a voice information collection operation on a user to be verified to output user audio to be identified corresponding to the user to be verified, the user audio to be identified comprising a plurality of user audio frames; performing an audio segmentation operation on the user audio to be identified to form a plurality of user audio segments to be identified corresponding to the user audio to be identified, each of the plurality of user audio segments to be identified comprising a plurality of user audio frames; performing an identification operation on each of the plurality of user audio segments to be identified according to a pre-trained voiceprint recognition model to output a plurality of audio identification results corresponding to the plurality of user audio segments to be identified, and then performing an identity verification operation on the user to be verified according to the plurality of audio identification results to output a user identity verification result, the user identity verification result being used to reflect whether the user to be verified belongs to a participant of a target surgery; the step of performing an audio segmentation operation on the user audio to be identified to form a plurality of user audio segments to be identified corresponding to the user audio to be identified comprises: performing an audio-to-text conversion operation on the user audio to be identified to output user text data to be identified corresponding to the user audio to be identified; performing an audio segmentation operation on the user audio to be identified according to the user text data to be identified to form a plurality of user audio segments to be identified corresponding to the user audio to be identified; the step of performing an audio segmentation operation on the user audio to be identified according to the user text data to be identified to form a plurality of user audio segments to be identified corresponding to the user audio to be identified comprises: performing a segmentation operation on the user text data to be identified to output a plurality of candidate text data segments corresponding to the user text data to be identified; for each of two adjacent candidate text data segments in the plurality of candidate text data segments, performing a text similarity calculation operation on the two candidate text data segments to output a text similarity between the two candidate text data segments, then performing a size comparison operation on the text similarity between the two candidate text data segments and a preconfigured text similarity reference value, and in the case that the text similarity is less than or equal to the text similarity reference value, marking the two candidate text data segments as non-associated candidate text data segments corresponding to each other; In a case where there are at least two candidate text data segments in the plurality of candidate text data segments which do not belong to mutually corresponding non-associated candidate text data segments, the step of performing the segmenting operation on the to-be-recognized user text data to output the plurality of candidate text data segments corresponding to the to-be-recognized user text data is executed in a case where each two adjacent candidate text data segments in the plurality of candidate text data segments belong to mutually corresponding non-associated candidate text data segments, each of the plurality of candidate text data segments currently output is marked as a target text data segment respectively to output a plurality of target text data segments; performing an audio segmentation operation on the to-be-recognized user audio according to the plurality of target text data segments to form a plurality of candidate to-be-recognized user audio segments corresponding to the to-be-recognized user audio; for each two adjacent candidate to-be-recognized user audio segments in the plurality of candidate to-be-recognized user audio segments, performing an audio similarity calculation operation on the two candidate to-be-recognized user audio segments to output an audio similarity between the two candidate to-be-recognized user audio segments; in a case where there is at least one audio similarity greater than or equal to a preset audio similarity comparison value among the audio similarities between each two adjacent candidate to-be-recognized user audio segments in the plurality of candidate to-be-recognized user audio segments, the step of performing the segmenting operation on the to-be-recognized user text data to output the plurality of candidate text data segments corresponding to the to-be-recognized user text data is executed, and in a case where there is no audio similarity greater than or equal to the preset audio similarity comparison value among the audio similarities between each two adjacent candidate to-be-recognized user audio segments in the plurality of candidate to-be-recognized user audio segments, each of the plurality of candidate to-be-recognized user audio segments currently formed is marked as a to-be-recognized user audio segment to form a plurality of to-be-recognized user audio segments; the step of performing the audio similarity calculation operation on each two adjacent candidate to-be-recognized user audio segments in the plurality of candidate to-be-recognized user audio segments to output an audio similarity between the two candidate to-be-recognized user audio segments, comprises: marking the two candidate to-be-recognized user audio segments corresponding to the audio similarity calculation operation as a first candidate to-be-recognized user audio segment and a second candidate to-be-recognized user audio segment respectively, and then performing a serialization operation on each first user audio frame corresponding to a first audio energy value included in the first candidate to-be-recognized user audio segment to output a first audio energy sequence corresponding to the first candidate to-be-recognized user audio segment, and performing a serialization operation on each second user audio frame corresponding to a second audio energy value included in the second candidate to-be-recognized user audio segment to output a second audio energy sequence corresponding to the second candidate to-be-recognized user audio segment; segmenting the first audio energy sequence according to the target segment number to output a plurality of first audio energy sequence segments, and segmenting the second audio energy sequence according to the target segment number to output a plurality of second audio energy sequence segments, wherein the number of the plurality of first audio energy sequence segments is equal to the target segment number, and the number of the plurality of second audio energy sequence segments is equal to the target segment number; for each first audio energy sequence segment, performing coordinate operation on each first audio energy value included in the first audio energy sequence segment and an audio frame time sequence of a corresponding first user audio frame to output a first audio feature coordinate set corresponding to the first audio energy sequence segment, wherein the first audio feature coordinate set includes a plurality of first audio feature coordinates; for each first audio energy sequence segment, performing pairwise connection operation on the plurality of first audio feature coordinates included in the first audio feature coordinate set corresponding to the first audio energy sequence segment to form a first connection line between each two first audio feature coordinates, performing traversal operation on the first connection line and the first audio feature coordinates according to a principle of minimum area to output a first region corresponding to the first audio energy sequence segment, and marking an audio energy value corresponding to a center point of the first region as a target first audio energy value corresponding to the first audio energy sequence segment, wherein each edge of the first region belongs to a first connection line, and vertices included in the first region coincide with the plurality of first audio feature coordinates; for each second audio energy sequence segment, performing pairwise connection operation on the plurality of second audio feature coordinates included in the second audio feature coordinate set corresponding to the second audio energy sequence segment to form a second connection line between each two second audio feature coordinates, performing traversal operation on the second connection line and the second audio feature coordinates according to the principle of minimum area to output a second region corresponding to the second audio energy sequence segment, and marking an audio energy value corresponding to a center point of the second region as a target second audio energy value corresponding to the second audio energy sequence segment, wherein each edge of the second region belongs to a second connection line, and vertices included in the second region coincide with the plurality of second audio feature coordinates; and for each first audio energy sequence segment, performing pairwise connection operation on the plurality of first audio feature coordinates included in the first audio feature coordinate set corresponding to the first audio energy sequence segment to form a first connection line between each two first audio feature coordinates, performing traversal operation on the first connection line and the first audio feature coordinates according to a principle of minimum area to output a first region corresponding to the first audio energy sequence segment, and marking an audio energy value corresponding to a center point of the first region as a target first audio energy value corresponding to the first audio energy sequence segment, wherein each edge of the first region belongs to a first connection line, and vertices included in the first region coincide with the plurality of first audio feature coordinates; for each second audio energy sequence segment, performing pairwise connection operation on the plurality of second audio feature coordinates included in the second audio feature coordinate set corresponding to the second audio energy sequence segment to form a second connection line between each two second audio feature coordinates, performing traversal operation on the second connection line and the second audio feature coordinates according to the principle of minimum area to output a second region corresponding to the second audio energy sequence segment, and marking an audio energy value corresponding to a center point of the second region as a target second audio energy value corresponding to the second audio energy sequence segment, wherein each edge of the second region belongs to a second connection line, and vertices included in the second region coincide with the plurality of second audio feature coordinates; and According to the target first audio energy value corresponding to each first audio energy sequence segment, a set construction operation is performed to output an ordered set of first audio energy values, and according to the target second audio energy value corresponding to each second audio energy sequence segment, a set construction operation is performed to output an ordered set of second audio energy values, and according to the difference between the target first audio energy value and the target second audio energy value corresponding to the set position, a similarity calculation operation is performed on the ordered set of first audio energy values and the ordered set of second audio energy values to output the audio similarity between the two candidate user audio segments to be identified.

2. The artificial intelligence-based voiceprint recognition method for surgical safety verification as claimed in claim 1, wherein, The step of performing voice information collection on the user to be checked to output the user audio to be identified corresponding to the user to be checked comprises: In the case of receiving the checking instruction of the participant of the target operation, a voice information collection instruction is generated, and the voice information collection instruction is issued to the target terminal device, and the target terminal device is used to perform voice information collection on the user to be checked corresponding to the target terminal device after receiving the voice information collection instruction, to generate the user audio to be identified corresponding to the user to be checked, and then report the user audio to be identified to the security checking server; After the voice information collection instruction is issued to the target terminal device, the user audio to be identified collected and reported by the target terminal device according to the voice information collection instruction is received.

3. The artificial intelligence-based voiceprint recognition method for surgical safety verification according to claim 2, wherein, The step of generating a voice information collection instruction and issuing the voice information collection instruction to a target terminal device in the case of receiving a checking instruction of a participant of a target operation comprises: In the case of receiving the checking instruction of the participant of the target operation, the checking instruction is parsed to output the operation content information of the target operation; According to the operation content information, the operation difficulty of the target operation is determined to output the operation difficulty information corresponding to the target operation; According to the operation difficulty information, a voice information collection instruction is generated, and the voice information collection instruction is issued to a target terminal device, and the operation difficulty information and the audio collection time length included in the voice information collection instruction have a positive correlation corresponding relationship.

4. The artificial intelligence-based voiceprint recognition method for surgical safety verification according to claim 2, wherein, The step of receiving the user audio to be identified collected and reported by the target terminal device according to the voice information collection instruction after the voice information collection instruction is issued to the target terminal device comprises: After the voice information collection instruction is issued to the target terminal device, the user audio to be identified collected and reported by the target terminal device according to the voice information collection instruction is received, and the audio effectiveness of the user audio to be identified is identified to output the audio effectiveness identification result corresponding to the user audio to be identified, and the audio effectiveness identification result is used to reflect whether the user audio to be identified is a voiceless audio; In a case where the audio validity identification result corresponding to the to-be-identified user audio reflects that the to-be-identified user audio belongs to non-audio, performing a discarding operation on the to-be-identified user audio, and in a case where the receiving of the checking instruction of the participant of the target surgery is received, generating a voice information collection instruction, and then issuing the voice information collection instruction to the target terminal device; In a case where the audio validity identification result corresponding to the to-be-identified user audio reflects that the to-be-identified user audio does not belong to non-audio, performing a retaining operation on the to-be-identified user audio.

5. The artificial intelligence voiceprint-based surgical safety verification method according to any one of claims 1-4, characterized in that, The step of performing an identification operation on each of the plurality of to-be-identified user audio segments according to the pre-trained voiceprint identification model to output a plurality of audio identification results corresponding to the plurality of to-be-identified user audio segments, and then performing an identity checking operation on the to-be-checked user according to the plurality of audio identification results to output a user identity checking result, comprises: performing an identification operation on each of the plurality of to-be-identified user audio segments according to the pre-trained voiceprint identification model to output a plurality of audio identification results corresponding to the plurality of to-be-identified user audio segments, each audio identification result being used to reflect user identity information of the to-be-checked user; performing a fusion operation on the plurality of audio identification results to output a target audio identification result corresponding to the plurality of audio identification results, and then performing an identity checking operation on the to-be-checked user according to the target audio identification result to output a user identity checking result.

6. The artificial intelligence-based voiceprint recognition method for surgical safety verification as claimed in claim 5 wherein, The step of performing a fusion operation on the plurality of audio identification results to output a target audio identification result corresponding to the plurality of audio identification results, and then performing an identity checking operation on the to-be-checked user according to the target audio identification result to output a user identity checking result, comprises: performing a classification operation on the plurality of audio identification results according to whether the reflected user identity information is the same, to output at least one identification result classification set corresponding to the plurality of audio identification results; for each of the at least one identification result classification set, performing a statistical operation on the number of audio identification results included in the identification result classification set to output a result number statistical value corresponding to the identification result classification set; marking the audio identification result corresponding to the identification result classification set with the maximum result number statistical value as the target audio identification result corresponding to the plurality of audio identification results; and performing an identity checking operation on the to-be-checked user according to the target audio identification result to output a user identity checking result. According to the target audio recognition result, identity verification is performed on the to-be-verified user to output a user identity verification result. If the user identity information reflected by the target audio recognition result does not match the identity information of the participant of the target surgery, the user identity verification result reflects that the to-be-verified user does not belong to the participant of the target surgery. If the user identity information reflected by the target audio recognition result matches the identity information of the participant of the target surgery, the user identity verification result reflects that the to-be-verified user belongs to the participant of the target surgery.

7. A surgical safety verification system based on artificial intelligence voiceprint recognition, characterized in that, The surgical safety verification system is applied to a safety verification server and includes: A voice information collection module is configured to collect voice information of a to-be-verified user to output to-be-identified user audio corresponding to the to-be-verified user. The to-be-identified user audio includes multiple frames of user audio frames. An audio segmentation module is configured to perform audio segmentation on the to-be-identified user audio to form multiple to-be-identified user audio segments corresponding to the to-be-identified user audio. Each to-be-identified user audio segment includes multiple frames of user audio frames. An identity verification module is configured to perform identification on each to-be-identified user audio segment in the multiple to-be-identified user audio segments according to a pre-trained voiceprint recognition model to output multiple audio recognition results corresponding to the multiple to-be-identified user audio segments. The identity verification module performs identity verification on the to-be-verified user according to the multiple audio recognition results to output a user identity verification result. The user identity verification result reflects whether the to-be-verified user belongs to a participant of a target surgery. The audio segmentation on the to-be-identified user audio to form multiple to-be-identified user audio segments corresponding to the to-be-identified user audio includes: Performing audio-to-text conversion on the to-be-identified user audio to output to-be-identified user text data corresponding to the to-be-identified user audio. Performing audio segmentation on the to-be-identified user audio according to the to-be-identified user text data to form multiple to-be-identified user audio segments corresponding to the to-be-identified user audio. The audio segmentation on the to-be-identified user audio according to the to-be-identified user text data to form multiple to-be-identified user audio segments corresponding to the to-be-identified user audio includes: Segmenting the to-be-identified user text data to output multiple candidate text data segments corresponding to the to-be-identified user text data. For each two adjacent candidate text data segments in the multiple candidate text data segments, a text similarity calculation is performed on the two candidate text data segments to output a text similarity between the two candidate text data segments. A size comparison is performed between the text similarity and a preconfigured text similarity reference value. If the text similarity is less than or equal to the text similarity reference value, the two candidate text data segments are marked as non-associated candidate text data segments corresponding to each other. In a case where there are at least two candidate text data segments in the plurality of candidate text data segments which do not belong to mutually corresponding non-associated candidate text data segments, the segmenting operation on the to-be-recognized user text data is performed again to output a plurality of candidate text data segments corresponding to the to-be-recognized user text data, in a case where each two adjacent candidate text data segments in the plurality of candidate text data segments belong to mutually corresponding non-associated candidate text data segments, each of the plurality of currently output candidate text data segments is marked as a target text data segment respectively to output a plurality of target text data segments; According to the plurality of target text data segments, an audio segmentation operation is performed on the to-be-recognized user audio to form a plurality of candidate to-be-recognized user audio segments corresponding to the to-be-recognized user audio; For each two adjacent candidate to-be-recognized user audio segments in the plurality of candidate to-be-recognized user audio segments, an audio similarity calculation operation is performed on the two candidate to-be-recognized user audio segments to output an audio similarity between the two candidate to-be-recognized user audio segments; In a case where there is at least one audio similarity greater than or equal to a preset audio similarity comparison value among the audio similarities between each two adjacent candidate to-be-recognized user audio segments in the plurality of candidate to-be-recognized user audio segments, the segmenting operation on the to-be-recognized user text data is performed again to output a plurality of candidate text data segments corresponding to the to-be-recognized user text data, in a case where there is no audio similarity greater than or equal to a preset audio similarity comparison value among the audio similarities between each two adjacent candidate to-be-recognized user audio segments in the plurality of candidate to-be-recognized user audio segments, each of the plurality of currently formed candidate to-be-recognized user audio segments is marked as a to-be-recognized user audio segment to form a plurality of to-be-recognized user audio segments; The audio similarity calculation operation on each two adjacent candidate to-be-recognized user audio segments in the plurality of candidate to-be-recognized user audio segments to output an audio similarity between the two candidate to-be-recognized user audio segments, comprises: The two candidate to-be-recognized user audio segments corresponding to the audio similarity calculation operation are marked as a first candidate to-be-recognized user audio segment and a second candidate to-be-recognized user audio segment respectively, and then a first audio energy sequence corresponding to the first candidate to-be-recognized user audio segment is output by performing a serialization operation on a first audio energy value corresponding to each frame of first user audio frame included in the first candidate to-be-recognized user audio segment, and a second audio energy sequence corresponding to the second candidate to-be-recognized user audio segment is output by performing a serialization operation on a second audio energy value corresponding to each frame of second user audio frame included in the second candidate to-be-recognized user audio segment. segmenting the first audio energy sequence according to the target segment number to output a plurality of first audio energy sequence segments, and segmenting the second audio energy sequence according to the target segment number to output a plurality of second audio energy sequence segments, wherein the number of the plurality of first audio energy sequence segments is equal to the target segment number, and the number of the plurality of second audio energy sequence segments is equal to the target segment number; for each first audio energy sequence segment, performing coordinate operation on each first audio energy value included in the first audio energy sequence segment and an audio frame time sequence of a corresponding first user audio frame to output a first audio feature coordinate set corresponding to the first audio energy sequence segment, wherein the first audio feature coordinate set includes a plurality of first audio feature coordinates; for each first audio energy sequence segment, performing pairwise connection operation on the plurality of first audio feature coordinates included in the first audio feature coordinate set corresponding to the first audio energy sequence segment to form a first connection line between each two first audio feature coordinates, performing traversal operation on the first connection line and the first audio feature coordinates according to a principle of minimum area to output a first region corresponding to the first audio energy sequence segment, and marking an audio energy value corresponding to a center point of the first region as a target first audio energy value corresponding to the first audio energy sequence segment, wherein each edge of the first region belongs to a first connection line, and vertices included in the first region coincide with the plurality of first audio feature coordinates; for each second audio energy sequence segment, performing pairwise connection operation on the plurality of second audio feature coordinates included in the second audio feature coordinate set corresponding to the second audio energy sequence segment to form a second connection line between each two second audio feature coordinates, performing traversal operation on the second connection line and the second audio feature coordinates according to the principle of minimum area to output a second region corresponding to the second audio energy sequence segment, and marking an audio energy value corresponding to a center point of the second region as a target second audio energy value corresponding to the second audio energy sequence segment, wherein each edge of the second region belongs to a second connection line, and vertices included in the second region coincide with the plurality of second audio feature coordinates; and for each first audio energy sequence segment, performing pairwise connection operation on the plurality of first audio feature coordinates included in the first audio feature coordinate set corresponding to the first audio energy sequence segment to form a first connection line between each two first audio feature coordinates, performing traversal operation on the first connection line and the first audio feature coordinates according to a principle of minimum area to output a first region corresponding to the first audio energy sequence segment, and marking an audio energy value corresponding to a center point of the first region as a target first audio energy value corresponding to the first audio energy sequence segment, wherein each edge of the first region belongs to a first connection line, and vertices included in the first region coincide with the plurality of first audio feature coordinates; for each second audio energy sequence segment, performing pairwise connection operation on the plurality of second audio feature coordinates included in the second audio feature coordinate set corresponding to the second audio energy sequence segment to form a second connection line between each two second audio feature coordinates, performing traversal operation on the second connection line and the second audio feature coordinates according to the principle of minimum area to output a second region corresponding to the second audio energy sequence segment, and marking an audio energy value corresponding to a center point of the second region as a target second audio energy value corresponding to the second audio energy sequence segment, wherein each edge of the second region belongs to a second connection line, and vertices included in the second region coincide with the plurality of second audio feature coordinates; and According to the target first audio energy value corresponding to each of the first audio energy sequence segments, a set construction operation is performed to output an ordered set of first audio energy values. According to the target second audio energy value corresponding to each of the second audio energy sequence segments, a set construction operation is performed to output an ordered set of second audio energy values. According to the difference between the target first audio energy value and the target second audio energy value corresponding to the set position, a similarity calculation operation is performed on the ordered set of first audio energy values and the ordered set of second audio energy values to output the audio similarity between the two candidate user audio segments to be identified.

Citation Information

Patent Citations

  • Operation safety control system

    CN107103204A

  • Speech recognition processing method and device and storage medium

    CN112053692A

  • Voiceprint recognition method and device, electronic equipment and storage medium

    CN114141252A