A voice-regulated surgical microscope focusing method and system
Through speech sample analysis and similarity calculation, the target command is determined and the focal length of the surgical microscope is adjusted, which solves the problems of cumbersome focus operation and difficulty in operating in a sterile environment, and improves surgical efficiency and safety.
Patent Information
- Application Number
- CN202410991892.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-07-23
AI Technical Summary
When performing fundus surgery, the focus operation of the surgical microscope needs to be performed through the foot switch, which causes the doctor to frequently operate multiple foot switches under a narrow surgical bed, which affects the surgical efficiency and is difficult to operate accurately in a sterile operating environment.
The surgical microscope focusing method is adopted with a speech adjustment method. By obtaining the speech sample of the target user, the energy of the speech sample and the energy of the pre-stored target sample are calculated, the target integrity and the target speech sample are determined, and the similarity between the speech sample and the target speech sample is calculated, the target command is determined, and the focal length step adjustment of the surgical microscope is performed based on the target command.
Microscopic focus operation is carried out through voice commands, which improves surgical efficiency and safety, reduces the operating burden of multiple foot switches under the surgical bed, and ensures accurate operation in a sterile environment.
Smart Images

Figure CN119002026B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microscopes, and in particular to a voice-regulated focusing method and system for a surgical microscope. Background Art
[0002] A non-contact wide-angle lens system is used in conjunction with an ophthalmic surgical microscope. The focus of the microscope is electrically adjusted and requires a foot switch to trigger. Usually, there are two foot switch configurations: an independent foot switch and a foot switch that is shared with the surgical microscope.
[0003] When using a surgical microscope to perform fundus surgery, in addition to the microscope that must be adjusted in real time using the microscope foot switch, there are often foot switches for other treatment equipment, which will further increase the number of foot switches under the doctor's feet, further adding to the burden on the already narrow space under the operating table and between the doctor's chair. In addition, the operation is a sterile operation, and the relevant area is covered with a sterile sheet. The doctor cannot easily see the situation under his feet, and the foot switch can only be replaced based on advance observation and the sense of the field of view, which affects the efficiency of the operation. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for focusing a surgical microscope regulated by voice to improve the above problems. In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0005] In a first aspect, the present application provides a method for focusing a surgical microscope using voice control, comprising:
[0006] Obtain a voice sample of the target user;
[0007] Calculating the energy of the speech sample and the energies of a plurality of pre-stored target samples to obtain a target completeness and a corresponding target speech sample, wherein the target completeness is the most suitable completeness among a plurality of completenesses calculated from the speech sample and the plurality of target samples, and the target speech sample is the target sample corresponding to the calculated target completeness;
[0008] When the target completeness is greater than a first set threshold, calculating the similarity between the speech sample and the target speech sample to obtain a target similarity;
[0009] When the target similarity is greater than a second set threshold, determining a target instruction from the voice sample;
[0010] The focal length step length of the surgical microscope is adjusted accordingly based on the target instruction.
[0011] In a second aspect, the present application also provides a voice-adjustable surgical microscope focusing system, comprising:
[0012] A first acquisition unit, used to acquire a voice sample of a target user;
[0013] A first calculation unit is used to calculate the energy of the speech sample and the energies of multiple target samples stored in advance to obtain a target completeness and a corresponding target speech sample, wherein the target completeness is the most suitable completeness among multiple completenesses calculated from the speech sample and multiple target samples, and the target speech sample is the target sample corresponding to the calculated target completeness;
[0014] A similarity calculation unit, configured to calculate the similarity between the speech sample and the target speech sample to obtain a target similarity when the target completeness is greater than a first set threshold;
[0015] A first determining unit, configured to determine a target instruction from the voice sample when the target similarity is greater than a second set threshold;
[0016] An adjustment unit is used to adjust the focal length step of the surgical microscope accordingly based on the target instruction.
[0017] The beneficial effects of the present invention are:
[0018] The present invention performs focusing operations on the microscope by means of voice commands, and in order to prevent personnel other than the surgeon from accidentally issuing related control commands and affecting the normal operation of the microscope, the voiceprint in the voice command is recognized to ensure that the corresponding focusing operation can be successfully recognized only under the voice command of the setting personnel, which greatly improves the convenience and safety of the microscope focusing operation.
[0019] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or be understood by implementing the embodiments of the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 It is a schematic flow chart of the method for focusing a surgical microscope by voice adjustment described in an embodiment of the present invention;
[0022] Figure 2It is a schematic structural diagram of a surgical microscope focusing system adjusted by voice according to an embodiment of the present invention;
[0023] Markings in the figure: 10, first acquisition unit; 20, first calculation unit; 30, similarity calculation unit; 40, first determination unit; 50, adjustment unit. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0025] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0026] Embodiment 1:
[0027] This embodiment provides a voice-regulated focusing method for a surgical microscope.
[0028] See also Figure 1 , the figure shows that the method includes step S10, step S20, step S30, step S40 and step S50.
[0029] Step S10. Obtain a voice sample of the target user;
[0030] Step S20. Calculate the energy of the speech sample and the energy of multiple target samples stored in advance to obtain a target completeness and a corresponding target speech sample, wherein the target completeness is the most suitable completeness among multiple completenesses calculated from the speech sample and the multiple target samples, and the target speech sample is the target sample corresponding to the calculated target completeness;
[0031] Specifically, each target user with voice control permission has a corresponding standard sample stored. It is necessary to verify whether the current voice-controlled user is the target user by comparing the relationship between the current sample and the standard sample, thereby ensuring the security of the control operation.
[0032] Under normal circumstances, when the same user issues the same voice command, the duration and energy of the voice sample should be roughly the same. Since there are multiple standard samples for different control instructions, it is necessary to determine the standard sample that may correspond to the current voice sample from multiple preset standard samples by judging the voice duration and energy size.
[0033] Specifically, step S20 specifically includes step S21, step S22, step S23, step S24, step S25, step S26, step S27 and step S28:
[0034] Step S21. Obtain a corresponding first waveform image based on the speech sample, wherein the first waveform image is an image with time as the horizontal axis and amplitude as the vertical axis;
[0035] Step S22. Based on the multiple target samples, a corresponding plurality of second waveform graphs are obtained, where the second waveform graph is an image with time as the horizontal axis and amplitude as the vertical axis;
[0036] Step S23. Obtain the speech cut-off time in the first waveform and the plurality of second waveforms to obtain the first cut-off time and the plurality of second cut-off times;
[0037] Step S24. Calculate the difference between the first deadline and the second deadline, and determine a plurality of initial samples based on the difference result, where the initial samples are target samples corresponding to the difference result within a preset range;
[0038] Specifically, the duration of the speech sample is first considered, and a standard sample with a time difference of less than one second from the current speech sample is determined from a plurality of preset standard samples as an initial sample, and subsequent energy size comparisons are continued.
[0039] Step S25. Calculate the energy of the speech sample based on the first waveform to obtain a first energy;
[0040] Step S26. Calculate the energy of multiple initial samples based on the second waveform to obtain multiple second energies;
[0041] Step S27. Calculate the ratio of the first energy to the plurality of second energies to obtain a plurality of completenesses;
[0042] Step S28. Determine the most suitable completeness from multiple completenesses as the target completeness, and use the initial sample corresponding to the most suitable completeness as the target speech sample;
[0043] Specifically, the energy of a speech sample is an important feature of speech, reflecting the intensity or power of a speech signal in a specific time period. For the same user, in the absence of strong emotional fluctuations, the energy of the same speech is basically the same.
[0044] The energy calculation formula of speech samples is:
[0045]
[0046] Where E is the energy of the speech sample; x i is the amplitude value of the i-th sampling point in the waveform graph; n is the total number of sampling points in the waveform graph.
[0047] Calculate the ratio of the speech sample energy to the initial sample energy. Usually, the energy difference between the same speech samples is small, and the obtained ratio result fluctuates around the value 1. From multiple ratios, determine the ratio result closest to the value 1 as the target completeness. The sample corresponding to the target completeness is the target speech sample.
[0048] Step S30. When the target completeness is greater than the first set threshold, the similarity between the speech sample and the target speech sample is calculated to obtain the target similarity;
[0049] Specifically, step S30 specifically includes step S31, step S32, step S33 and step S34:
[0050] Step S31. Divide the speech sample into a plurality of first syllable units based on the pauses of the target user when uttering the speech sample;
[0051] Specifically, in the embodiments of the present application, only Chinese speech is used as an example. However, in actual use, recognition of other languages may be involved. The language characteristics considered in the present application may also be extended to other languages. There is no special limitation here. For example, the Chinese speech of "adjust upward" can be divided into four syllable units of "toward", "up", "adjust" and "rectify" based on the pauses between the pronunciations of each word.
[0052] Step S32. Calculate the difference between the number of the first syllable unit and the number of syllable units preset in the target speech sample to obtain a syllable difference;
[0053] Specifically, due to different pronunciation habits, different people may use different pause methods for the same sentence, which may lead to connected speech. Therefore, the speech can be divided into multiple syllable units based on the pause points during pronunciation. The number of syllable units can be used for preliminary screening to determine whether the user currently pronouncing the sentence is the target user.
[0054] Step S33. When the syllable difference is less than the third set threshold, calculate the local similarity between each first syllable unit and the corresponding syllable unit in the target voice sample respectively to obtain multiple local similarity values;
[0055] Specifically, when the difference calculated by the syllable unit is greater than the set threshold, it can be considered that the current user is not the target user himself, and the current voice sample is discarded; when the difference of the syllable unit is less than the set threshold, it is necessary to further calculate the local similarity between the current voice sample and the target voice sample, so as to make an accurate determination as to whether the user of the current control operation is the target user himself. The third set threshold is usually set to 1 and can be set correspondingly based on the performance of the voice acquisition device, and no special restrictions are made here.
[0056] Specifically, step S33 specifically includes step S331, step S332, step S333, step S334, step S335, step S336, step S337, step S338, step S339, step S3310, step S3311 and step S3312:
[0057] Step S331. Divide the multiple first syllable units based on the tone characteristics to obtain multiple second syllable units. The tone characteristics are used to split the syllables in the first syllable unit that do not use the same tone;
[0058] Specifically, considering that each Chinese character has different tones and there are differences in the tone readings of the same Chinese character by different people, in order to accurately compare the pronunciation characteristics of syllable units, it is necessary to split them again based on intonation characteristics. For example, for the voice of "adjust upward", since the tone of "整" is the third tone, "整" needs to be split. According to the change of intonation, it is split into two parts before and after, and five syllable units of "向", "上", "调", "整(前部分)", and "整(后部分)" can be obtained.
[0059] Step S332. Sort all the second syllable units in chronological order to obtain the first sorting;
[0060] Step S333. Sort the syllable units corresponding to the second syllable units in the target voice sample correspondingly to obtain the second sorting;
[0061] Step S334. Based on a preset pseudo-random number calculation formula, determine some syllable units from the first sorting and the second sorting to obtain the initial syllable units;
[0062] Specifically, for the split second syllable units, the ratio between pronunciation times also needs to be considered.
[0063] Specifically, step S334 specifically includes step S3341, step S3342, step S3343, step S3344, step S3345, step S3346 and step S3347:
[0064] Step S3341. Obtain the number of second syllable units to obtain the number of syllables;
[0065] Step S3342. Calculate the product of the number of syllables and a preset penalty factor to obtain a first product;
[0066] Step S3343. Calculate the difference between the preset maximum number of syllables and the number of syllables to obtain a first difference;
[0067] Step S3344. Calculate the product of the first difference and the penalty factor to obtain the second product
[0068] Step S3345. Calculate the difference between the first cut-off time and the second product to obtain a second difference;
[0069] Step S3346. Generate multiple pseudo-random values within a numerical range between the first product and the second difference based on the random function, where the number of pseudo-random values is the value generated for the first time by the random function;
[0070] Step S3347. Determine some syllable units from the first sort and the second sort based on the pseudo-random value to obtain an initial syllable unit;
[0071] Specifically, the pseudo-random number calculation formula is:
[0072] S=radmon([(n*i),t-(mn)*i])
[0073] Among them, radmon() is a random function, n is the number of syllables; i is the penalty factor; t is the first cutoff time; m is the maximum number of syllables; and S is the calculated pseudo-random value.
[0074] The number of initial syllable units is determined by a numerical value randomly generated by a random function. In an embodiment of the present application, a first pseudo-random numerical value is first generated by a random function. The first pseudo-random numerical value determines the number of initial syllables to be selected. Subsequently, the number of pseudo-random numbers corresponding to the first pseudo-random numerical value is obtained, and the syllable units corresponding to the numerical values are selected in order from the first sort and the second sort to form the initial syllable units.
[0075] Step S335. Calculate the time correlation between the initial syllable units to obtain the syllable correlation;
[0076] Specifically, step S335 specifically includes step S3351, step S3352, step S3353, step S3354 and step S3355:
[0077] Step S3351. Determine a plurality of first units from the first sorting based on the pseudo-random value;
[0078] Step S3352. Determine a plurality of second units from the second sorting based on the pseudo-random value;
[0079] Step S3353. Calculate the ratio of the sounding time of the corresponding first unit and the second unit to obtain a plurality of first values;
[0080] Step S3354. Sum all first values to obtain a second value.
[0081] Step S3355. Calculate the ratio of the second value to the number of pseudo-random values to obtain the syllable correlation;
[0082] Specifically, it is assumed that when a user issues a command, regardless of the volume of the voice, the pronunciation habit remains unchanged, and the pronunciation time for each word is basically the same, so the time ratio between the selected initial syllable units is roughly the same.
[0083] Step S336. When the syllable correlation is greater than the fourth set threshold, the pronunciation time of each second syllable unit is determined based on the first waveform of the sample unit;
[0084] Step S337. Determine the frequency value and envelope value of each second syllable unit in the speech sample based on Fourier transform and Hilber transform;
[0085] Step S338. Randomly determine a plurality of calculation points on the second syllable unit, and determine the frequency change speed, frequency change acceleration, envelope change speed and envelope change acceleration of each calculation point;
[0086] Step S339. Calculate the first difference, the second difference, the third difference and the fourth difference between each calculation point on the plurality of second syllable units and the corresponding calculation point in the target speech sample, the respective frequency change speed, the respective frequency change acceleration, the respective envelope change speed and the respective envelope change acceleration based on the dynamic time warping algorithm;
[0087] Step S3310. Determine a first local difference of the corresponding second syllable unit based on the first difference and the second difference;
[0088] Step S3311. Determine a second local difference of the corresponding second syllable unit based on the third difference and the fourth difference;
[0089] Step S3312: Determine the local similarity based on the first local differences and the second local differences of all second syllable units;
[0090] Specifically, when the user issues a command, the frequency change and envelope change of each syllable unit are basically the same, and the corresponding frequency change speed, frequency change acceleration, envelope change speed and envelope change acceleration are also basically the same. If someone deliberately imitates the pronunciation of the target user, the pronunciation time is usually roughly the same, but there are large differences in the frequency change parameters on each syllable unit and the envelope change parameters at certain points. Therefore, the frequency and envelope change states at each calculation point can be used to accurately determine whether the user who currently issues the voice control command is the target user.
[0091] Step S34. Sum multiple local similarity values to obtain target similarity;
[0092] Specifically, after obtaining the local similarities of multiple syllable units, the overall similarity of the speech sample can be determined by summing up the similarities, thereby determining the difference between the speech sample and the target speech sample.
[0093] Step S40. When the target similarity is greater than a second set threshold, determining the target instruction from the voice sample;
[0094] Specifically, when the similarity between the voice sample and the target voice sample reaches a set threshold, it is considered that the current voice command is issued by the target user, and the command in the voice sample needs to be further determined so as to execute the command operation.
[0095] Specifically, step S40 specifically includes step S41, step S42, step S43 and step S44:
[0096] Step S41. De-noising the speech sample to obtain a first sample;
[0097] Step S42: normalize the first sample to obtain a second sample;
[0098] Step S43. Extract the feature vector of the second sample to obtain the target feature vector;
[0099] Step S44: Input the target feature vector into the trained multilingual recognition model to obtain the target instruction;
[0100] Specifically, in the process of recognizing speech samples, preprocessing is first performed, and the preprocessing step includes denoising and normalizing the speech samples. Denoising usually uses a filter to remove background noise in the audio, and normalization is used to adjust the amplitude of the audio signal to be within a standard range.
[0101] Parameters that can reflect the speech characteristics are extracted from the preprocessed speech sample as feature vectors, and the extracted feature vectors are input into the trained multilingual recognition model to recognize the instructions corresponding to the speech, wherein the multilingual recognition model refers to a recognition model that can recognize multiple languages.
[0102] Step S50. Adjusting the focal length of the surgical microscope according to the target instruction;
[0103] Specifically, usually, in order to ensure the convenience of command-based adjustment of the microscope, some fixed control commands are set, and different control commands correspond to different focal length adjustment amounts.
[0104] Regarding the control instruction setting and the focal length adjustment amount corresponding to each control instruction, corresponding settings can be made according to the parameters of the microscope itself. The specific setting steps include step S01, step S02, step S03 and step S04:
[0105] Step S01. Acquire parameter information of the surgical microscope, and calculate the geometric depth of field and the physical depth of field based on the parameter information;
[0106] Step S02. Determine the maximum focal length range based on the geometric depth of field, the physical depth of field and the preset adjustment depth of field;
[0107] Step S03: dividing the maximum focusing range based on a preset step length to obtain a plurality of step lengths;
[0108] Step S04. Determine the number of instructions and the step length corresponding to each instruction based on the step length;
[0109] Specifically, the total magnification of the universal non-contact wide-angle lens system considered in this patent is usually around 20 times, the magnification during normal operation is around 10 times, the numerical aperture of the microscope is around 0.05, and through the relevant parameters of the microscope, it is calculated that the geometric depth of field of the surgical microscope is approximately 0.72 mm, and the physical depth of field of the surgical microscope is approximately 0.22 mm. The geometric depth of field and physical depth of field of the microscope together are approximately 0.94 mm, that is, the geometric depth of field and physical depth of field of the surgical microscope on the observing eye side together are equivalent to approximately 1.5D depth of field.
[0110] The total depth of field of the entire surgical microscope visual system is the sum of the accommodative depth of field formed by the human eye's refractive accommodation ability, the geometric depth of field formed by the optical characteristics of the microscope, and the physical depth of field. The values of the human eye's refractive accommodation ability that change with age are shown in Table 1. Taking into account the age differences of surgeons, the total depth of field is approximately 2.5D-6.0D.
[0111] Table 1
[0112] Age (years) 30 35 40 45 50 55 60 Adjustment range 7.0D 5.5D 4.5D 3.5D 2.5D 1.75D 1D
[0113] The maximum focusing range of the universal non-contact wide-angle lens system considered in this patent is 20D, and the entire focal length adjustment range can be divided into 8 steps, each step corresponding to a 2.5D adjustment amount.
[0114] The voice commands and corresponding adjustment steps set in the embodiment of the present application are shown in Table 2, in which only Chinese commands and English commands are identified. The other language commands can be obtained according to the actual language environment, and there is no special limitation here.
[0115] Table 2
[0116]
[0117] The voice commands designed in this application are all voice commands of three or more syllables, and there is basically no overlap between the text and common words, which can effectively avoid confusion and misrecognition. In actual usage environments, more voice commands can be set, and there is no special restriction here.
[0118] Among them, the "previous step" instruction and the "previous two steps" instruction refer to moving the focusing mirror to the upper half of the corresponding step length.
[0119] The "next step" and "next two steps" instructions refer to moving the focusing mirror to the corresponding step length to the lower half.
[0120] The "center center" command means to move the focusing lens to the center of the focusing range.
[0121] The specific operation of the upper half scan corresponding to the "U (optimal) scan" instruction is that when the focusing mirror is in the center position, it runs at high speed from the center position to the upper limit, and then returns to the center position. If the focusing mirror is not currently in the center position, the centering operation is performed first, and then it runs at high speed from the center position to the upper limit, and then returns to the center position.
[0122] The specific operation of the second half scan corresponding to the "D (ground) scan" instruction is that when the focusing mirror is in the center position, it runs at high speed from the center position to the lower limit, and then returns to the center position. If the focusing mirror is not currently in the center position, the centering operation is performed first, and then it runs at high speed from the center position to the lower limit, and then returns to the center position.
[0123] The human eye has the ability to capture information at high speed. In actual microscope operation, the upper half of the scan can be performed first, allowing the doctor to quickly screen the patient's fundus in sequence. If the best focus is in the upper half, the surgeon can know the approximate position of the best clarity that the fundus image of the patient can achieve, providing a basis for subsequent precise focusing. If the best focus is not found in the upper half of the scan, the second half of the scan will be performed until the best focus is determined.
[0124] Based on the clear image state at the best focus obtained in the first half of the scan, if it is determined that the current center position is close to the best state, a "previous step" voice command is issued, and the position of the best focus point continues to be observed during the adjustment of the focusing lens; if it is determined that the current center is far from the best state, a "two steps up" voice command is issued to move the focusing lens to the 1 / 2 position of the first half. If the best focus state is observed during the operation process, after stopping the action, a "next step" voice command is issued to reach the best focus position and end the focusing.
[0125] If after issuing the "two steps up" voice command, the image gradually becomes clearer but has not yet reached the best focus point, the clear image status at the best focus point obtained in the upper half of the scan can be used to determine approximately how many steps of focus are needed. Depending on the situation, issue the "one step up" voice command or the "two steps up" voice command. During the entire adjustment process, the image clarity status at the best focus point obtained through the upper half of the scan provides an important basis for subsequent focusing.
[0126] When the best focus point is in the second half, you can follow the above steps and use the "next step" voice command and "next two steps" voice command to adjust the focus until the best focus point is reached.
[0127] Embodiment 2:
[0128] like Figure 2 As shown, this embodiment provides a voice-adjustable surgical microscope focusing system, the system comprising:
[0129] A first acquisition unit 10 is used to acquire a voice sample of a target user;
[0130] A first calculation unit 20 is used to calculate the energy of the speech sample and the energy of multiple target samples stored in advance to obtain a target completeness and a corresponding target speech sample, wherein the target completeness is the most suitable completeness among multiple completenesses calculated from the speech sample and the multiple target samples, and the target speech sample is the target sample corresponding to the calculated target completeness;
[0131] A similarity calculation unit 30, configured to calculate the similarity between the speech sample and the target speech sample to obtain a target similarity when the target completeness is greater than a first set threshold;
[0132] A first determining unit 40 is used to determine a target instruction from the voice sample when the target similarity is greater than a second set threshold;
[0133] The adjustment unit 50 is used to adjust the focal length step of the surgical microscope according to the target instruction.
[0134] In a specific implementation disclosed in the present application, the first computing unit 20 includes:
[0135] A first obtaining unit is used to obtain a corresponding first waveform image based on the speech sample, where the first waveform image is an image with time as the horizontal axis and amplitude as the vertical axis;
[0136] A second obtaining unit is used to obtain a plurality of corresponding second waveform graphs based on the plurality of target samples, wherein the second waveform graph is an image with time as the horizontal axis and amplitude as the vertical axis;
[0137] A second acquisition unit is used to acquire the speech cut-off time in the first waveform and the plurality of second waveforms to obtain the first cut-off time and the plurality of second cut-off times;
[0138] A second calculation unit is used to calculate the difference between the first deadline and the second deadline, and determine a plurality of initial samples based on the difference result, where the initial samples are target samples corresponding to the difference result within a preset range;
[0139] A third obtaining unit is used to calculate the energy of the speech sample based on the first waveform diagram to obtain a first energy;
[0140] a fourth obtaining unit, configured to calculate the energies of the plurality of initial samples based on the second waveform diagram to obtain a plurality of second energies;
[0141] A third calculation unit, used for calculating the ratio of the first energy to the plurality of second energies to obtain a plurality of completenesses;
[0142] The second determining unit is used to determine the most suitable completeness from multiple completenesses as the target completeness, and use the initial sample corresponding to the most suitable completeness as the target speech sample.
[0143] In a specific implementation disclosed in the present application, the similarity calculation unit 30 includes:
[0144] A first division unit, configured to divide the speech sample into a plurality of first syllable units based on a pause of the target user when uttering the speech sample;
[0145] a fourth calculation unit, configured to calculate a difference between the number of the first syllable units and the number of syllable units preset in the target speech sample to obtain a syllable difference;
[0146] a fifth calculation unit, configured to calculate the local similarity between each first syllable unit and the corresponding syllable unit in the target speech sample when the syllable difference is less than a third set threshold value, to obtain a plurality of local similarity values;
[0147] The first summing calculation unit is used to perform summing calculation on multiple local similarity values to obtain target similarity.
[0148] In a specific implementation manner disclosed in the present application, the fifth computing unit includes:
[0149] A second division unit, configured to divide the plurality of first syllable units into a plurality of second syllable units based on tone characteristics, wherein the tone characteristics are used to split syllables in the first syllable units that do not use the same tone;
[0150] A first sorting unit, used for sorting all the second syllable units based on time sequence to obtain a first sorting;
[0151] A second sorting unit is used to sort the syllable units corresponding to the second syllable unit in the target speech sample to obtain a second sorting;
[0152] A third determining unit is used to determine some syllable units from the first sorting and the second sorting based on a preset pseudo-random number calculation formula to obtain an initial syllable unit;
[0153] a sixth calculation unit, used to calculate the time correlation between the initial syllable units to obtain the syllable correlation;
[0154] a seventh calculation unit, configured to determine the pronunciation time of each second syllable unit based on the first waveform of the sample unit when the syllable correlation is greater than a fourth set threshold;
[0155] A fourth determining unit, configured to determine a frequency value and an envelope value of each second syllable unit in the speech sample based on Fourier transform and Hilber transform;
[0156] A random determination unit, used to randomly determine a plurality of calculation points on the second syllable unit, and determine the frequency change speed, frequency change acceleration, envelope change speed and envelope change acceleration of each calculation point;
[0157] an eighth calculation unit, for respectively calculating, based on a dynamic time warping algorithm, a first difference, a second difference, a third difference, and a fourth difference between a frequency change speed, a frequency change acceleration, an envelope change speed, and an envelope change acceleration of each calculation point on the plurality of second syllable units and a corresponding calculation point in the target speech sample;
[0158] A ninth calculation unit, configured to determine a first local difference of the corresponding second syllable unit based on the first difference and the second difference;
[0159] a fifth determining unit, configured to determine a second local difference of the corresponding second syllable unit based on the third difference and the fourth difference;
[0160] The sixth determining unit is configured to determine the local similarity based on the first local differences and the second local differences of all the second syllable units.
[0161] In a specific implementation manner disclosed in the present application, the third determining unit includes:
[0162] A third obtaining unit is used to obtain the number of the second syllable units to obtain the number of syllables;
[0163] a tenth calculation unit, configured to calculate the product of the number of syllables and a preset penalty factor to obtain a first product;
[0164] an eleventh calculating unit, configured to calculate a difference between a preset maximum number of syllables and the number of syllables to obtain a first difference;
[0165] A twelfth calculation unit is used to calculate the product of the first difference and the penalty factor to obtain a second product
[0166] a thirteenth calculating unit, configured to calculate a difference between the first cut-off time and the second product to obtain a second difference;
[0167] A generating unit, configured to generate a plurality of pseudo-random values within a numerical range between the first product and the second difference based on a random function, wherein the number of the pseudo-random values is a value generated for the first time by the random function;
[0168] The seventh determining unit is used to determine some syllable units from the first sorting and the second sorting based on the pseudo-random value to obtain an initial syllable unit.
[0169] In a specific implementation disclosed in the present application, the sixth calculation unit includes:
[0170] an eighth determining unit, configured to determine a plurality of first units from the first sequence based on the pseudo-random value;
[0171] a ninth determining unit, configured to determine a plurality of second units from the second sorting based on the pseudo-random value;
[0172] A fourteenth calculation unit, used to calculate the ratio of the sounding time of the corresponding first unit and the second unit to obtain a plurality of first values;
[0173] The second summing calculation unit is used to sum all the first values to obtain a second value.
[0174] The fifteenth calculation unit is used to calculate the ratio of the second numerical value to the number of pseudo-random numerical values to obtain the syllable correlation.
[0175] In a specific implementation disclosed in the present application, the first determining unit 40 includes:
[0176] A denoising processing unit, used for performing denoising processing on the speech sample to obtain a first sample;
[0177] A normalization unit, used for performing normalization processing on the first sample to obtain a second sample;
[0178] An extraction unit, used for extracting a feature vector of the second sample to obtain a target feature vector;
[0179] The input unit is used to input the target feature vector into the trained multilingual recognition model to obtain the target instruction.
[0180] In a specific embodiment disclosed in the present application, the system includes:
[0181] A fourth acquisition unit, used to acquire parameter information of the surgical microscope, and calculate a geometric depth of field and a physical depth of field based on the parameter information;
[0182] a tenth determining unit, configured to determine a maximum focal length range based on a geometric depth of field, a physical depth of field, and a preset adjusted depth of field;
[0183] A third dividing unit, configured to divide the maximum focusing range based on a preset step length to obtain a plurality of step lengths;
[0184] The eleventh determining unit is used to determine the number of instructions and the step length corresponding to each instruction based on the step length.
[0185] It should be noted that, regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0186] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0187] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for focusing a surgical microscope using voice control, characterized in that: include: Obtain voice samples of target users; Calculating the energy of the speech sample and the energies of a plurality of pre-stored target samples to obtain a target completeness and a corresponding target speech sample, wherein the target completeness is the most suitable completeness among a plurality of completenesses calculated from the speech sample and the plurality of target samples, and the target speech sample is the target sample corresponding to the calculated target completeness; When the target completeness is greater than a first set threshold, calculating the similarity between the speech sample and the target speech sample to obtain a target similarity; When the target similarity is greater than a second set threshold, determining a target instruction from the voice sample; Based on the target instruction, the surgical microscope is adjusted in a corresponding focal length step; The energy of the speech sample and the energy of a plurality of pre-stored target samples are calculated to obtain a target integrity and a corresponding target speech sample, including: Based on the voice sample, a corresponding first waveform image is obtained, wherein the first waveform image is an image with time as the horizontal axis and amplitude as the vertical axis; Based on the multiple target samples, a corresponding plurality of second waveform graphs are obtained, wherein the second waveform graph is an image with time as the horizontal axis and amplitude as the vertical axis; Acquire the speech cut-off time in the first waveform and the plurality of second waveforms to obtain a first cut-off time and a plurality of second cut-off times; Calculating a difference between the first deadline and the second deadline, and determining a plurality of initial samples based on the difference result, wherein the initial samples are target samples corresponding to the difference result within a preset range; Calculate the energy of the speech sample based on the first waveform to obtain a first energy; Calculate the energy of a plurality of initial samples based on the second waveform diagram to obtain a plurality of second energies; Calculating a ratio of the first energy to a plurality of second energies to obtain a plurality of completenesses; Determine the most suitable completeness from multiple completenesses as the target completeness, and use the initial sample corresponding to the most suitable completeness as the target speech sample; Wherein, when the target completeness is greater than a first set threshold, calculating the similarity between the speech sample and the target speech sample to obtain the target similarity includes: Dividing the speech sample into a plurality of first syllable units based on a pause of the target user when uttering the speech sample; Calculating a difference between the number of the first syllable units and the number of syllable units preset in the target speech sample to obtain a syllable difference; When the syllable difference is less than a third set threshold, respectively calculating the local similarity between each first syllable unit and the same syllable unit corresponding to the target speech sample to obtain a plurality of local similarity values; The target similarity is obtained by summing up the multiple local similarity values.
2. The method for focusing a surgical microscope using voice control according to claim 1, characterized in that When the syllable difference is less than the third set threshold, the local similarity between each first syllable unit and the corresponding syllable unit in the target speech sample is calculated respectively to obtain multiple local similarity values, including: Dividing the plurality of first syllable units into a plurality of second syllable units based on tone characteristics, wherein the tone characteristics are used to split syllables in the first syllable units that do not use the same tone; Sorting all second syllable units based on time order to obtain a first sorting; Sorting the syllable units corresponding to the second syllable unit in the target speech sample to obtain a second sorting; Based on a preset pseudo-random number calculation formula, a partial syllable unit is determined from the first sorting and the second sorting to obtain an initial syllable unit; Calculating the time correlation between the initial syllable units to obtain syllable correlation; When the syllable correlation is greater than a fourth set threshold, determining the pronunciation time of each second syllable unit based on the first waveform of the speech sample; Determine the frequency value and envelope value of each second syllable unit in the speech sample based on Fourier transform and Hilber transform; Randomly determine a plurality of calculation points on the second syllable unit, and determine the frequency change speed, frequency change acceleration, envelope change speed and envelope change acceleration of each calculation point; Based on the dynamic time warping algorithm, respectively calculate the first difference, the second difference, the third difference and the fourth difference between each calculation point on the plurality of second syllable units and the corresponding calculation point in the target speech sample, the respective frequency change speed, the respective frequency change acceleration, the respective envelope change speed and the respective envelope change acceleration; Determine a first local difference of the corresponding second syllable unit based on the first difference and the second difference; Determine a second local difference of the corresponding second syllable unit based on the third difference and the fourth difference; The local similarity value is determined based on the first local differences and the second local differences of all second syllable units.
3. The method for focusing a surgical microscope using voice control according to claim 2, characterized in that , based on a preset pseudo-random number calculation formula, determining some syllable units from the first sorting and the second sorting to obtain an initial syllable unit, including: Obtain the number of the second syllable units to obtain the number of syllables; Calculating the product of the number of syllables and a preset penalty factor to obtain a first product; Calculating a difference between a preset maximum number of syllables and the number of syllables to obtain a first difference; Calculate the product of the first difference and the penalty factor to obtain a second product Calculating a difference between the first cut-off time and the second product to obtain a second difference; Generate a plurality of pseudo-random values within a numerical range between the first product and the second difference based on a random function, wherein the number of the pseudo-random values is the value generated for the first time by the random function; Based on the pseudo-random value, some syllable units are determined from the first sorting and the second sorting to obtain the initial syllable unit.
4. The method for focusing a surgical microscope using voice control according to claim 1, characterized in that Before obtaining the voice sample of the target user, the method includes: Acquiring parameter information of the surgical microscope, and calculating a geometric depth of field and a physical depth of field based on the parameter information; Determining a maximum focal length range based on the geometric depth of field, the physical depth of field, and a preset adjusted depth of field; Dividing the maximum focal length range based on a preset step length to obtain a plurality of step lengths; The number of instructions and the step size corresponding to each instruction are determined based on the step size.
5. The method for focusing a surgical microscope using voice control according to claim 3, characterized in that , the pseudo-random number calculation formula is: S=radmon([(n*i),t-(mn)*i]) Among them, radmon() is a random function, n is the number of syllables; i is the penalty factor; t is the first cutoff time; m is the maximum number of syllables; and S is the calculated pseudo-random value.
6. A voice-controlled focusing system for a surgical microscope, characterized in that: include: A first acquisition unit, used to acquire a voice sample of a target user; A first calculation unit is used to calculate the energy of the speech sample and the energies of multiple target samples stored in advance to obtain a target completeness and a corresponding target speech sample, wherein the target completeness is the most suitable completeness among multiple completenesses calculated from the speech sample and multiple target samples, and the target speech sample is the target sample corresponding to the calculated target completeness; A similarity calculation unit, configured to calculate the similarity between the speech sample and the target speech sample to obtain a target similarity when the target completeness is greater than a first set threshold; A first determining unit, configured to determine a target instruction from the voice sample when the target similarity is greater than a second set threshold; An adjustment unit, used for adjusting the focal length step of the surgical microscope according to the target instruction; Wherein, the first calculation unit includes: A first obtaining unit, configured to obtain a corresponding first waveform image based on the speech sample, wherein the first waveform image is an image with time as the horizontal axis and amplitude as the vertical axis; A second obtaining unit is used to obtain a plurality of corresponding second waveform graphs based on the plurality of target samples, wherein the second waveform graph is an image with time as the horizontal axis and amplitude as the vertical axis; A second acquisition unit, used to acquire the speech cut-off time in the first waveform and the plurality of second waveforms, to obtain a first cut-off time and a plurality of second cut-off times; a second calculation unit, configured to calculate a difference between the first deadline and the second deadline, and determine a plurality of initial samples based on the difference result, wherein the initial samples are target samples corresponding to the difference result within a preset range; A third obtaining unit, configured to calculate the energy of the speech sample based on the first waveform diagram to obtain a first energy; a fourth obtaining unit, configured to calculate the energies of a plurality of initial samples based on the second waveform diagram to obtain a plurality of second energies; A third calculation unit, configured to calculate a ratio of the first energy to a plurality of second energies to obtain a plurality of completenesses; A second determining unit is used to determine the most suitable completeness from multiple completenesses as the target completeness, and use the initial sample corresponding to the most suitable completeness as the target speech sample; Wherein, the similarity calculation unit includes: A first division unit, configured to divide the speech sample into a plurality of first syllable units based on a pause of the target user when uttering the speech sample; a fourth calculating unit, configured to calculate a difference between the number of the first syllable units and the number of syllable units preset in the target speech sample to obtain a syllable difference; a fifth calculation unit, configured to calculate the local similarity between each first syllable unit and the corresponding same syllable unit in the target speech sample respectively, when the syllable difference value is less than a third set threshold value, to obtain a plurality of local similarity values; The first summation calculation unit is used to perform summation calculation on multiple local similarity values to obtain the target similarity.
Citation Information
Patent Citations
User-defined keyword recognition method and system based on similar pair comparative learning
CN115410552A
Voice activated microscope
US4989253A