Three-dimensional scanning device, control method, computer program product, and storage medium

By integrating voice acquisition and processing modules into 3D scanning equipment and using voice recognition models to identify and train user voice data, the problems of cumbersome user operations and cross-infection are solved, more efficient and safer voice control is achieved, and the user experience is improved.

WO2025195392A1PCT designated stage Publication Date: 2025-09-25SHINING 3D TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/083306
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-19
Filing Date
2025-03-19
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

When using a 3D scanning device, users need to hold the device for scanning and operate the terminal device at the same time, which makes the operation cumbersome and prone to cross infection, especially in scenarios such as oral scanning.

Method used

The voice acquisition module and processing module are integrated into the 3D scanning device, and the voice recognition model is used to identify the user's voice data and convert it into control instructions. Training is carried out in combination with feedback information to improve recognition accuracy, and training samples are automatically constructed to realize voice control functions.

Benefits of technology

It improves the accuracy of the voice control function of 3D scanning equipment, reduces the user operation steps, and enhances the user experience, especially reducing the risk of cross-infection in scenarios such as oral scanning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025083306_25092025_PF_FP_ABST
    Figure CN2025083306_25092025_PF_FP_ABST
Patent Text Reader

Abstract

A three-dimensional scanning device, a control method, a computer program product, and a storage medium. During the process of a user controlling a three-dimensional scanning device by using speech, the three-dimensional scanning device recognizes speech data input by the user, and converts same into a corresponding control instruction, and the user controls the three-dimensional scanning device to execute a corresponding operation. During the process of the three-dimensional scanning device executing the corresponding operation, feedback information of the user is collected, whether the speech data input by the user matches a target control instruction recognized by a speech recognition model is determined on the basis of the feedback information, and a label corresponding to the speech data is then determined on the basis of a matching result, such that a training sample that carries the label is automatically constructed; and the speech recognition model is then continuously trained by using the training sample automatically constructed during speech control, such that the speech recognition model performs continuous learning, thereby improving the accuracy of the speech recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional scanning device, control method, computer program product and storage medium

[0001] This disclosure claims priority to a Chinese patent application filed with the Patent Office of China on March 19, 2024, with application number 202410320872.X, entitled “Three-dimensional scanning device, control method, computer program product and storage medium,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present disclosure relates to the field of three-dimensional scanning technology, and in particular to a three-dimensional scanning device, a control method, a computer program product, and a storage medium. Background Art

[0003] Currently, when users use 3D scanning devices to scan target objects to reconstruct a 3D model of the target object, they usually connect the 3D scanning device to a terminal device installed with scanning software, and then use the scanning software on the terminal device to set some parameters in the scanning process, or issue control instructions to control the 3D scanning device or scanning software.

[0004] During the scanning process, the user must not only hold the 3D scanning device to scan the target object, but also operate the scanning software to issue control commands. This requires the user to manage both the 3D scanning device and the scanning software on the terminal device, which is cumbersome and inconvenient. Therefore, a solution is needed to make it easier for users to control the 3D scanning device or scanning software. Summary of the Invention

[0005] The present disclosure provides a three-dimensional scanning device, a control method, a computer program product, and a storage medium.

[0006] According to a first aspect of an embodiment of the present disclosure, a three-dimensional scanning device is provided, wherein a voice acquisition module and a processing module are provided in the three-dimensional scanning device, the voice acquisition module is configured to acquire voice data of a user, the processing module is configured to recognize the voice data using a preset voice recognition model, parse the voice data to obtain a target control instruction, and control the three-dimensional scanning device to perform a corresponding operation based on the target control instruction, and obtain feedback information from the user regarding the operation during the process of the three-dimensional scanning device performing the corresponding operation, determine whether the target control instruction matches the voice data based on the feedback information, determine a label corresponding to the voice data based on the matching result, and construct a training sample using the voice data and the label corresponding to the voice data to train the voice recognition model using the constructed training sample, wherein the label is used to indicate whether the control instruction corresponding to the voice data is the target control instruction.

[0007] According to a second aspect of an embodiment of the present disclosure, a control method for a three-dimensional scanning device is provided, the method comprising: obtaining voice data input by a user, recognizing the voice data using a preset voice recognition model, obtaining a target control instruction corresponding to the voice data, and controlling the three-dimensional scanning device to perform a corresponding operation based on the target control instruction; while the three-dimensional scanning device is performing the corresponding operation, obtaining feedback information from the user regarding the operation; determining whether the target control instruction matches the voice data based on the feedback information; determining a label corresponding to the voice data based on the matching result; and constructing a training sample using the voice data and the label corresponding to the voice data to train the voice recognition model using the constructed training sample, wherein the label is used to indicate whether the control instruction corresponding to the voice data is the target control instruction.

[0008] According to a third aspect of an embodiment of the present disclosure, a computer program product is provided, wherein the computer program product includes computer instructions, and when the processor executes the computer instructions, the method of the second aspect described above is implemented.

[0009] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored. When the computer instructions are executed, the method of the second aspect described above is implemented.

[0010] In an embodiment of the present disclosure, a three-dimensional scanning device is provided that can be integrated with a voice control function. When a user uses voice to control the three-dimensional scanning device, the three-dimensional scanning device can recognize the voice data input by the user, convert it into corresponding control instructions, and control the three-dimensional scanning device to perform the corresponding operation. Furthermore, during the process of the three-dimensional scanning device performing the corresponding operation, user feedback information can be collected. Based on this feedback information, it is determined whether the voice data input by the user matches the target control instruction recognized by the voice recognition model. Based on the matching result, a label corresponding to the voice data can be determined, thereby automatically constructing training samples with labels. The training samples automatically constructed during the voice control process can then be used to continuously train the voice recognition model, allowing the voice recognition model to continuously learn, thereby improving the accuracy of the voice recognition model. In this way, the accuracy of the voice control function of the three-dimensional scanning device can be improved, and the user experience when using the voice control function of the three-dimensional scanning device can be enhanced.

[0011] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0013] FIG1 is a schematic diagram of a three-dimensional scanning device and scanning software used in combination according to an embodiment of the present disclosure.

[0014] FIG2 is a schematic structural diagram of a three-dimensional scanning device according to an embodiment of the present disclosure.

[0015] FIG3 is a schematic diagram of voice control of a three-dimensional scanning device according to an embodiment of the present disclosure.

[0016] FIG4 is a schematic diagram of verifying a recognized control instruction based on a workflow of a three-dimensional scanning device according to an embodiment of the present disclosure.

[0017] FIG5 is a schematic structural diagram of a three-dimensional scanning device according to an embodiment of the present disclosure.

[0018] FIG6 is a schematic diagram of automatically retrieving pre-stored configuration parameters based on a user's voice features according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0020] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a", "the" and "the" used in this disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items. In addition, the term "at least one" herein means any combination of at least two of any one or more of a plurality of.

[0021] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."

[0022] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present disclosure and to make the above-mentioned purposes, features and advantages of the embodiments of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure are further described in detail below with reference to the accompanying drawings.

[0023] Currently, when users use 3D scanning devices to scan a target object and reconstruct a 3D model of the target object, they typically connect the 3D scanning device to a terminal device installed with scanning software. Because 3D scanning devices are often small and have limited control buttons, the scanning software on the terminal device can be used to set certain parameters during the scanning process or issue control commands to control the 3D scanning device or scanning software. For example, this can control the operation of certain functional modules on the 3D scanning device or the way the scanning software performs 3D reconstruction using data collected by the 3D scanning device.

[0024] For example, as shown in Figure 1, the 3D scanning device is typically connected to a user's computer, which has scanning software installed. The 3D scanning device can send the collected oral scan data to the scanning software on the computer, which can then reconstruct a 3D model of the oral cavity based on the received scan data. Because the 3D scanning device is small and has a limited number of buttons, the scanning software on the computer can be used to set parameters during the scanning process or issue control commands to control the 3D scanning device or scanning software.

[0025] During the scanning process, users must both hold a 3D scanning device to scan the target object and operate the scanning software to issue control commands, making the operation cumbersome and inconvenient. Furthermore, for oral scanning scenarios, if users must manually operate the terminal device (e.g., the user's computer) while also performing operations inside the user's mouth, cross-infection can easily occur due to improper hand washing.

[0026] To facilitate user operation, the applicants considered integrating voice control functionality into 3D scanning devices, allowing users to control the 3D scanning device or scanning software via voice, eliminating the need to manually operate the terminal device to issue control commands during the scanning process. However, in voice control scenarios, accurately recognizing the user's input voice data and converting it into the user's intended control commands to accurately control the 3D scanning device or scanning software is particularly critical, directly affecting the user's experience.

[0027] Based on this, an embodiment of the present disclosure provides a three-dimensional scanning device that can integrate a voice control function. When a user uses voice to control the three-dimensional scanning device, the three-dimensional scanning device can recognize the voice data input by the user, convert it into corresponding control instructions, and control the three-dimensional scanning device to perform the corresponding operation. Furthermore, during the process of the three-dimensional scanning device performing the corresponding operation, user feedback information can be collected. Based on this feedback information, it is determined whether the voice data input by the user matches the target control instruction recognized by the voice recognition model. Then, based on the matching result, a label corresponding to the voice data can be determined, thereby automatically constructing training samples with labels. The training samples automatically constructed during the voice control process can then be used to continuously train the voice recognition model, allowing the voice recognition model to continuously learn, thereby improving the accuracy of the voice recognition model. In this way, the accuracy of the voice control function of the three-dimensional scanning device can be improved, and the user experience when using the voice control function of the three-dimensional scanning device can be enhanced.

[0028] The three-dimensional scanning device in the embodiment of the present disclosure can be an oral scanner, a facial scanner, an industrial scanner, or a professional scanner, which can be used to perform three-dimensional scanning and three-dimensional reconstruction of objects such as teeth, faces, bodies, industrial products, industrial equipment, cultural relics, artworks, prostheses, medical devices, and buildings.

[0029] As shown in Figure 2, in addition to the hardware structure necessary to realize the three-dimensional scanning function (such as lasers, cameras, fill lights, etc.), the three-dimensional scanning device can also include a voice acquisition module 11 and a processing module 12. Among them, the voice acquisition module 11 can be a hardware device with audio capture function such as a microphone. The processing module 12 can be a processing chip or embedded system built into the three-dimensional scanning device, which is mainly configured to control the three-dimensional scanning device. Of course, it is not difficult to understand that for different types of three-dimensional scanning devices, other hardware functional modules can be configured based on their functions, such as indicator lights, buttons, etc., and the embodiments of the present disclosure are not limited thereto.

[0030] In order to accurately recognize and analyze the voice data input by the user, a voice recognition model can be built into the three-dimensional scanning device. The voice recognition model can be a language model with natural language processing capabilities such as a BERT model and a ChatGPT model.

[0031] When the user needs to control the three-dimensional scanning device or scanning software, he can issue instructions through voice. The voice acquisition module 11 can collect the user's voice data and transmit it to the processing module 12. The processing module 12 can use a preset voice recognition model to recognize the voice data, and parse the voice data to obtain target control instructions, and then control the three-dimensional scanning device to perform corresponding operations based on the target control instructions.

[0032] Controlling the 3D scanning device to perform corresponding operations can include controlling components on the 3D scanning device to perform corresponding operations, such as turning on a fill light, turning on a defogger function, adjusting the intensity of the structured light projected by the 3D scanning device, and so on. Controlling the 3D scanning device to perform corresponding operations can also include controlling the scanning software through the 3D scanning device to control the scanning software to adjust the reconstruction method during the 3D reconstruction process. For example, the 3D scanning device can send control instructions to the scanning software so that the scanning software can adjust parameters during the 3D reconstruction process, such as adjusting the enhancement level when enhancing image brightness, adjusting the maximum stitching error allowed during image stitching, adjusting the type and extent of deleted miscellaneous data, and so on.

[0033] As shown in Figure 3, in order for the voice recognition model to accurately recognize the voice data and parse it to obtain the correct control instructions, during the process of the 3D scanning device performing the corresponding operation, the user's feedback information regarding the operation can be obtained, and based on the feedback information, it can be determined whether the parsed target control instructions match the voice data. The user's feedback information can be used to describe some of the user's operational behaviors during the process of the 3D scanning device performing the operation, such as whether the user interrupted the operation or corrected the operation. For example, suppose the user inputs the voice data "turn on the fill light". If the 3D scanning device accurately recognizes the control instruction and turns on the fill light, the user will continue to use the 3D scanning device to scan without interrupting the scanning operation. If the 3D scanning device does not accurately recognize the control instruction at this time, for example, the 3D scanning device does not "turn on the fill light" but turns on the defog function, the user knows that the 3D scanning device does not accurately recognize the control instruction, then the user will most likely interrupt the current scan, or re-issue the control instruction through the buttons on the 3D scanning device or the buttons on the scanning software. At this time, the 3D scanning device can record the user's operation behavior and generate information corresponding to the operation behavior as feedback information, and then determine whether the target control instruction recognized by the 3D scanning device matches the voice data input by the user based on the user's feedback information, and determine the label corresponding to the voice data based on the matching result, wherein the label is used to indicate whether the control instruction corresponding to the voice data is the target control instruction, and then a training sample can be constructed based on the voice data and its label, so that the voice recognition model can be trained using the constructed training sample.

[0034] For example, assuming that the voice data input by the user is voice A (i.e., turning on the fill light), if the voice recognition module recognizes it as target control instruction B (i.e., turning on the fill light), at this time, the three-dimensional scanning device does not detect the user's interruption operation, that is, the user has been performing the scanning operation, then it means that the control instruction corresponding to voice A is instruction B, then a training sample voice A-instruction B can be constructed, and then the training sample can be used to perform supervised training on the voice recognition model.

[0035] In an embodiment of the present disclosure, when a user uses voice to control a three-dimensional scanning device, the three-dimensional scanning device can determine, based on the user's feedback information, whether the voice data input by the user matches the target control instruction recognized by the voice recognition model based on the voice data, and then determine the label corresponding to the voice data based on the matching result, so as to automatically construct a training sample with the label, and then use the training samples automatically constructed during the voice control process to continuously train the voice recognition model, so that the voice recognition model continues to learn, thereby improving the accuracy of the voice recognition model.

[0036] In some embodiments, as shown in FIG4 , considering that when a user uses a three-dimensional scanning device to scan a target object, the three-dimensional scanning device usually works according to a predetermined workflow (e.g., operation A-operation B-operation C), that is, when the user controls the three-dimensional scanning device, under normal circumstances, the user will also issue control instructions according to the instruction flow corresponding to the workflow. However, there are some scenarios in which the voice control instructions (i.e., the input voice data) issued by the user may be incorrect, or the voice recognition model may misrecognize the voice data, resulting in an incorrect parsed control instruction. If the three-dimensional scanning device is controlled according to the incorrect control instruction, an erroneous operation will occur, affecting the scanning efficiency. In order to reduce the above problems, the target control instruction parsed based on the user's current input voice data can be verified in combination with the context information of the user's input voice data to verify the accuracy of the target control instruction. For example, after the processing module 12 obtains the voice data currently input by the user and identifies the target control instruction corresponding to the voice data, it can obtain the previous segment or multiple segments of voice data of the current voice data by using the voice recognition model to identify the historical control instructions, and then predict what the current control instruction should be based on the workflow of the three-dimensional scanning device and the historical control instructions, and then verify the accuracy of the currently parsed target control instruction.

[0037] For example, when a user performs a scanning task using a 3D scanning device, the workflow is typically as follows: create order - start scanning - end scanning. Therefore, if the control instruction recognized based on the user's current voice input is "stop the 3D scanning device's scanning operation," and the control instruction corresponding to the user's previous voice input is "create order," it is clear that "create order - end scanning" does not conform to the 3D scanning device's set workflow. Therefore, it can be determined that the control instruction recognized based on the current voice input is likely incorrect. In this case, the control instruction can be rejected and the user can be prompted.

[0038] There are two possible reasons for inaccurate target control commands: 1. The user inputs incorrect voice data, meaning the user's voice command is incorrect. 2. The user's voice command is correct, but the voice recognition model misidentifies it. By verifying target control commands in the above manner, inaccurate target control commands caused by either of these two reasons can be detected.

[0039] In some embodiments, if the target control instruction is determined to be accurate based on the above method, the three-dimensional scanning device can be controlled based on the target control instruction.

[0040] In some embodiments, if the target control instruction is determined to be inaccurate based on the above method, for example, the target control instruction is inconsistent with the set workflow, a prompt can be given to the user to remind the user that the current target control instruction may be incorrect. There are many ways to prompt, for example, the prompt information can be output through the scanning software. In some embodiments, as shown in Figure 5, the three-dimensional scanning device may also include a voice output module 13. If the target control instruction identified based on the verification result is indicated to be inaccurate, a voice prompt information is issued through the voice output module 13. For example, the prompt "The current control instruction is to turn on the fill light, which conflicts with the previous control instruction" and so on. The voice output module 13 can be various hardware modules with audio output and playback functions, for example, it can be a speaker.

[0041] In some embodiments, the 3D scanning device can communicate with scanning software installed on a terminal device. The scanning software can be configured to receive data collected by the 3D scanning device, then use the data to perform real-time 3D reconstruction of the target object, and display the real-time reconstructed 3D model through a user interface. Furthermore, before scanning the target object using the 3D scanning device, the user can also set certain configuration parameters through the scanning software.

[0042] In some scenarios, these configuration parameters can be used to guide the 3D scanning device in scanning the target object. For example, these configuration parameters can be the operating status or operating parameters of various components in the 3D scanning device. For example, these configuration parameters can be the activation status of the fill light in the 3D scanning device, the intensity of the laser projected by the 3D scanning device, the scanning speed of the 3D scanning device, the frame rate of the image captured by the camera in the 3D scanning device, etc.

[0043] In some scenarios, these configuration parameters can also be used to guide scanning software in performing 3D reconstruction of a target object using data collected by a 3D scanning device. For example, these configuration parameters can include processing parameters for the data collected during 3D reconstruction. For example, these configuration parameters can include the brightness enhancement level of images collected by the 3D scanning device, the maximum allowable stitching error when stitching images collected by the 3D scanning device, the accuracy of the 3D model obtained by 3D reconstruction, whether to delete noise data during the 3D reconstruction process, and the type of noise data to be deleted.

[0044] Currently, each time a user uses a 3D scanning device to scan a target object, they need to manually open the scanning software on the terminal device, enter the configuration parameter setting interface, and then manually set various configuration parameters. This is quite cumbersome. In addition, for oral scanning scenarios, if the user needs to manually operate the terminal device while scanning or performing other operations (such as placing ink pad) in the oral cavity, it is easy to cause handover infection. Considering that different users often have the same operating habits or desired 3D reconstruction effects when using a 3D scanning device to scan a target object, the configuration parameters set by the same user for each scan are generally the same. In order to avoid the need to re-set the configuration parameters each time the user scans and reduce the user's manual operation, as shown in Figure 6, in some embodiments, the configuration parameters associated with different user identities can be stored in the 3D scanning device. For example, for different users, a set of configuration parameters that conform to the user's operating habits can be pre-determined, and then the set of configuration parameters can be bound to the user identity (i.e., user ID) of the user. Before a user uses a three-dimensional scanning device to scan a target object, the three-dimensional scanning device can collect the user's voice data through the voice collection module 11. The processing module 12 can obtain the voice data and extract the user's voice features from the voice data. The voice features can be features that can uniquely identify the user's identity, for example, they can be voiceprint features. The user's identity can then be determined based on the voice features.

[0045] For example, a three-dimensional scanning device may include a voice feature library that can store voice features of different users. After extracting the user's voice features from the voice data input by the current user, they can be compared one by one with each voice feature stored in the voice feature library, and the user identity of the user can be determined based on the comparison results. A set of configuration parameters bound to the user identity can then be obtained from the stored configuration parameters. Of course, if the configuration parameters are parameters used to guide the three-dimensional scanning device to scan the target object, the three-dimensional scanning device can be controlled to scan the target object based on the configuration parameters. If the configuration parameters are parameters used to guide the scanning software to perform three-dimensional reconstruction, the configuration parameters can be sent to the scanning software so that the scanning software reconstructs a three-dimensional model of the target object based on the configuration parameters.

[0046] By extracting the user's voice features from the voice data input by the user, and then automatically calling the configuration parameters that conform to the user's operating habits based on the user's voice features, the user does not need to manually set the configuration parameters in the scanning software, reducing the user's operations and making the setting of configuration parameters more convenient.

[0047] In some embodiments, the configuration parameters may include one or more of the following: parameters for indicating the operating mode of the functional modules in the three-dimensional scanning device (for example, the on state of the fill light in the three-dimensional scanning device, the intensity of the laser projected by the laser in the three-dimensional scanning device, the scanning speed of the three-dimensional scanning device, the frame rate of the image captured by the camera in the three-dimensional scanning device), parameters for indicating how the scanning software processes the data collected by the three-dimensional scanning device (for example, the brightness enhancement amplitude of the image collected by the three-dimensional scanning device, the maximum stitching error allowed when stitching the images collected by the three-dimensional scanning device, the accuracy of the three-dimensional model obtained by three-dimensional reconstruction, whether miscellaneous data needs to be deleted during the three-dimensional reconstruction process, and the type of miscellaneous data to be deleted), and parameters for indicating how the scanning software displays the reconstructed three-dimensional model (for example, whether to display the three-dimensional model without texture information or the three-dimensional model with texture information).

[0048] In some embodiments, for different users, in order to more accurately determine configuration parameters that conform to the user's operating habits, configuration parameters that conform to the user's operating habits can be predicted based on the user's historical operating behavior data and / or the 3D reconstruction results corresponding to the historical scanning tasks performed by the user using the 3D scanning device. The user's historical operating behavior data includes the user's historical control operations on the 3D scanning device and / or scanning software. For example, it can be the user's behavior of setting configuration parameters through the scanning software in historical scanning tasks, or the user's behavior of modifying configuration parameters through the 3D scanning device or scanning software. The behavior of modifying configuration parameters can be manual modification or modification through voice commands, or can also be some configuration parameter-related behavior triggered by the user through the 3D scanning device or scanning software. The present embodiment does not limit this. In addition, configuration parameters that conform to the user's operating habits can also be determined based on the 3D reconstruction results corresponding to historical scanning tasks. For example, by analyzing the 3D reconstruction results in multiple historical scanning tasks, the user's personal preferences for the reconstructed 3D model can be predicted, such as whether the reconstructed 3D model is expected to contain miscellaneous data, the user's preferred display method for the reconstructed 3D model, whether the reconstructed 3D model is expected to carry texture information, etc., and then the configuration parameters can be continuously adjusted based on the analysis results to make the configuration parameters more in line with the user's operating habits, so that the user does not need to make excessive adjustments to the configuration parameters during use.

[0049] The prediction of configuration parameters that conform to the user's operating habits based on the user's historical operating behavior data and / or the 3D reconstruction results corresponding to historical scanning tasks can be implemented by a model set in the 3D scanning device. For example, a prediction model can be pre-trained. During the user's use of the 3D scanning device, the 3D scanning device can continuously collect the user's historical operating behavior data and 3D reconstruction results so that the prediction model can continuously learn the user's operating behavior habits based on this data and update the configuration parameters to configuration parameters that are suitable for the user. Of course, the processing module 12 in the 3D scanning device can also analyze the historical operating behavior data and 3D reconstruction results to determine the optimal configuration parameters.

[0050] In the embodiment of the present disclosure, the user's operating behavior habits can be predicted based on the user's operating behavior data in historical scanning tasks and the three-dimensional reconstruction results corresponding to the historical scanning tasks to determine configuration parameters that conform to the user's operating habits. When the user subsequently uses the three-dimensional scanning device, the predetermined configuration parameters can be directly called without the user having to manually set them. The number of times the user adjusts the called configuration parameters can be reduced, which can improve scanning efficiency, reduce the user's manual operations, and enhance user experience.

[0051] Considering that for each scanning task, even if the target objects being scanned are of the same type, there may be individual differences, or the current scanning environment may vary. Therefore, the configuration parameters pre-stored in the 3D scanning device that conform to the user's operating habits may not necessarily fully meet the requirements of the current scanning task. To ensure that the automatically set configuration parameters are more consistent with the requirements of the current scanning task, in some embodiments, after determining the user's identity based on the extracted voice features and obtaining the configuration parameters associated with the user's identity, while the user uses the 3D scanning device to scan the target object, the processing module 12 can also identify target keywords from the voice data collected by the voice acquisition module 11. These target keywords are used to describe the current state of the target object. The configuration parameters can then be adjusted based on the identified target keywords, so that the 3D scanning device can be controlled to scan the target object based on the adjusted configuration parameters, and / or the adjusted configuration parameters can be sent to the scanning software so that the scanning software can reconstruct a 3D model of the target object based on the adjusted configuration parameters. For example, the target keywords can be pre-set keywords used to describe the state of the target object. For different target keywords, the corresponding configuration parameters and the adjustment method of the configuration parameters can be pre-set.

[0052] For example, taking the target object as the patient's mouth and the configuration parameter as scanning accuracy, the keyword "tooth decay" can be set in advance. Usually, when the doctor uses the oral 3D scanning equipment to scan the patient's mouth, he will tell the patient about the disease of his teeth. Therefore, if the keyword "tooth decay" is recognized from the collected doctor's voice data, in order to perform a more detailed scan of the diseased area, the scanning accuracy in the configuration parameters can be set to a larger value to improve the scanning accuracy of the 3D scanning equipment. Of course, the specific increase can be set in advance, or the severity of "tooth decay" can be identified based on the doctor's voice data, and the increase can be automatically adjusted based on the severity.

[0053] For another example, taking the target object as an industrial pipeline and the configuration parameter as scanning accuracy, the keyword "corrosion" can be set in advance. If the keyword "corrosion" is recognized from the collected user's voice data, in order to perform a more detailed scan of the pipeline corrosion area, the scanning accuracy in the configuration parameter can be set to a larger value to improve the scanning accuracy of the three-dimensional scanning equipment.

[0054] In some embodiments, taking the target object as the patient's mouth and the configuration parameter as the scanning mode as an example, the keyword "implant rod" or "metal teeth" can be pre-set. Usually, when the doctor uses the oral 3D scanning equipment to scan the patient's mouth, he will tell the patient about the disease of his teeth. Therefore, if the keyword "implant rod" or "metal teeth" is recognized from the collected doctor's voice data, in order to be able to perform a more suitable scanning mode for the area of ​​interest, the scanning mode in the configuration parameter can be switched, such as switching to: "implant rod scanning mode" or "metal scanning mode". Of course, the specific scanning mode can be set in advance, and the switching of the scanning mode can also be automatically turned on based on the tooth position number recognition of the scanning area.

[0055] In some embodiments, the 3D scanning device is an oral 3D scanning device, the target object is a patient's oral cavity, and the target keywords include keywords used to describe the oral condition, such as "cavities," "dentate jaw," "edentulous jaw," etc. Different keywords can be associated with one or more configuration parameters. Upon recognition of these keywords, the configuration parameters associated with these keywords can be automatically adjusted (e.g., enabling infrared light mode, modifying the criteria for determining when to delete miscellaneous data). The adjustment method can be pre-set or automatically determined based on the collected voice data.

[0056] In other embodiments, the three-dimensional scanning device is an oral three-dimensional scanning device, the target object is the patient's oral cavity, and the target keyword includes a keyword used to describe the current status of the identified target object, such as the target object's age status (patient's age) or racial status (patient's race). For example, the keyword can be "child", "infant", "X years old", "white", "black", etc.

[0057] In some embodiments, the user is a first user. During the process of the first user using the three-dimensional scanning device to scan the target object, the voice features of a second user other than the first user can also be obtained, wherein the voice features of the second user can be used to determine the current state of the scanned target object. For example, the first user can be a doctor and the second user can be a patient. The oral three-dimensional scanning device uses the timbre, voice line, voiceprint or accent of the second person other than the doctor as the voice features of the second user, and then identifies the current state of the scanned target object based on the voice features of the second user, such as the age of the target object (patient age), the region where the target object usually resides (patient region), or the race of the target object (patient race).

[0058] Among them, different target keywords or voice features of the second user can be associated with one or more configuration parameters. After identifying these target keywords or the voice features of the second user, the configuration parameters associated with these keywords (such as adult tooth templates, child tooth templates, black tooth templates, Caucasian tooth templates, yellow tooth templates, adult disease templates, child disease templates, or different disease templates formed according to the dietary habits of different regions) can be automatically adjusted. The adjustment method can be pre-set or automatically determined based on the collected voice data. Based on the adjusted configuration parameters, the three-dimensional scanning device is controlled to scan the target object, and / or the adjusted configuration parameters are sent to the scanning software, so that the scanning software reconstructs a three-dimensional model of the target object based on the adjusted configuration parameters or performs an analysis operation on the three-dimensional model, the analysis operation at least including tooth position number recognition, tooth feature point recognition or tooth measurement based on the tooth template, or disease detection based on the disease template.

[0059] In some embodiments, a user can use voice control to control a 3D scanning device to scan a specified area of ​​a target object. For example, the user can input voice data, and the target control instruction corresponding to the voice data is used to instruct the 3D scanning device to scan the specified area. After receiving the target control instruction, the processing module 12 of the 3D scanning device can determine whether the current scanning area of ​​the 3D scanning device is within the specified area based on the image data collected by the 3D scanning device. If it is outside the specified area, a prompt message can be issued to inform the user that the current scanning area exceeds the specified area, so that the user can readjust the position or orientation of the 3D scanning device. In some embodiments, the 3D scanning device can include a voice output module 13, which can output a voice prompt message to the user when providing prompts to the user.

[0060] Considering that when a user scans a target object using a 3D scanning device, if the scanning speed is too fast, the accuracy of the 3D reconstruction result may be affected, while if the scanning speed is too slow, the scanning rate may be affected. Therefore, the user needs to control the scanning speed to achieve accurate 3D reconstruction results while improving scanning efficiency. Therefore, in some embodiments, the 3D scanning device also includes a voice output module 13, and the processing module 12 is further configured to determine whether the current scanning speed of the 3D scanning device meets a preset condition; if not, output a voice prompt message via the voice output module. The current scanning speed of the 3D scanning device meeting the preset condition may mean that the current scanning speed of the 3D scanning device is within a preset speed range. The preset speed range can be set based on experience. When the scanning speed is within this preset speed range, the 3D reconstruction result can be highly accurate and the scanning efficiency can be maintained. In some scenarios, the 3D scanning device may include a built-in inertial measurement unit, and the speed of the 3D scanning device can be determined based on data measured by the inertial measurement unit. For example, the inertial measurement unit can be used to measure the linear acceleration and angular acceleration of the 3D scanning device in a certain direction, and the linear acceleration and angular acceleration can be integrated to determine the current scanning speed. In some scenarios, the current scanning speed of the 3D scanning device can also be determined based on the size of the overlapping area between two adjacent frames of images captured by the 3D scanning device. For example, the smaller the overlapping area, the faster the user moves the 3D scanning device, and vice versa.

[0061] Of course, in some embodiments, whether the current scanning speed of the 3D scanning device meets the preset conditions may also refer to whether two adjacent frames of images captured by the 3D scanning device can be successfully stitched together during the 3D reconstruction process. If successful, the scanning speed meets the preset conditions; otherwise, it does not. For example, if the scanning speed is too fast, there may be little overlap between two adjacent frames of images from the 3D scanning device, and thus a small number of feature points, making accurate stitching impossible. This situation may affect the accuracy of the 3D reconstruction. In this case, a prompt message may be issued to the user, prompting the user to reduce the scanning speed.

[0062] In some embodiments, the three-dimensional scanning device can be an oral three-dimensional scanning device, and the doctor can use the oral three-dimensional scanning device to scan the patient's oral cavity. Considering that when the doctor scans the patient's oral cavity, he usually tells the patient about the patient's oral disease condition and the general treatment plan. Therefore, the voice acquisition module 11 of the three-dimensional scanning device can be used to collect the conversation data between the doctor and the patient, and then the conversation data can be recognized and analyzed, and the patient's oral disease condition (for example, the type of disease, the severity of the disease) and the treatment plan can be extracted from the conversation data. The extracted disease condition and treatment plan can then be sent to the scanning software, and the scanning software automatically generates a diagnostic report based on the disease condition, treatment plan, and a preset diagnostic report template. Subsequently, the doctor can directly make fine adjustments on the automatically generated diagnostic report to obtain the final diagnostic report, which is then output to the patient. In this way, it is possible to assist the doctor in issuing a diagnostic report without the doctor having to manually write or enter the entire diagnostic report, reducing the doctor's workload. Before using the voice acquisition module 11 of the three-dimensional scanning device to collect the conversation data between the doctor and the patient, a prompt message can be issued to obtain the user's explicit consent. When identifying and analyzing the conversation data, sensitive information can be desensitized or encrypted, and sensitive information such as biometrics: name, address, country, etc. can be processed using an anonymization algorithm before reuse.

[0063] In addition, an embodiment of the present disclosure further provides a method for controlling a three-dimensional scanning device, which can be applied to the three-dimensional scanning device. The method includes:

[0064] Get the voice data input by the user;

[0065] Recognizing the voice data using a preset voice recognition model to obtain a target control instruction corresponding to the voice data, and controlling the three-dimensional scanning device to perform a corresponding operation based on the target control instruction;

[0066] During the process of the three-dimensional scanning device performing a corresponding operation, feedback information from the user regarding the operation is obtained, and based on the feedback information, whether the target control instruction matches the voice data is determined, and a label corresponding to the voice data is determined based on the matching result, and a training sample is constructed using the voice data and the label corresponding to the voice data, so as to train the voice recognition model using the constructed training sample, wherein the label is used to indicate whether the control instruction corresponding to the voice data is the target control instruction.

[0067] It is not difficult to understand that the solutions described in the above embodiments can be freely combined to obtain new solutions when there is no conflict. Due to space reasons, they are not listed one by one in the embodiments of this disclosure.

[0068] Accordingly, an embodiment of the present disclosure further provides a computer program product, which includes computer instructions. When the processor executes the computer instructions, the method mentioned in the above embodiment is implemented.

[0069] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in any of the aforementioned embodiments when the program is executed by a processor.

[0070] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0071] Through the description of the above implementation methods, it can be seen that those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the embodiments of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments of the present disclosure.

[0072] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0073] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the functions of each module can be implemented in the same one or more software and / or hardware when implementing the embodiment of the present disclosure. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the embodiment. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0074] The above is only a specific implementation of the embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the embodiment of the present disclosure. These improvements and modifications should also be regarded as the scope of protection of the embodiment of the present disclosure. Industrial Applicability

[0075] The three-dimensional scanning device provided by the present disclosure integrates a voice control function. When a user uses voice to control the three-dimensional scanning device, the three-dimensional scanning device can recognize the voice data input by the user, convert it into corresponding control instructions, and control the three-dimensional scanning device to perform the corresponding operation. In addition, during the process of the three-dimensional scanning device performing the corresponding operation, user feedback information can be collected. Based on this feedback information, it is determined whether the voice data input by the user matches the target control instruction recognized by the voice recognition model. Then, based on the matching result, the label corresponding to the voice data can be determined, thereby automatically constructing training samples with labels. The training samples automatically constructed during the voice control process can then be used to continuously train the voice recognition model, allowing the voice recognition model to continuously learn, thereby improving the accuracy of the voice recognition model. In this way, the accuracy of the voice control function of the three-dimensional scanning device can be improved, and the user experience when using the voice control function of the three-dimensional scanning device can be enhanced, which has strong industrial practicality.

Claims

1. A three-dimensional scanning device, wherein: The three-dimensional scanning device is provided with a voice acquisition module and a processing module; The voice collection module is configured to collect voice data of the user; The processing module is configured to recognize the voice data using a preset voice recognition model, parse the voice data to obtain a target control instruction, and control the three-dimensional scanning device to perform a corresponding operation based on the target control instruction, and During the process of the three-dimensional scanning device performing a corresponding operation, user feedback information regarding the operation is obtained, and based on the feedback information, whether the target control instruction matches the voice data is determined, and a label corresponding to the voice data is determined based on the matching result, and a training sample is constructed using the voice data and the label corresponding to the voice data, so as to train the voice recognition model using the constructed training sample, wherein the label is used to indicate whether the control instruction corresponding to the voice data is the target control instruction.

2. The three-dimensional scanning device according to claim 1, wherein: Before controlling the three-dimensional scanning device to perform a corresponding operation based on the target control instruction, the processing module is further configured to: Obtaining historical control instructions corresponding to each of the previous segment or segments of speech data of the speech data recognized by the speech recognition model; Verifying the accuracy of the target control instruction obtained by parsing based on the historical control instruction and the workflow of the three-dimensional scanning device; When the processing module is configured to control the three-dimensional scanning device to perform a corresponding operation based on the target control instruction, it is specifically configured to: When the verification result indicates that the target control instruction obtained by parsing is accurate, the three-dimensional scanning device is controlled to perform a corresponding operation based on the target control instruction.

3. The three-dimensional scanning device according to claim 2, wherein: The three-dimensional scanning device further includes a voice output module, and the processing module is further configured to: When the verification result indicates that the target control instruction obtained by parsing is inaccurate, a voice prompt message is issued through the voice output module.

4. The three-dimensional scanning device according to claim 1, wherein: The three-dimensional scanning device is in communication with scanning software installed on a terminal device. The three-dimensional scanning device stores configuration parameters associated with different user identities, and the configuration parameters are used to guide the three-dimensional scanning device to scan a target object and / or to guide the scanning software to perform three-dimensional reconstruction of the target object using data collected by the three-dimensional scanning device. The processing module is further configured to: Before the user scans the target object using a three-dimensional scanning device, obtaining voice data of the user; Extract the user's voice features from the voice data, determine the user's identity based on the voice features, obtain configuration parameters associated with the user identity, control the three-dimensional scanning device to scan the target object based on the configuration parameters, and / or send the configuration parameters to the scanning software so that the scanning software reconstructs a three-dimensional model of the target object based on the configuration parameters.

5. The three-dimensional scanning device according to claim 4, wherein: The configuration parameters include one or more of the following: parameters for indicating the operating mode of the functional modules in the three-dimensional scanning device, parameters for indicating the processing method of the scanning software on the data collected by the three-dimensional scanning device, and parameters for indicating the display method of the scanning software on the reconstructed three-dimensional model.

6. The three-dimensional scanning device according to claim 4, wherein: The user is a first user. After obtaining the configuration parameters associated with the user identity, the processing module is further configured to: During the process of the first user scanning the target object using the three-dimensional scanning device, identifying a target keyword or a voice feature of the second user from the voice data collected by the voice collection module, wherein the target keyword or the voice feature of the second user is used to describe the current state of the target object; The configuration parameters are adjusted based on the target keyword or the voice characteristics of the second user, the three-dimensional scanning device is controlled to scan the target object based on the adjusted configuration parameters, and / or the adjusted configuration parameters are sent to the scanning software so that the scanning software reconstructs or analyzes the three-dimensional model of the target object based on the adjusted configuration parameters.

7. The three-dimensional scanning device according to claim 6, wherein: The three-dimensional scanning device is an oral three-dimensional scanning device, the target object is a patient's oral cavity, and the target keywords include keywords used to describe the patient's oral cavity condition; and / or The voice feature of the second user is the voice feature of the patient.

8. The three-dimensional scanning device according to claim 4, wherein: The configuration parameters corresponding to different user identities are obtained based on the user's historical operation behavior data and / or the three-dimensional reconstruction results corresponding to the historical scanning tasks performed by the user using the three-dimensional scanning device, wherein the historical operation behavior data includes the user's control operations on the three-dimensional scanning device and / or the scanning software.

9. The three-dimensional scanning device according to claim 1, wherein: The three-dimensional scanning device further includes a voice output module, the target control instruction is used to instruct the three-dimensional scanning device to scan a specified area, and the processing module is further configured to: Determining whether a current scanning area of ​​the three-dimensional scanning device is within the designated area based on image data collected by the three-dimensional scanning device; If not, the voice prompt information is output through the voice output module.

10. The three-dimensional scanning device according to claim 1, wherein: The three-dimensional scanning device further includes a voice output module, and the processing module is further configured to: Determining whether a current scanning speed of the three-dimensional scanning device meets a preset condition; If not, outputting a voice prompt message through the voice output module; Wherein, determining that the current scanning speed of the three-dimensional scanning device meets the preset condition includes any of the following situations: The current scanning speed of the three-dimensional scanning device is within a preset speed range, wherein the current speed of the three-dimensional scanning device is determined by an inertial measurement unit provided in the three-dimensional scanning device, or the current scanning speed of the three-dimensional scanning device is determined by the size of an overlapping area between two adjacent frames of images acquired by the three-dimensional scanning device; and / or During the process of performing three-dimensional reconstruction using the images acquired by the three-dimensional scanning device, two adjacent frames of images acquired by the three-dimensional scanning device are successfully spliced ​​together.

11. A method for controlling a three-dimensional scanning device, wherein: The method comprises: Get the voice data input by the user; Recognizing the voice data using a preset voice recognition model to obtain a target control instruction corresponding to the voice data, and controlling the three-dimensional scanning device to perform a corresponding operation based on the target control instruction; During the process of the three-dimensional scanning device performing a corresponding operation, feedback information from the user regarding the operation is obtained, and based on the feedback information, whether the target control instruction matches the voice data is determined, and a label corresponding to the voice data is determined based on the matching result, and a training sample is constructed using the voice data and the label corresponding to the voice data, so as to train the voice recognition model using the constructed training sample, wherein the label is used to indicate whether the control instruction corresponding to the voice data is the target control instruction.

12. A computer program product, wherein: The computer program product comprises computer instructions, which implement the method according to claim 11 when executed by a processor.

13. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, which implements the method according to claim 11 when executed by a processor.

Citation Information

Patent Citations

  • Medical information feedback method and device, equipment and readable storage medium

    CN111813946A

  • Language model training method and device, language model application method and device, equipment and storage medium

    CN112466295A

  • Voice model training method, apparatus and device, and computer readable storage medium

    CN114399995A

  • Scanning system, method and equipment based on voice control and storage medium

    CN114898845A

  • Three-dimensional scanning apparatus, control method, computer program product, and storage medium

    CN118248126A