X-ray imaging system

The X-ray imaging system addresses pronunciation and accent variations by playing back sample audio of keywords, enhancing voice recognition accuracy and control execution.

WO2026018862A1PCT designated stage Publication Date: 2026-01-22SHIMADZU CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/025452
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-18
Filing Date
2025-07-16
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing X-ray imaging systems face challenges in accurately recognizing user keywords due to variations in pronunciation and accent, which can lead to misinterpretation of voice commands.

Method used

The system includes an audio output unit that plays back and outputs sample audio of keywords for voice recognition, allowing operators to easily understand the pronunciation and accent of keywords through voice recognition.

Benefits of technology

Enables accurate and consistent voice recognition by providing sample audio of keywords, ensuring operators can perform predetermined controls effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025025452_22012026_PF_FP_ABST
    Figure JP2025025452_22012026_PF_FP_ABST
Patent Text Reader

Abstract

This X-ray imaging system (100) comprises: a voice input unit (5) that receives voice input; a voice output unit (15); and a control unit (8) that, on the basis of the voice received by the voice input unit (5), performs voice recognition of a keyword (86), thereby performing control based on the keyword (86). The control unit (8) is configured to reproduce and output a sample voice of the keyword (86), by using the voice output unit (15), on the basis of an operation performed by an operator.
Need to check novelty before this filing date? Find Prior Art

Description

X-ray imaging system

[0001] The present invention relates to an X-ray imaging system.

[0002] 2. Description of the Related Art Conventionally, an X-ray imaging apparatus has been known, and such an X-ray imaging apparatus is disclosed, for example, in U.S. Pat. No. 9,271,687.

[0003] The above-mentioned U.S. Patent No. 9,271,687 discloses an X-ray imaging device (X-ray imaging system) including an X-ray irradiation unit, an X-ray detection unit that detects X-rays irradiated from the X-ray irradiation unit, a display unit (alert unit) that displays an X-ray image, and an audio input unit provided on the display unit. The audio input unit includes a microphone that accepts audio input from a user (operator). The user's audio input to the audio input unit may include keywords that cause an area selection unit (control unit) to execute a specific function. An example of the specific function is an area selection function that selects a specific area within an X-ray image. That is, the area selection unit is configured to select a specific area within an X-ray image based on the user's audio accepted by the audio input unit.

[0004] U.S. Patent No. 9,271,687

[0005] However, the accent and intonation of a keyword uttered by a user may vary from user to user depending on the user's place of birth, place of residence, native language, etc. Furthermore, for example, in words composed of alphabetic characters, the pronunciation of the alphabet may vary depending on the word. Furthermore, if a keyword is a coined word, the user may not know the pronunciation of the coined keyword because the coined keyword is not a known keyword. Therefore, although not explicitly stated in the above-mentioned U.S. Patent No. 9,271,687 specification, even if a user (operator) utters a keyword to cause the area selection unit (control unit) to perform keyword-based control, the area selection unit may not recognize the keyword uttered by the user because the pronunciation or accent of the keyword uttered by the user differs from the pronunciation or accent that the area selection unit can recognize. Therefore, it is desirable for an operator to be able to easily understand the pronunciation and accent of a keyword used to perform a predetermined control through speech recognition.

[0006] The present invention has been made to solve the above-mentioned problems, and one object of the present invention is to provide an X-ray imaging system that allows an operator to easily understand the pronunciation and accent of keywords used to perform specified controls through voice recognition.

[0007] An X-ray imaging system according to one aspect of the present invention includes an X-ray irradiation unit, an X-ray detection unit that detects X-rays irradiated from the X-ray irradiation unit, an audio input unit that accepts audio input, an audio output unit, and a control unit that performs keyword-based control by voice recognition of keywords based on the audio accepted by the audio input unit, and the control unit is configured to play back and output sample audio of the keywords via the audio output unit based on the operation of an operator.

[0008] In the X-ray imaging system according to the above aspect, as described above, the control unit is configured to, based on an operation by the operator, play and output, by the audio output unit, a sample audio of a keyword for performing keyword-based control through voice recognition. Since the sample audio of the keyword is played and output by the audio output unit, the operator can listen to and hear the sample audio of the keyword for performing predetermined control in the X-ray imaging system through voice recognition. Therefore, the operator can easily understand the pronunciation and accent of the keyword for performing predetermined control through voice recognition.

[0009] FIG. 1 is a schematic diagram showing the overall configuration of an X-ray imaging system according to an embodiment; FIG. 2 is a functional block diagram of an X-ray imaging system according to an embodiment; FIG. 3 is a front view showing a display device; FIG. 4 is a diagram showing an example of a setting screen; FIG. 5 is a diagram showing an example of a test start keyword recognition image; FIG. 6 is a diagram showing an example of a test instruction keyword recognition image; FIG. 7 is a diagram showing an example of a keyword unrecognized image; FIG. 8 is a diagram for explaining an example of instruction keywords and control based on the instruction keywords; FIG. 9 is a flowchart for explaining sample audio playback output processing by a control unit; and FIG. 10 is a sequence diagram for explaining voice recognition test processing in a voice recognition test mode by a control unit.

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.

[0011] 1 performs fluoroscopy, which captures (fluoroscopically views) images of the interior of a subject's body by irradiating the subject with X-rays while the subject has a medical device inserted therein. For example, the X-ray imaging system 100 captures images (moving images) for confirming the state of the subject's body when performing percutaneous coronary intervention (PCI). The device includes, for example, a stent placed in a blood vessel in the subject's heart.

[0012] In percutaneous coronary intervention, an operator such as a doctor or technician inserts a device such as a stent into the body of a subject while checking the internal condition of the subject by visually viewing fluoroscopic images as real-time moving images displayed on the display device 10. The X-ray imaging system 100 is configured to display on the display device 10 a superimposed image in which a fluoroscopic image and a blood vessel image captured at an imaging angle corresponding to the imaging angle of the fluoroscopic image are superimposed on each other, in order to place a device such as a stent at a stenotic portion of a coronary artery using a catheter or the like while checking the shape of the blood vessel in the fluoroscopic image.

[0013] Here, the vascular image is a DSA (Digital Subtraction Angiography) image generated by digitally subtracting a non-contrast image from a contrast image. The non-contrast image is an image captured at multiple imaging angles when a contrast agent has not been administered to the subject (when no contrast agent is present in the blood vessels). The contrast image is an image captured at multiple imaging angles similar to those used to capture the non-contrast image when a contrast agent has been administered to the subject (when contrast agent remains in the blood vessels).

[0014] The X-ray imaging system 100 also captures images of the subject's lower limbs to identify the treatment location of the blood vessels in the subject's lower limbs. Because the blood vessels in the lower limbs extend from the base of the subject's feet to the toes, it is not possible to capture all of them in a single imaging session. Therefore, the X-ray imaging system 100 generates a long image by stitching together multiple images. The long image is generated by stitching together multiple vascular images obtained by subtracting multiple contrast agent images and multiple non-contrast agent images that have the same relative position coordinates. The contrast agent image is an image captured multiple times while the X-ray irradiator 40 and X-ray detector 41 are moved in the longitudinal direction of the tabletop 3 and the relative positions of the X-ray irradiator 40 and X-ray detector 41 and the subject are changed after the subject has been administered a contrast agent. In addition, the non-contrast agent image is an image that is taken multiple times under the same imaging conditions as when the contrast agent image is taken, with no contrast agent administered to the subject, by moving the X-ray irradiation unit 40 and the X-ray detection unit 41 in the longitudinal direction of the tabletop 3 and changing the relative positions of the X-ray irradiation unit 40 and the X-ray detection unit 41 and the subject.

[0015] (Configuration of X-Ray Imaging System) The configuration of an X-ray imaging system 100 according to an embodiment of the present invention will be described with reference to FIGS.

[0016] 1, the X-ray imaging system 100 includes an X-ray imaging device 200, a notification unit 1 including a display device 10, and a display device moving mechanism 2. The X-ray imaging device 200 includes a tabletop 3, an imaging unit 4, a voice input unit 5, an operation unit 6, an image processing unit 7 (see FIG. 2), a control unit 8 (see FIG. 2), and a storage unit 9 (see FIG. 2).

[0017] The tabletop 3 is configured as a bed on which the subject lies.

[0018] The imaging unit 4 is configured to perform X-ray imaging of the subject. The imaging unit 4 is also configured to perform fluoroscopic imaging of the subject while changing the imaging angle. That is, the X-ray imaging system 100 can capture X-ray images as still images (X-ray imaging) and can also capture fluoroscopic images as moving images (fluoroscopic imaging). The imaging unit 4 is configured to sequentially perform fluoroscopic imaging to capture moving images of internal parts of the subject's body (e.g., the heart or lower limbs) at each of a plurality of imaging angles. The imaging unit 4 includes an X-ray irradiation unit 40, an X-ray detection unit 41, and a holding unit 42. In this specification, the term "image" may include both X-ray images as still images and fluoroscopic images as moving images.

[0019] The X-ray irradiator 40 is configured to irradiate the subject with X-rays. The X-ray irradiator 40 includes an X-ray tube that irradiates X-rays when power is supplied. The X-ray tube is configured such that the irradiated X-rays are controlled by controlling the voltage and current applied thereto.

[0020] The X-ray detection unit 41 is configured to detect X-rays that are irradiated from the X-ray irradiation unit 40 and have passed through the subject. The X-ray detection unit 41 outputs a detection signal based on the detected X-rays. The X-ray detection unit 41 includes, for example, an FPD (Flat Panel Detector).

[0021] The holder 42 is configured to hold the X-ray irradiator 40 and the X-ray detector 41 so as to be able to change the imaging angle of fluoroscopic imaging by the imaging unit 4. Specifically, the holder 42 holds the X-ray irradiator 40 and the X-ray detector 41 so that they face each other across the tabletop 3 on which the subject lies.

[0022] The notification unit 1 includes a display device 10. As shown in FIG. 3, the display device 10 includes a display unit 11, a handle 12, and an audio output unit 15.

[0023] The display unit 11 is, for example, a monitor such as a liquid crystal display. The display unit 11 includes a first display area 13 and a second display area 14 different from the first display area 13. The first display area 13 is configured to display information on X-ray imaging conditions, including information on the tube voltage and tube current applied to the X-ray tube, as well as various images such as fluoroscopic images, superimposed images, and long images. The second display area 14 is configured to switch between a screen displaying the various images and a setting screen 80 (see FIG. 4 ) for setting the voice recognition enabled mode, etc. The display device 10 is suspended from a ceiling 90 in an imaging room 91 in which the imaging unit 4 is installed via a display device moving mechanism 2.

[0024] The handle 12 is configured to move the display unit 11 via the display device moving mechanism 2 (see FIG. 1 ). For example, the handle 12 is a rod-shaped member attached along the left side, bottom, and right side of the display unit 11. Here, the left side of the display unit 11 refers to the left side of the display unit 11 when an operator views the display unit 11 from the front, and the right side of the display unit 11 refers to the right side of the display unit 11 when an operator views the display unit 11 from the front. The operator can move the display unit 11 via the display device moving mechanism 2 by gripping and moving the handle 12. Note that in this specification, the term "operator" includes doctors, technicians, and the like, as well as service personnel who install and maintain the X-ray imaging system 100.

[0025] The audio output unit 15 is configured to play back and output a sample audio of the keyword 86 (see FIG. 4 ). The audio output unit 15 is configured to play back and output the sample audio of the keyword 86 based on an operation by the operator. The audio output unit 15 includes, for example, a speaker provided in the display unit 11.

[0026] As shown in FIG. 1 , the display device moving mechanism 2 is configured to movably support the display device 10. The display device moving mechanism 2 includes a rail 20, a ceiling suspension device 21, and a support member 22. The rail 20 is attached to the ceiling 90 inside the imaging room 91. The ceiling suspension device 21 is configured to be movable in the horizontal direction (directions along the longitudinal direction of the top plate 3 and along the lateral direction of the top plate 3) using the rail 20. The ceiling suspension device 21 is configured to support a support member 22. The support member 22 is configured to support the display device 10. The support member 22 includes a first support member 23 and a second support member 24. The first support member 23 is supported by the ceiling suspension device 21 and is configured to rotate the second support member 24 and the display device 10 around a vertical axis 92. The second support member 24 is also configured to be rotatable in the vertical direction relative to the first support member 23, thereby allowing the display device 10 to be moved in the vertical direction. In addition, the display device moving mechanism 2 is not limited to the above configuration as long as it is configured to be able to move the display device 10 horizontally (in the direction along the longitudinal direction of the top plate 3 and in the direction along the short side direction of the top plate 3) and vertically.

[0027] The operator can move the display device 10 horizontally and vertically via the handle 12. The operator can also rotate the display device 10 around a vertical axis 92 via the handle 12. The display device moving mechanism 2 has a motor 25 and the like, so the operator can easily move and rotate the display device 10 via the display device moving mechanism 2. The display device 10 may be configured to be movable to a predetermined position in accordance with the position of the imaging unit 4 selected by the operator through an input operation of the operation unit 6. The display device 10 and the voice input unit 5 provided on the display device 10 are configured to be movable relative to the operator who is standing on the longitudinal side of the tabletop 3 and performing the procedure.

[0028] 3 , the voice input unit 5 is configured to receive voice input from the operator. The voice input unit 5 includes, for example, a microphone. The voice input unit 5 is provided in the display device 10. That is, the voice input unit 5 and the voice output unit 15 are provided in the display device 10.

[0029] The audio input unit 5 is provided on the handle 12 of the display device 10. The audio input unit 5 is provided with an attachment member 5a that is detachably attached to the handle 12. The attachment member 5a is, for example, a clip. The audio input unit 5 is attached via the clip to the handle 12 that faces the left side of the display unit 11. The audio input unit 5, together with the display device 10, is configured to be movable in the horizontal and vertical directions by the display device movement mechanism 2 and to be rotatable around a vertical axis 92.

[0030] 1, the operation unit 6 is configured to accept various input operations related to X-ray imaging and fluoroscopic imaging by the operator, as well as various input operations on a setting screen 80 (see FIG. 4) by the operator. The operation unit 6 is provided on the tabletop 3. The operation unit 6 includes, for example, a touch panel and a lever.

[0031] The operation unit 6 accepts, for example, an input operation in the first display area 13 related to a fluoroscopic image to be stored in the storage unit 9, an input operation related to the display of a fluoroscopic image on the display device 10, and an input operation related to a selection for generating a long image by stitching together a plurality of vascular images captured after administering a contrast agent. The operation unit 6 also accepts, for example, an input operation in the setting screen 80 (see FIG. 4 ) displayed in the second display area 14 related to a selection between a voice-recognition enabled mode and a voice-recognition disabled mode for disabling the voice-recognition enabled mode, an input operation related to settings for voice recognition in the voice-recognition enabled mode, an input operation for playing and outputting a sample voice of a keyword 86 (see FIG. 4 ) by the voice output unit 15, and an input operation for starting and ending a voice recognition test mode for testing whether the voice of the keyword 86 has been recognized.

[0032] 2, the image processing unit 7 is configured to generate an image based on a detection signal output from the X-ray detection unit 41. The image processing unit 7 is configured, for example, by a processor such as a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) configured for image processing. The image processing unit 7 is also configured to generate a fluoroscopic image and a vascular image of the subject including an image of a contrast agent based on the detection signal output from the X-ray detection unit 41. The image processing unit 7 is also configured to generate a long image by stitching together multiple vascular images captured after administering a contrast agent.

[0033] The control unit 8 is configured to perform control based on the keyword 86 by voice-recognizing the keyword 86 (see FIG. 4) based on the voice received by the voice input unit 5. The keyword 86 includes a command keyword 86b (see FIG. 4) for executing a corresponding predetermined function, and a start keyword 86a (see FIG. 4) that serves as a trigger for voice recognition of the command keyword 86b. The control unit 8 is configured to perform control to start voice recognition for the command keyword 86b by voice-recognizing the start keyword 86a based on the voice received by the voice input unit 5, and control to execute the function corresponding to the command keyword 86b by voice-recognizing the command keyword 86b based on the voice received by the voice input unit 5.

[0034] The control unit 8 is also configured to reproduce and output sample voice of a keyword 86 (see FIG. 4) from the voice output unit 15 based on the operation of the operator.

[0035] Furthermore, in the voice recognition test mode, the control unit 8 is configured to determine whether or not the test voice of the keyword 86 (see FIG. 4 ) that is reproduced and output from the voice output unit 15 based on an operation by the operator and that is accepted by the voice input unit 5 has been recognized. Specifically, in the voice recognition test mode, the control unit 8 is configured to determine whether or not the sample voice of the keyword 86 included in the test voice that is reproduced and output from the voice output unit 15 based on an operation by the operator and that is accepted by the voice input unit 5 has been recognized.

[0036] Furthermore, in the voice recognition test mode, the control unit 8 is configured to cause the notification unit 1 to notify the result of voice recognition of the sample voice that is reproduced and output from the voice output unit 15 and accepted by the voice input unit 5. Specifically, in the voice recognition test mode, the control unit 8 is configured to cause the display device 10 to display the result of voice recognition of the sample voice that is reproduced and output from the voice output unit 15 and accepted by the voice input unit 5.

[0037] The control unit 8 includes a CPU, a ROM (Read Only Memory), a RAM (Random Access Memory), a GPU, or an FPGA configured for image processing.

[0038] The storage unit 9 stores various images generated by the image processing unit 7 and various programs executed by the control unit 8. The storage unit 9 also stores sample audio of keywords 86 (see FIG. 4 ). The storage unit 9 also stores keyword information corresponding to the keywords 86 that are speech-recognized based on the speech received by the speech input unit 5. The storage unit 9 is, for example, a non-volatile storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive).

[0039] (Setting Screen) The setting screen 80 will be described. FIG. 4 is a diagram showing an example of the setting screen 80. The setting screen 80 is a screen displayed in the second display area 14 of the display unit 11 when the operator performs settings related to the voice recognition enabled mode, etc. The operator can perform settings related to the voice recognition enabled mode, etc. by performing input operations via the operation unit 6 on various GUIs (Graphical User Interfaces) displayed on the setting screen 80. Note that the input operations on the various GUIs displayed on the setting screen 80 may be performed by the operation unit 6 or may be performed by a separately provided input device including a keyboard, a mouse, etc.

[0040] The setting screen 80 includes a voice recognition enable mode setting checkbox 81, a keyword language selection field 82, an amplitude display area 83, a start button 84 and an end button (not shown) for the voice recognition test mode, and a keyword list 85.

[0041] The voice recognition enable mode setting check box 81 is a GUI for setting the voice recognition enable mode. The operator operates the operation unit 6 to check the voice recognition enable mode setting check box 81, thereby setting the voice recognition enable mode.

[0042] The keyword language selection field 82 is a GUI for selecting the language of the voice of the keyword 86 to be voice-recognized in the voice-recognition enabled mode. As an example, either English (American) or Japanese can be selected as the language of the voice of the keyword 86 to be voice-recognized in the voice-recognition enabled mode. The operator operates the operation unit 6 to select either "English" or "Japanese" in the keyword language selection field 82, and voice recognition is performed on the voice of the keyword 86 in the selected language.

[0043] The amplitude display area 83 is configured to display the amplitude of the voice input by the voice input unit 5. For example, in the voice recognition test mode, the amplitude display area 83 displays the amplitude of the sample voice that is reproduced and output from the voice output unit 15 and received by the voice input unit 5.

[0044] The voice recognition test mode start button 84 and end button are GUIs for starting or ending the voice recognition test mode. The voice recognition test mode start button 84 is displayed on the setting screen 80 before the voice recognition test mode is started, as shown in FIG. 4 . Furthermore, an end button (not shown) for the voice recognition test mode is displayed on the setting screen 80 while the voice recognition test mode is in progress, instead of the start button 84. The operator starts the voice recognition test mode by operating the operation unit 6 to press the start button 84. Furthermore, the operator ends the voice recognition test mode by operating the operation unit 6 to press the end button.

[0045] The keyword list 85 includes a plurality of keywords 86, keyword check boxes 87 corresponding to each of the plurality of keywords 86, and play buttons 88 corresponding to each of the plurality of keywords 86 and for playing back sample audio for the plurality of keywords 86. In the keyword list 85, the keyword check boxes 87, the plurality of keywords 86, and the play buttons 88 corresponding to each of the plurality of keywords 86 are displayed side by side. Note that Fig. 4 shows an example of the keyword list 85 when "Japanese" is set in the keyword language selection field 82; when "English" is set in the keyword language selection field 82, English (American) keywords 86 corresponding to the Japanese keywords 86 are displayed in the keyword list 85.

[0046] The plurality of keywords 86 includes one start keyword 86 a and a plurality of instruction keywords 86 b. As the screen of the keyword list 85 is scrolled, each of the plurality of keywords 86 is displayed in sequence together with a keyword check box 87 and a play button 88.

[0047] The start keyword 86a is a so-called wake-up word. The start keyword 86a is a keyword 86 that is uttered by the operator before the operator utters the instruction keyword 86b. One example of the start keyword 86a is "Trinious." Note that the word for the start keyword 86a is not limited to "Trinious," and other words may be set.

[0048] The command keyword 86b is a keyword 86 for command input, and is issued by the operator after the operator issues the start keyword 86a, "Torius." The command keyword 86b includes a plurality of command keywords 86b corresponding to corresponding functions. As an example, one of the plurality of command keywords 86b is "next image," and the predetermined function corresponding to "next image" is to display, on the display device 10, a fluoroscopic image of a moving image that is stored one image after the plurality of fluoroscopic images. The command keyword 86b and the predetermined function corresponding to the command keyword 86b will be described later. The fluoroscopic images include images obtained by fluoroscopy (observed at a low dose, and images are not saved unless a fluoroscopic saving function is executed) and images obtained by radiography (observed at a higher dose than fluoroscopy, and images are saved in the storage unit 9).

[0049] The keyword check box 87 is a GUI for setting a keyword 86 for executing a corresponding function of the X-ray imaging system 100 by recognizing the voice of the keyword 86 uttered by the operator through voice recognition processing in the voice recognition enabled mode. The operator operates the operation unit 6 to check or uncheck the keyword check box 87 corresponding to the desired keyword 86, thereby setting the keyword 86 for executing the corresponding function of the X-ray imaging system 100 through voice recognition.

[0050] The play button 88 corresponds to each of the plurality of keywords 86, and is a GUI for playing and outputting sample audio of the plurality of keywords 86. When the operator operates the operation unit 6 to press the play button 88, the sample audio of the keyword 86 corresponding to the play button 88 is played and output by the audio output unit 15.

[0051] (Functional Blocks of the Control Unit) The functional blocks included in the control unit 8 will be described with reference to Fig. 2. The control unit 8, which is made up of a CPU, GPU, and the like as hardware, includes a playback output control unit 8a, a display control unit 8b, and a keyword determination processing unit 8c as software (program) functional blocks. The control unit 8 functions as the playback output control unit 8a, the display control unit 8b, and the keyword determination processing unit 8c by executing a program stored in the storage unit 9. The playback output control unit 8a, the display control unit 8b, and the keyword determination processing unit 8c may each be configured individually as hardware by providing a dedicated processor (processing circuit).

[0052] (Playback output of sample voice of keyword) The playback output of sample voice of keyword 86 will be described with reference to Fig. 4. The playback output control unit 8a (see Fig. 2) is configured to play back and output the sample voice of keyword 86 by the voice output unit 15 based on an operation by the operator.

[0053] The sample audio of the keyword 86 includes a sample audio of the start keyword 86 a and a sample audio of the instruction keyword 86 b. The sample audio of the start keyword 86 a and the sample audio of the instruction keyword 86 b are stored in advance in the storage unit 9 (see FIG. 2 ).

[0054] The sample voice of the start keyword 86a is a voice sample of the start keyword 86a spoken with standard pronunciation and a standard accent for each language that can be voice-recognized by the keyword determination processing unit 8c (see FIG. 2). In the present embodiment, as an example, either English (American) or Japanese can be selected as the language of the voice of the keyword 86 to be voice-recognized in the voice-recognition-enabled mode. Therefore, the sample voice of the start keyword 86a includes a voice sample of the English (American) start keyword 86a spoken with standard English (American) pronunciation and a standard English (American) accent, and a voice sample of the Japanese start keyword 86a spoken with standard Japanese pronunciation and a standard Japanese accent. The storage unit 9 (see FIG. 2) stores a sample voice of "Trinious" as an example of the start keyword 86a.

[0055] The sample voice of the instruction keyword 86b is a sample voice of the instruction keyword 86b uttered with standard pronunciation and standard accent for each language that can be voice-recognized by the keyword determination processing unit 8c (see FIG. 2). As an example, the sample voice of the instruction keyword 86b includes a sample voice of the English (American) instruction keyword 86b uttered with standard pronunciation and standard accent for English (American), and a sample voice of the Japanese instruction keyword 86b uttered with standard pronunciation and standard accent for Japanese. The memory unit 9 (see FIG. 2) stores all sample voices of the multiple instruction keywords 86b. The memory unit 9 stores a sample voice of "next image" as an example of the multiple instruction keywords 86b.

[0056] The sample audio of the start keyword 86a and the sample audio of the instruction keyword 86b are played back and output by the audio output unit 15 (see FIG. 2) based on an operator's operation on the operation unit 6. Specifically, the sample audio of the start keyword 86a and the sample audio of the instruction keyword 86b are played back and output from the audio output unit 15 under the control of the playback output control unit 8a (see FIG. 2) based on an operator's operation of pressing a play button 88 displayed on the setting screen 80.

[0057] (Voice Recognition Enabled Mode) The voice recognition enabled mode is described below. The voice recognition enabled mode is a mode in which a keyword 86 is recognized based on the operator's voice through a voice recognition process, and a function of the X-ray imaging system 100 corresponding to the keyword 86 is executed.

[0058] The control unit 8 shown in Fig. 2 is configured to execute the voice-recognition enabled mode when the voice-recognition enabled mode is set by checking a voice-recognition enabled mode setting checkbox 81 on a setting screen 80 (see Fig. 4). In the voice-recognition enabled mode, the control unit 8 is configured to start voice recognition for a command keyword 86b (see Fig. 4) by voice-recognizing a start keyword 86a (see Fig. 4) based on the operator's voice, and to execute a function corresponding to the command keyword 86b by voice-recognizing the command keyword 86b based on the operator's voice.

[0059] The following describes voice recognition by the control unit 8. The keyword determination processing unit 8c is configured to perform control based on the keyword 86 by voice-recognizing a keyword 86 (see FIG. 4) based on the voice received by the voice input unit 5. The keyword determination processing unit 8c is configured to start voice recognition for a command keyword 86b (see FIG. 4) by voice-recognizing a start keyword 86a (see FIG. 4) based on the voice received by the voice input unit 5. The keyword determination processing unit 8c is also configured to execute a function corresponding to the command keyword 86b by voice-recognizing the command keyword 86b based on the voice received by the voice input unit 5.

[0060] Specifically, the keyword determination processing unit 8c determines a keyword 86 (see FIG. 4 ) based on the voice data acquired from the voice input unit 5. As an example, the keyword determination processing unit 8c performs a voice recognition process on the voice data acquired from the voice input unit 5 to convert the voice data into text data. Then, the keyword determination processing unit 8c determines control corresponding to the converted text data by referring to keyword information stored in the storage unit 9. Then, the keyword determination processing unit 8c executes processing corresponding to the determined control. Note that the process of determining control based on the keyword 86 by voice recognizing the keyword 86 based on the operator's voice is not limited to the above example, and known techniques can be applied.

[0061] More specifically, for example, if the keyword 86 (see FIG. 4) is "Trinious" as the start keyword 86a (see FIG. 4), the keyword determination processing unit 8c performs a voice recognition process on the voice data including "Trinious" as the start keyword 86a, converts the voice data into text data, and, with reference to the keyword information, determines that the control is to start voice recognition for the command keyword 86b. Then, the keyword determination processing unit 8c starts voice recognition for the command keyword 86b.

[0062] In addition, when the keyword determination processing unit 8c recognizes the voice of "Trinious", which is the start keyword 86a (see Figure 4), the display control unit 8b displays an image of "(check mark):Trinious", which is an execution start keyword recognition image (not shown), in the first display area 13.

[0063] For example, if the keyword 86 is "next image" as the instruction keyword 86b (see FIG. 4), the keyword determination processing unit 8c executes a voice recognition process on the voice data including the instruction keyword 86b to convert the voice data into text data, and also refers to the keyword information to determine that the control is to display the fluoroscopic image of the moving image that is stored one image after the multiple fluoroscopic images on the display device 10. Then, the keyword determination processing unit 8c causes the fluoroscopic image of the moving image that is stored one image after the multiple fluoroscopic images on the display device 10.

[0064] In addition, when the keyword determination processing unit 8c recognizes the voice of ``next image'', which is the instruction keyword 86b (see Figure 4), the display control unit 8b displays an image of ``(check mark): next image'', which is an execution instruction keyword recognition image (not shown), in the first display area 13.

[0065] (Voice Recognition Test Mode) The voice recognition test mode is described below. The voice recognition test mode is a mode in which a voice recognition test is performed to test whether or not the sample voice of the keyword 86 that is played back and output from the voice output unit 15 and input via the voice input unit 5 has been recognized by the control unit 8.

[0066] The control unit 8 is configured to start the voice recognition test mode when an operation of pressing a start button 84 on the setting screen 80 (see FIG. 4) is performed. In the voice recognition test mode, the control unit 8 is configured to concurrently perform control to play and output a sample voice from the voice output unit 15 based on an operation by the operator, and control to determine whether or not the sample voice played and output from the voice output unit 15 and received by the voice input unit 5 has been recognized.

[0067] That is, in the voice recognition test, based on the operator's operation of pressing a play button 88 displayed on the setting screen 80 (see FIG. 4), the playback output control unit 8a causes the voice output unit 15 to play back and output a start keyword 86a (see FIG. 4) and a command keyword 86b (see FIG. 4) corresponding to the pressed play button 88. Then, the sample voices of the start keyword 86a and the command keyword 86b played back and output by the voice output unit 15 are input to the voice input unit 5. Then, the keyword determination processing unit 8c determines whether the sample voices of the start keyword 86a and the command keyword 86b input by the voice input unit 5 have been voice recognized.

[0068] During the execution of the voice recognition test mode, even if the keyword determination processing unit 8c recognizes the start keyword 86a (see Figure 4) and the instruction keyword 86b (see Figure 4) based on the sample voice input by the voice input unit 5, it does not control the execution of the function corresponding to the instruction keyword 86b.

[0069] In the voice recognition test, when the keyword determination processing unit 8c performs voice recognition on the sample voice of the start keyword 86a (see FIG. 4 ) that is played back and output from the voice output unit 15 and input via the voice input unit 5, the display control unit 8b is configured to control the display device 10 to display a result of the voice recognition of the sample voice of the start keyword 86a. That is, when the keyword determination processing unit 8c performs voice recognition on the sample voice of "Trinious," which is the start keyword 86a that is played back and output from the voice output unit 15 and input via the voice input unit 5, the display control unit 8b displays a test start keyword recognition image 70 (see FIG. 5 ) indicating the voice recognition result in the first display area 13. Specifically, an image of "TEST:Trinious" is displayed in the first display area 13 as the test start keyword recognition image 70.

[0070] Furthermore, in the voice recognition test, when the keyword determination processing unit 8c performs voice recognition on the sample voice of the instruction keyword 86b (see FIG. 4 ), which is reproduced and output from the voice output unit 15 and input via the voice input unit 5, the display control unit 8b is configured to control the display device 10 to display a result of the voice recognition of the sample voice of the instruction keyword 86b. That is, as an example, when the keyword determination processing unit 8c performs voice recognition on the sample voice of "next image," which is the instruction keyword 86b, which is reproduced and output from the voice output unit 15 and input via the voice input unit 5, the display control unit 8b displays a test instruction keyword recognition image 71 (see FIG. 6 ), which indicates the result of the voice recognition, in the first display area 13. Specifically, as an example, an image of "TEST: next image" is displayed in the first display area 13 as the test start keyword recognition image 70.

[0071] In the voice recognition test, if the keyword determination processing unit 8c fails to recognize the sample voice of the start keyword 86a (see FIG. 4) or the instruction keyword 86b (see FIG. 4) before a predetermined time has elapsed, the display control unit 8b is configured to control the display device 10 to display a result indicating that the sample voice of the start keyword 86a or the instruction keyword 86b has not been recognized. That is, as an example, if the keyword determination processing unit 8c fails to recognize the sample voice of the start keyword 86a or the instruction keyword 86b, the display control unit 8b displays a keyword unrecognized image 72 (see FIG. 7) indicating the result of the voice unrecognition in the first display area 13. Specifically, as an example, an image of "X Sorry, I couldn't hear you" is displayed in the first display area 13 as the keyword unrecognized image 72.

[0072] When a voice recognition test is performed in the voice recognition test mode, the operator places the voice input unit 5 in a predetermined position for the voice recognition test. Specifically, the operator attaches the voice input unit 5 to the handle 12 of the display device 10 so that the sound collection unit of the voice input unit 5 faces the voice output unit 15 provided on the display device 10. When a voice recognition test is performed in the voice recognition test mode, the operator sets the volume of the voice output unit 15 to a predetermined volume for the voice recognition test. By placing the voice input unit 5 in a predetermined position for the voice recognition test and setting the volume of the voice output unit 15 to a predetermined volume for the voice recognition test, the voice recognition test can be performed under substantially the same conditions with respect to the pronunciation, accent, and volume of the voice based on the playback output of the sample voice of the keyword 86 (see FIG. 4 ), even when the X-ray imaging system 100 is installed in different environments.

[0073] The timing of executing the voice recognition test mode will now be described. The operator can start the voice recognition test before, during, or after an examination of a subject by the X-ray imaging system 100 by pressing the start button 84 on the setting screen 80 (see FIG. 4).

[0074] For example, when an operator such as a doctor or technician wants to check whether the voice recognition processing (voice recognition function) by the control unit 8 is operating normally before an examination on a subject using the X-ray imaging system 100, the operator can start a voice recognition test by pressing the start button 84 on the setting screen 80 (see FIG. 4). If a test start keyword recognition image 70 (see FIG. 5) and a test instruction keyword recognition image 71 (see FIG. 6) are displayed in the first display area 13 after the voice recognition test, the operator can check that the voice recognition processing (voice recognition function) by the control unit 8 is operating normally.

[0075] Furthermore, for example, an operator such as a serviceman who installs and maintains the X-ray imaging system 100 can start a voice recognition test by pressing the start button 84 on the setting screen 80 (see FIG. 4 ) to check whether the voice recognition processing (voice recognition function) by the control unit 8 operates normally when installing or maintaining the X-ray imaging system 100. If a test start keyword recognition image 70 (see FIG. 5 ) and a test instruction keyword recognition image 71 (see FIG. 6 ) are displayed in the first display area 13 after the voice recognition test, the operator can check that the voice recognition processing (voice recognition function) by the control unit 8 operates normally.

[0076] (Control of the voice recognition test mode by the control unit) The control unit 8 is configured to control, in the voice recognition test mode, to play back and output a sample voice of the start keyword 86a (see FIG. 4) from the voice output unit 15 based on an operation by the operator. Specifically, the playback output control unit 8a plays back and outputs the sample voice of "Trinious" from the voice output unit 15 based on the operator pressing the play button 88 for the sample voice of "Trinious" as the start keyword 86a.

[0077] The control unit 8 is configured to control the start of voice recognition for the instruction keyword 86b (see FIG. 4) by performing voice recognition on the sample voice of the start keyword 86a (see FIG. 4) reproduced and output from the voice output unit 15. Specifically, the keyword determination processing unit 8c controls the start of voice recognition for the instruction keyword 86b by performing voice recognition on the sample voice of "Trinious" as the start keyword 86a reproduced and output from the voice output unit 15 and input to the voice input unit 5.

[0078] The control unit 8 is configured to control the display device 10 to display the result of speech recognition of the sample speech of the start keyword 86 a (see FIG. 4 ) reproduced and output from the speech output unit 15. Specifically, when the keyword determination processing unit 8 c performs speech recognition on the sample speech of "Trinious," which is the start keyword 86 a, the display control unit 8 b displays an image of "TEST:Trinious," which is a test start keyword recognition image 70 (see FIG. 5 ), in the first display area 13. Note that when the keyword determination processing unit 8 c performs speech recognition on the sample speech of "Trinious," which is the start keyword 86 a, the display control unit 8 b may display the test start keyword recognition image 70 in the first display area 13 and may also cause the speech output unit 15 to output an alert sound indicating that the sample speech of the start keyword 86 a has been recognized.

[0079] The control unit 8 is configured to control the audio output unit 15 to play back and output the sample audio of the instruction keyword 86b (see FIG. 4) based on an operation by the operator. Specifically, as an example, the playback output control unit 8a plays back and outputs the sample audio of "next image" through the audio output unit 15 based on the operator pressing the playback button 88 of the sample audio of "next image" as the instruction keyword 86b.

[0080] The control unit 8 is configured to perform voice recognition on the sample voice of the instruction keyword 86b (see FIG. 4 ) reproduced and output from the voice output unit 15, and thereby control the display device 10 to display a result of the voice recognition of the sample voice of the instruction keyword 86b. Specifically, when the keyword determination processing unit 8c has voice recognized the sample voice of "next image," which is the instruction keyword 86b, the display control unit 8b displays an image of "TEST: next image," which is a test instruction keyword recognition image 71 (see FIG. 6 ), in the first display area 13. Note that when the keyword determination processing unit 8c has voice recognized the sample voice of "next image," which is the instruction keyword 86b, the display control unit 8b may display the test instruction keyword recognition image 71 in the first display area 13 and may also cause the voice output unit 15 to output an alarm sound indicating that the sample voice of the instruction keyword 86b has been voice recognized.

[0081] (Command Keyword 86b and Control Based on Command Keyword 86b) The command keyword 86b (see FIG. 4) and control based on the command keyword 86b in this embodiment will be described with reference to Fig. 8. Note that the command keyword 86b and control based on the command keyword 86b described below are examples, and are not limited to the examples below.

[0082] The command keywords 86 b include keywords 86 relating to the fluoroscopic images to be stored in the storage unit 9 , and the control based on the command keywords 86 b includes control relating to the fluoroscopic images to be stored in the storage unit 9 .

[0083] For example, control based on the command keyword 86b is control to store the fluoroscopic images generated by the image processing unit 7 in the storage unit 9 in the format of a moving image. In this case, an operator such as a doctor or technician utters "Trinious" as the start keyword 86a and then utters "Save" as the command keyword 86b. The control unit 8 controls based on the command keyword 86b determined by voice recognition processing to store the fluoroscopic images generated by the image processing unit 7 in the storage unit 9 in the format of a moving image.

[0084] Further, for example, control based on the command keyword 86b is control to store in the storage unit 9 in the form of a moving image fluoroscopic images to be generated by the image processing unit 7 as a result of fluoroscopic imaging being performed by the imaging unit 4. In this case, an operator such as a doctor or technician utters "Trinious" as the start keyword 86a and then utters "Save now" as the command keyword 86b. As control based on the command keyword 86b determined by voice recognition processing, the control unit 8 stores in the storage unit 9 in the form of a moving image fluoroscopic images to be generated by the image processing unit 7 as a result of fluoroscopic imaging being performed by the imaging unit 4.

[0085] Furthermore, for example, control based on the command keyword 86b is control for storing the final frame image of the fluoroscopic images generated by the image processing unit 7 as a result of fluoroscopic imaging in the storage unit 9. In this case, an operator such as a doctor or technician utters "Trinious" as the start keyword 86a and then utters "Save one image" as the command keyword 86b. The control unit 8 controls based on the command keyword 86b determined by voice recognition processing to store the final frame image of the fluoroscopic images generated by the image processing unit 7 as a result of fluoroscopic imaging in the storage unit 9.

[0086] Furthermore, for example, control based on the instruction keyword 86b is control to register a frame selected in a moving image as a reference image. In this case, an operator such as a doctor or technician utters "Torienious" as the start keyword 86a, and then utters "Reference registration" as the instruction keyword 86b. The control unit 8 registers the selected image as a reference image as control based on the instruction keyword 86b determined by voice recognition processing.

[0087] Furthermore, for example, control based on the command keyword 86b is control to enable an imaging mode in which a moving image captured while a contrast agent is administered and the tabletop 3 is moved is converted into a single long image. In this case, an operator such as a doctor or technician utters "Trinious" as the start keyword 86a and then utters "Score Chase" as the command keyword 86b. As control based on the command keyword 86b determined by voice recognition processing, the control unit 8 enables an imaging mode in which a moving image captured while a contrast agent is administered and the tabletop 3 is moved is converted into a single long image.

[0088] The instruction keywords 86b include keywords 86 relating to the display of the fluoroscopic image on the display device 10, and the control based on the instruction keywords 86b includes control relating to the display of the fluoroscopic image on the display device 10.

[0089] For example, control based on the command keyword 86b is control to enable an imaging mode in which a device fixed image in which the position of a device such as a stent is fixed and displayed based on the position data of a marker in a fluoroscopic image is displayed on the display device 10. In this case, an operator such as a doctor or technician utters "Trinious" as the start keyword 86a and then utters "Stent View" as the command keyword 86b. As control based on the command keyword 86b determined by voice recognition processing, the control unit 8 enables an imaging mode in which a device fixed image in which the position of a device such as a stent is fixed and displayed based on the position data of a marker in a fluoroscopic image is displayed on the display device 10.

[0090] Furthermore, for example, control based on the command keyword 86b is control of an operation related to the display of fluoroscopic images. In this case, an operator such as a doctor or technician utters "Trinious" as the start keyword 86a, and then utters "Previous Image" as the command keyword 86b. As control based on the command keyword 86b determined by voice recognition processing, the control unit 8 causes the display device 10 to display the fluoroscopic image of the moving image that is stored one image earlier among the multiple fluoroscopic images. Furthermore, when the operator such as a doctor or technician utters "Next Image" as the command keyword 86b, the control unit 8 causes the display device 10 to display the fluoroscopic image of the moving image that is stored one image later among the multiple fluoroscopic images.

[0091] Furthermore, an operator such as a doctor or technician may utter "trinous" as a start keyword 86a, and then utter "previous frame" as a command keyword 86b. The control unit 8 controls the display device 10 to display the image of the previous frame in the fluoroscopic images of the moving image based on the command keyword 86b determined by the voice recognition process. Furthermore, when the operator such as a doctor or technician utters "next frame" as the command keyword 86b, the control unit 8 controls the display device 10 to display the image of the next frame in the fluoroscopic images of the moving image.

[0092] Furthermore, an operator such as a doctor or technician may utter "Trinious" as a start keyword 86a, and then utter "Play" as a command keyword 86b. The control unit 8 plays back the fluoroscopic images of the moving image as control based on the command keyword 86b determined by the voice recognition process. Furthermore, when the operator such as a doctor or technician utters "Stop" as the command keyword 86b, the control unit 8 stops playing back the fluoroscopic images of the moving image.

[0093] (Sample Audio Reproduction Output Processing) The sample audio reproduction output processing by the control unit 8 will be described with reference to Fig. 9. The order of the processing steps can be reversed or executed simultaneously as long as there are no contradictions.

[0094] In step S1, if the control unit 8 (playback output control unit 8a) receives an input operation to press the play button 88 for the sample audio of the start keyword 86a or the instruction keyword 86b by operating the operation unit 6 (Yes in step S1), the processing proceeds to step S2; if the control unit 8 (playback output control unit 8a) does not receive an input operation to press the play button 88 for the sample audio of the start keyword 86a or the instruction keyword 86b by operating the operation unit 6 (No in step S1), the processing proceeds to step S1.

[0095] In step S2, the control unit 8 (playback output control unit 8a) plays back and outputs the sample audio of the start keyword 86a or instruction keyword 86b corresponding to the pressed play button 88 through the audio output unit 15. Then, the process ends.

[0096] (Speech recognition test processing in speech recognition test mode) The speech recognition test processing in the speech recognition test mode by the control unit 8 will be described with reference to Fig. 10. The speech recognition test processing in the speech recognition test mode by the control unit 8 is started when the user presses the start button 84 on the setting screen 80, and is ended when the user presses the end button on the setting screen 80. The order of the processing steps can be reversed or executed simultaneously as long as there are no contradictions between them.

[0097] In step S10, the control unit 8 (playback output control unit 8a) accepts an input operation of pressing the playback button 88 of the sample audio of the start keyword 86a by operating the operation unit 6.

[0098] In step S11, the control unit 8 (playback output control unit 8a) causes the audio output unit 15 to play back and output the sample audio of the start keyword 86a.

[0099] In step S12, if the control unit 8 (keyword determination processing unit 8c) has recognized the sample voice of the start keyword 86a that was played back and output by the voice output unit 15 and input to the voice input unit 5 before the predetermined time has elapsed, the control unit 8 (keyword determination processing unit 8c) proceeds to step S13. If the control unit 8 (keyword determination processing unit 8c) has not recognized the sample voice of the start keyword 86a before the predetermined time has elapsed, the control unit 8 proceeds to step S14.

[0100] In step S13 , the control unit 8 (display control unit 8 b ) displays the test start keyword recognition image 70 in the first display area 13 .

[0101] In step S14 , the control unit 8 (display control unit 8 b ) displays the keyword unrecognized image 72 in the first display area 13 .

[0102] In step S15, the control unit 8 (playback output control unit 8a) accepts an input operation of pressing the playback button 88 of the sample audio of the instruction keyword 86b by operating the operation unit 6.

[0103] In step S16, the control unit 8 (playback output control unit 8a) causes the audio output unit 15 to play back and output a sample audio of the instruction keyword 86b corresponding to the pressed instruction keyword 86b.

[0104] In step S17, if the control unit 8 (keyword determination processing unit 8c) has recognized the sample voice of the instruction keyword 86b that has been played back and output by the voice output unit 15 and input to the voice input unit 5 before the predetermined time has elapsed, the process proceeds to step S18. On the other hand, if the control unit 8 (keyword determination processing unit 8c) has not recognized the sample voice of the instruction keyword 86b before the predetermined time has elapsed, the process proceeds to step S19.

[0105] In step S18 , the control unit 8 (display control unit 8 b ) displays the test instruction keyword recognition image 71 in the first display area 13 .

[0106] In step S19 , the control unit 8 (display control unit 8 b ) displays the keyword unrecognized image 72 in the first display area 13 .

[0107] (Effects of this embodiment) In this embodiment, the following effects can be obtained.

[0108] In this embodiment, as described above, the control unit 8 is configured to, based on an operation by the operator, play back and output, via the audio output unit 15, a sample audio of the keyword 86 for performing control based on the keyword 86 by performing voice recognition. As a result, the sample audio of the keyword 86 is played back and output by the audio output unit 15, so that the operator can listen to and hear the sample audio of the keyword 86 for performing predetermined control in the X-ray imaging system 100 by voice recognition. Therefore, the operator can easily understand the pronunciation and accent of the keyword 86 for performing predetermined control by voice recognition.

[0109] Furthermore, in this embodiment, the following additional effects can be obtained by the following configuration.

[0110] That is, in this embodiment, as described above, the keyword 86 includes the command keyword 86b for executing the corresponding predetermined function and the start keyword 86a that serves as a trigger for voice recognition of the command keyword 86b, and the sample audio of the keyword 86 includes a sample audio of the start keyword 86a and a sample audio of the command keyword 86b. This allows the operator to easily grasp the pronunciation and accent of both the sample audio of the start keyword 86a that serves as a trigger for voice recognition of the command keyword 86b and the command keyword 86b for executing the corresponding predetermined function.

[0111] Furthermore, in this embodiment, as described above, the display device 10 is provided which displays, side by side, a plurality of keywords 86 and play buttons 88 that correspond to each of the plurality of keywords 86 and play back sample audio for each of the plurality of keywords 86. As a result, the display device 10 displays, side by side, a plurality of keywords 86 and play buttons 88 that correspond to each of the plurality of keywords 86, thereby improving the operability of playing back sample audio for keywords 86 whose pronunciation or accent is to be understood.

[0112] Furthermore, in this embodiment, as described above, the control unit 8 is configured to determine whether the test voice of the keyword 86, which is played back and output from the voice output unit 15 based on an operator's operation and received by the voice input unit 5, has been recognized in the voice recognition test mode in which a test is performed to determine whether the voice of the keyword 86 has been recognized. As a result, in the voice recognition test mode, it is determined whether the test voice of the keyword 86, which is played back and output from the voice output unit 15 and received by the voice input unit 5, has been recognized. Therefore, even if the X-ray imaging system 100 is installed in a different environment, a voice recognition test can be performed under substantially the same conditions with respect to the pronunciation and accent of the voice based on the playback output of the test voice of the keyword 86. Therefore, based on the voice recognition test, it is possible to properly check whether the voice recognition process (voice recognition function) is operating normally, and to distinguish whether a malfunction of the voice recognition process (voice recognition function) is caused by a defect in the voice recognition process program or by the voice of the operator or the environment of the X-ray imaging system 100.

[0113] Furthermore, in this embodiment, as described above, the test audio includes sample audio, and the control unit 8 is configured to, in the voice recognition test mode, concurrently perform control to play and output the sample audio from the voice output unit 15 based on an operation by the operator, and control to determine whether or not the sample audio played and output from the voice output unit 15 and received by the voice input unit 5 has been recognized. This allows the control to play and output the sample audio and the control to determine whether or not the sample audio received by the voice input unit 5 has been recognized to be performed concurrently. Therefore, in the voice recognition test, the determination of whether or not voice recognition has been performed can be made based on the sample audio played and output from the voice output unit 15 and received by the voice input unit 5, rather than determining which sample audio has been played and output by internal processing of the control unit 8. Therefore, based on the voice recognition test, it is possible to accurately check whether or not the voice recognition process (voice recognition function) is operating normally.

[0114] Furthermore, in this embodiment, as described above, the notification unit 1 is provided, and the control unit 8 is configured to cause the notification unit 1 to notify the result of speech recognition of the sample speech that is played back and output from the speech output unit 15 and received by the speech input unit 5 in the speech recognition test mode. As a result, the notification unit 1 notifies the result of speech recognition of the sample speech in the speech recognition test mode, and therefore, based on the notification by the notification unit 1, it can be easily known that the speech recognition process (speech recognition function) is operating normally.

[0115] Furthermore, in this embodiment, as described above, the notification unit 1 includes the display device 10, and the sample audio of the keyword 86 includes a sample audio of the start keyword 86a and a sample audio of the instruction keyword 86b. The control unit 8 is configured to perform the following controls in the voice recognition test mode: control to have the voice output unit 15 play and output the sample audio of the start keyword 86a based on an operation by the operator; control to start voice recognition for the instruction keyword 86b by voice recognizing the sample audio of the start keyword 86a played and output from the voice output unit 15; control to have the display device 10 display a result of the voice recognition of the sample audio of the start keyword 86a played and output from the voice output unit 15; control to have the voice output unit 15 play and output the sample audio of the instruction keyword 86b based on an operation by the operator; and control to have the display device 10 display a result of the voice recognition of the sample audio of the instruction keyword 86b by voice recognizing the sample audio of the instruction keyword 86b played and output from the voice output unit 15. As a result, the display device 10 displays the results of voice recognition of the sample voices of the start keyword 86a and the instruction keyword 86b reproduced and output from the voice output unit 15, making it easier to understand that the voice recognition process (voice recognition function) is operating normally for each of the start keyword 86a and the instruction keyword 86b.

[0116] In this embodiment, as described above, the voice input unit 5 and the voice output unit 15 are provided in the display device 10. This allows the sample voices of the start keyword 86 a and the instruction keyword 86 b reproduced and output from the voice output unit 15 to be easily input to the voice input unit 5.

[0117] [Modifications] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the description of the above-mentioned embodiments, and further includes all modifications (modifications) within the meaning and scope of the claims.

[0118] For example, in the above embodiment, the sample audio of the keyword reproduced and output by the audio output unit includes the sample audio of the start keyword and the sample audio of the instruction keyword, but the present invention is not limited to this. For example, the sample audio of the keyword reproduced and output by the audio output unit may be either the sample audio of the start keyword or the sample audio of the instruction keyword.

[0119] In the above embodiment, the display device displays a plurality of keywords and a play button for playing back a sample sound corresponding to each of the plurality of keywords, but the present invention is not limited to this. For example, the display device may be configured to display a plurality of keywords and a single play button, and, when playing back a sample sound corresponding to the keyword, the operator may operate the operation unit to select one of the plurality of keywords and then press the play button.

[0120] In the above embodiment, the test audio includes a sample audio, and the control unit is configured to determine whether or not the sample audio of the keyword has been recognized in the audio recognition test mode. However, the present invention is not limited to this. For example, the test audio may include audio for a speech recognition test that is different from the sample audio of the keyword and is stored in the storage unit, and the control unit may be configured to determine whether or not the speech recognition test audio that is played back and output from the audio output unit and received by the audio input unit in response to an operation by the operator has been recognized in the audio recognition test mode.

[0121] In the above embodiment, the notification unit includes a display device, and the control unit is configured to cause the display device to display the result of speech recognition of the sample speech in the speech recognition test mode, but the present invention is not limited to this. For example, the notification unit may be a speaker separate from the audio output unit, earphones worn in the user's ears, or headphones worn on the user's head, and the control unit may be configured to notify the result of speech recognition of the sample speech by sound via the speaker, earphones, or headphones.

[0122] In the above embodiment, the control unit is configured to start speech recognition for the command keyword by performing speech recognition on the sample speech of the start keyword in the speech recognition test mode, but the present invention is not limited to this. For example, the control unit may be configured to start speech recognition for the command keyword in the speech recognition test mode without performing speech recognition on the sample speech of the start keyword. That is, in the speech recognition test, the playback output of the sample speech of the start keyword, which serves as a trigger for speech recognition of the command keyword, may be omitted, and the control unit may be configured to determine whether the sample speech of the command keyword received by the speech input unit has been recognized based only on the playback output of the sample speech of the command keyword.

[0123] In addition, although the above embodiment shows an example in which the audio input unit and the audio output unit are provided in the display device, the present invention is not limited to this. For example, one or both of the audio input unit and the audio output unit may not be provided in the display device.

[0124] In the above embodiment, the audio languages ​​of the keywords to be voice-recognized in the voice-recognition-enabled mode are English (American) and Japanese, and the sample audio of the start keyword and the instruction keyword are English (American) and Japanese, but the present invention is not limited to this. For example, the audio languages ​​of the keywords to be voice-recognized in the voice-recognition-enabled mode and the sample audio may be English (British), Chinese, Korean, Spanish, or Portuguese, without being limited to English (American) and Japanese.

[0125] Aspects It will be appreciated by those skilled in the art that the exemplary embodiments described above are examples of the following aspects.

[0126] (Item 1) An X-ray imaging system comprising: an X-ray irradiation unit; an X-ray detection unit that detects X-rays irradiated from the X-ray irradiation unit; an audio input unit that accepts audio input; an audio output unit; and a control unit that performs control based on the keyword by voice recognition of the keyword based on the audio accepted by the audio input unit, wherein the control unit is configured to play and output a sample audio of the keyword by the audio output unit based on an operation by an operator.

[0127] (Item 2) The X-ray imaging system described in Item 1, wherein the keywords include instruction keywords for executing corresponding predetermined functions and start keywords that trigger voice recognition of the instruction keywords, and the sample voices of the keywords include the sample voice of the start keywords and the sample voice of the instruction keywords.

[0128] (Item 3) The X-ray imaging system according to item 1 or 2, further comprising a display device that displays a plurality of the keywords and playback buttons that correspond to each of the plurality of keywords and play back the sample audio for each of the plurality of keywords.

[0129] (Item 4) The X-ray imaging system according to any one of Items 1 to 3, wherein the control unit is configured to determine whether test voice of the keyword, which is played back and output from the voice output unit based on an operation by the operator and received by the voice input unit, has been recognized in a voice recognition test mode for testing whether voice recognition of the keyword has been achieved.

[0130] (Item 5) The X-ray imaging system described in Item 4, wherein the test audio includes the sample audio, and the control unit is configured to, in the voice recognition test mode, concurrently perform control to play and output the sample audio from the audio output unit based on the operation of the operator, and control to determine whether the sample audio that is played and output from the audio output unit and received by the voice input unit has been voice recognized.

[0131] (Item 6) The X-ray imaging system according to Item 5, further comprising an alarm unit, wherein the control unit is configured to, in the voice recognition test mode, cause the alarm unit to notify the result of voice recognition of the sample voice that is played back and output from the voice output unit and received by the voice input unit.

[0132] (Item 7) The X-ray imaging system according to Item 6, wherein the notification unit includes a display device, and the sample audio of the keyword includes the sample audio of a start keyword and the sample audio of a command keyword, and the control unit is configured to perform the following controls in the voice recognition test mode: control to have the voice output unit play and output the sample audio of the start keyword based on an operation by the operator, control to start the voice recognition for the command keyword by performing voice recognition on the sample audio of the start keyword played and output from the voice output unit, control to have the display device display a result of the voice recognition of the sample audio of the start keyword played and output from the voice output unit, control to have the voice output unit play and output the sample audio of the command keyword based on an operation by the operator, and control to have the display device display a result of the voice recognition of the sample audio of the command keyword by performing voice recognition on the sample audio of the command keyword played and output from the voice output unit.

[0133] (Item 8) The X-ray imaging system according to Item 7, wherein the audio input unit and the audio output unit are provided in the display device.

[0134] REFERENCE SIGNS LIST 1 Notification unit 5 Audio input unit 7 Image processing unit 8 Control unit 9 Storage unit 10 Display device 15 Audio output unit 40 X-ray irradiation unit 41 X-ray detection unit 86 Keyword 86a Start keyword 86b Instruction keyword 88 Play button 100 X-ray imaging system

Claims

1. An X-ray imaging system comprising: an X-ray irradiation unit; an X-ray detection unit that detects X-rays irradiated from the X-ray irradiation unit; an audio input unit that accepts audio input; an audio output unit; and a control unit that performs control based on the keyword by voice recognition of the keyword based on the audio accepted by the audio input unit, wherein the control unit is configured to play and output a sample audio of the keyword through the audio output unit based on operation by an operator.

2. The X-ray imaging system of claim 1, wherein the keywords include an instruction keyword for executing a corresponding predetermined function and a start keyword that triggers voice recognition of the instruction keyword, and the sample audio of the keywords includes the sample audio of the start keyword and the sample audio of the instruction keyword.

3. The X-ray imaging system according to claim 1, further comprising a display device that displays a plurality of the keywords and a playback button that corresponds to each of the plurality of keywords and plays back the sample audio of the plurality of keywords.

4. The X-ray imaging system of claim 1, wherein the control unit is configured to determine whether or not the test voice of the keyword, which is played back and output from the voice output unit based on the operator's operation and received by the voice input unit, has been recognized in a voice recognition test mode that tests whether or not the voice of the keyword has been recognized.

5. The X-ray imaging system of claim 4, wherein the test audio includes the sample audio, and the control unit is configured to, in the voice recognition test mode, concurrently control the audio output unit to play and output the sample audio based on the operator's operation, and control the audio output unit to determine whether or not the sample audio that has been played and output and received by the voice input unit has been recognized.

6. An X-ray imaging system as described in claim 5, further comprising an alarm unit, wherein the control unit is configured to, in the voice recognition test mode, cause the alarm unit to notify the result of voice recognition of the sample voice that is played back and output from the voice output unit and received by the voice input unit.

7. The X-ray imaging system of claim 6, wherein the notification unit includes a display device, and the sample audio of the keyword includes the sample audio of a start keyword and the sample audio of a command keyword, and the control unit is configured to perform the following controls in the voice recognition test mode: control to play and output the sample audio of the start keyword by the voice output unit based on an operation by the operator; control to start the voice recognition for the command keyword by performing voice recognition on the sample audio of the start keyword played and output from the voice output unit; control to display on the display unit a result of the voice recognition of the sample audio of the start keyword played and output from the voice output unit; control to play and output the sample audio of the command keyword by the voice output unit based on an operation by the operator; and control to display on the display unit a result of the voice recognition of the sample audio of the command keyword by performing voice recognition on the sample audio of the command keyword played and output from the voice output unit.

8. The X-ray imaging system according to claim 7, wherein the audio input section and the audio output section are provided in the display device.

Citation Information

Patent Citations

  • Information terminal device

    JP2020005158A

  • Voice recognition input device, voice recognition input program, and medical image capturing system

    JP2020089641A

  • Language independent and voice operated information management system

    US20030033152A1

  • System and method for voice control of cabinet x-ray systems

    US20180228010A1