Information Processing Apparatus, Information Processing Method, and Program
The information processing apparatus addresses the challenge of ambiguous voice commands by using user-specific speech pattern parameters to ensure accurate and consistent voice operation in devices.
Patent Information
- Application Number
- JP2022509520
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-23
- Filing Date
- 2021-03-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-03-09
AI Technical Summary
Existing voice-controlled devices struggle to accurately interpret voice commands that include ambiguous words like 'more' and 'extremely', leading to erratic device operations.
An information processing apparatus that includes a command processing unit, which executes processing according to voice commands containing predetermined ambiguous words by using parameters based on the user's way of speaking, thereby adjusting operations accordingly.
Enables seamless voice operations using natural expressions with ambiguous words, ensuring consistent and accurate device control by personalizing parameter settings based on the user's speech patterns.
Smart Images

Figure 0007697455000001 
Figure 0007697455000002 
Figure 0007697455000003
Abstract
Description
Technical Field
[0001] The present technology relates to an information processing apparatus, an information processing method, and a program, and more particularly to an information processing apparatus, an information processing method, and a program that enable voice operations using natural expressions.
Background Art
[0002] In recent years, devices that can be operated by voice have been increasing. For example, Patent Document 1 describes a television receiver incorporating a voice recognition device that analyzes the speech content of a user.
[0003] According to the television receiver described in Patent Document 1, a user can request the presentation of certain information by a voice command and view the information presented in response.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Generally, in natural conversations, people may express the degree of things using ambiguous words such as "more" and "extremely".
[0006] When such voice including ambiguous words is used as a voice command for a device equipped with a voice UI function, the variation in the operation of the device becomes large. Therefore, it is difficult to use such ambiguous words as voice commands.
[0007] The present technology has been made in view of such a situation, and enables voice operations using natural expressions.
Means for Solving the Problems
[0008] An information processing apparatus according to an aspect of the present technology includes a command processing unit that executes processing according to the voice command when a predetermined word determined to have an ambiguous degree of control is included in the voice command for instructing control of a device input by a user, using parameters according to the way of speaking of the user when the voice command is input.
[0009] In an aspect of the present technology, when a voice command for instructing control of a device input by a user includes a predetermined word determined to have an ambiguous degree of control, processing according to the voice command is executed using parameters according to the way of speaking of the user when the voice command is input.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0011] Hereinafter, embodiments for implementing the present technology will be described. The description will be carried out in the following order. 1. Voice operation using ambiguous words 2. Configuration of the imaging device 3. Operation of the imaging device 4. Regarding other embodiments 5. Regarding the computer
[0012] <1. Voice operation using ambiguous words> FIG. 1 is a diagram showing a usage example of an imaging device 11 according to an embodiment of the present technology.
[0013] The imaging device 11 is a camera that can be operated by a voice UI (User Interface). The imaging device 11 is provided with a microphone (not shown) for collecting the voice uttered by the user. The user can perform various operations such as setting shooting parameters by speaking to the imaging device 11 and inputting a voice command. The voice command is information for instructing the control of the imaging device 11.
[0014] In the example of FIG. 1, although the imaging device 11 is a camera, it is also possible to use other devices having an imaging function such as a smartphone, a tablet terminal, or a PC as the imaging device 11.
[0015] As shown in FIG. 1, a liquid crystal monitor 21 is provided on the back surface of the housing of the imaging device 11. On the liquid crystal monitor 21, for example, a live view image that displays the image captured by the imaging device 11 in real time before shooting a still image is displayed. The user who is the photographer can perform a shooting operation using a voice command while checking the shooting angle, color tone, etc. by looking at the live view image displayed on the liquid crystal monitor 21.
[0016] As shown in the speech bubble #1, for example, when the user says "Make the color of the cherry blossoms more pink," the imaging device 11 performs speech recognition and semantic analysis, and performs image processing to adjust the color tone of the cherry blossoms shown in the image to pink according to the user's speech.
[0017] In this way, in natural conversation, people sometimes use ambiguous words such as "more" and "extremely" to express a degree. Since ambiguous words are non-quantitative words such that the degree they represent varies from person to person, when a voice command containing such a word is input, usually the operation of the device becomes more erratic.
[0018] In the imaging device 11 of FIG. 1, words such as "more" and "extremely," whose degree of control is non-quantitative, are specified in advance as ambiguous designation words. When the voice command includes an ambiguous designation word, the imaging device 11 performs image processing using parameters set according to the user's way of speaking when the voice command was input.
[0019] For example, if the normal way of speaking is set as the reference way of speaking, image processing is performed using parameters set based on the difference between the user's way of speaking when the voice command was input and the normal way of speaking. In this way, the imaging device 11 functions as an information processing device that performs image processing using parameters set according to the user's way of speaking when the voice command was input.
[0020] FIG. 2 is a diagram showing an example of image processing according to the user's way of speaking.
[0021] The image processing shown in FIG. 2 is the processing when the user makes a statement of "Make the color of the cherry blossoms more pink," that is, when a voice command for adjusting the color is input. The voice command input by the user includes the ambiguous designation word "more."
[0022] When a voice command for color adjustment is input, in the imaging device 11, it is determined whether the way the user speaks when inputting the voice command is different from the normal way of speaking.
[0023] For example, as shown in A of FIG. 2, when it is determined that the way the user speaks is the same as the normal way of speaking, as shown at the tip of arrow A1, the imaging device 11 adjusts the color tone of the cherry blossoms shown in the image to pink by a predetermined degree according to the voice command. In A of FIG. 2, the fact that the light color is painted on the cherry blossoms indicates that the color tone of the cherry blossoms shown in the image is adjusted to pink by a predetermined degree.
[0024] On the other hand, as shown in B of FIG. 2, when it is determined that the way the user speaks is different from the normal way of speaking, as shown at the tip of arrow A2, the imaging device 11 extremely adjusts the color tone of the cherry blossoms shown in the image to pink according to the voice command.
[0025] That is, when the way the user speaks is different from the normal way of speaking, the imaging device 11 adjusts the color tone with an adjustment amount larger than the adjustment amount when the way the user speaks is the same as the normal way of speaking. In B of FIG. 2, the fact that the dark color is painted on the cherry blossoms indicates that the color tone of the cherry blossoms shown in the image is extremely adjusted to pink.
[0026] In this way, in the imaging device 11, a parameter representing the degree of image processing is set according to whether the way the user speaks when inputting the voice command is different from the normal way of speaking. Not only the color tone of the image, but also other settings such as the frame rate, the amount of blur, and the brightness can be adjusted in the same way using a voice command including an ambiguous designation word.
[0027] As a result, the user who is the photographer can operate the imaging device 11 by voice including natural expressions using ambiguous words such as "more" and "extremely" as if giving instructions to a camera assistant.
[0028] When the user adjusts the parameters related to shooting while observing the operation of the imaging device 11, the user can adjust the parameters without specifically specifying numerical values, making the operation easy.
[0029] The user can easily use voice commands related to adjusting sensory expressions such as color tone, frame rate, blurring, and brightness (luminance).
[0030] <2. Configuration of Imaging Device> FIG. 3 is a block diagram showing a configuration example of the imaging device 11.
[0031] As shown in FIG. 3, the imaging device 11 includes an operation input unit 31, a voice command processing unit 32, an imaging unit 33, a signal processing unit 34, an image data storage unit 35, a recording unit 36, and a display unit 37.
[0032] The operation input unit 31 is composed of buttons, a touch panel monitor, a controller, a remote controller, etc. The operation input unit 31 detects a camera operation by the user and outputs an operation instruction representing the content of the detected camera operation. The operation instruction output from the operation input unit 31 is appropriately supplied to each component of the imaging device 11.
[0033] The voice command processing unit 32 includes a voice command input unit 51, a voice signal processing unit 52, a voice command recognition unit 53, a voice command semantic analysis unit 54, a user feature determination unit 55, a user feature storage unit 56, a parameter value storage unit 57, and a voice command execution unit 58.
[0034] The voice command input unit 51 is composed of a sound collecting device such as a microphone. The voice command input unit 51 collects the voice uttered by the user and outputs a voice signal to the voice signal processing unit 52.
[0035] Note that a microphone different from the microphone mounted on the imaging device 11 may be used to collect the voice uttered by the user. It is possible to collect the voice uttered by the user by an external device connected to the imaging device 11, such as a pin microphone or a microphone provided in another device.
[0036] The voice signal processing unit 52 performs signal processing such as noise reduction on the voice signal supplied from the voice command input unit 51, and outputs the voice signal after the signal processing to the voice command recognition unit 53.
[0037] The voice command recognition unit 53 performs voice recognition on the voice signal supplied from the voice signal processing unit 52 to detect a voice command. The voice command recognition unit 53 outputs the detection result of the voice command and the voice signal to the voice command meaning analysis unit 54.
[0038] The voice command meaning analysis unit 54 analyzes the meaning of the voice command detected by the voice command recognition unit 53, and determines whether the voice command input by the user includes an ambiguous designation word.
[0039] When the voice command meaning analysis unit 54 determines that the voice command includes an ambiguous designation word, the voice command meaning analysis unit 54 outputs the analysis result of the meaning of the voice command and the voice signal supplied from the voice command recognition unit 53 to the user feature determination unit 55. In addition, the voice command meaning analysis unit 54 outputs the analysis result of the meaning of the voice command to the voice command execution unit 58.
[0040] Instead of determining whether the ambiguous designation word itself is included in the voice command, it may be determined whether a word similar to the ambiguous designation word is included in the voice command. For example, when "more" is designated as the ambiguous designation word, words such as "a little more" and "just a bit more" are determined as words similar to the ambiguous designation word.
[0041] When a word similar to the ambiguous designation word is included in the voice command, the same processing as when the ambiguous designation word is included in the voice command is performed in each part.
[0042] In this way, in the voice command meaning analysis unit 54, it is determined whether a predetermined word with an ambiguous degree of control, including the ambiguous designation word and a word similar to it, is included in the voice command.
[0043] The user feature determination unit 55 analyzes the voice signal supplied from the voice command meaning analysis unit 54 and extracts feature amounts. Also, the user feature determination unit 55 reads out the feature amounts of the reference voice signal from the user feature storage unit 56. In the user feature storage unit 56, for example, the feature amounts of the voice signal of the user's usual way of speaking are stored as the feature amounts of the reference voice signal.
[0044] The user feature determination unit 55 compares the feature amounts of the voice signal supplied from the voice command meaning analysis unit 54 with the feature amounts of the reference voice signal, and determines whether the way of speaking of the user when inputting the voice command is different from the usual way of speaking.
[0045] FIG. 4 is a diagram showing an example of a way of speaking different from the usual way of speaking.
[0046] The way of speaking is specified, for example, by intonation, emotion, and diction. Whether the intonation, emotion, and diction when inputting the voice command are different from the usual intonation, emotion, and diction is determined by the user feature determination unit 55.
[0047] Instead of using all of intonation, emotion, and diction, the way of speaking may be specified based on at least any one of intonation, emotion, and diction. The way of speaking may also be specified by other elements such as the user's expression and attitude.
[0048] The intonation is specified by, for example, the speed, volume, and tone of the voice. When the speed of the voice is different from the reference speed, when the volume of the voice is different from the reference volume, or when the tone of the voice is different from the reference tone, it is determined that the user's way of speaking is different from the normal way of speaking.
[0049] The intonation may be specified by, for example, the pitch represented by the frequency of the voice signal, the timbre represented by the waveform of the voice signal, and the like.
[0050] Emotions are specified by performing emotion estimation based on the voice signal. When it is specified that the user has negative emotions such as anger and anxiety, it is determined that the user's way of speaking is different from the normal way of speaking. The emotions of the user may be estimated based on an image obtained by imaging the state of the user when the voice command is input.
[0051] Word usage is specified based on the result of semantic analysis and the like. When it is specified that the user is using negative word usage such as "What?", "Don't you understand?", it is determined that the user's way of speaking is different from the normal way of speaking.
[0052] Based on such a determination result, the user feature determination unit 55 in FIG. 3 sets parameters used when executing processing according to the voice command, and stores the set values of the parameters in the parameter value storage unit 57. That is, the user feature determination unit 55 also functions as a parameter setting unit that sets parameters.
[0053] In addition, the user feature determination unit 55 stores the feature amount of the voice signal supplied from the voice command semantic analysis unit 54 in the user feature storage unit 56.
[0054] The feature amount of the voice signal stored in the user feature storage unit 56 is used for determination when the next voice command is input. The more the feature amounts stored in the user feature storage unit 56 increase, the higher the accuracy of the determination by the user feature determination unit 55.
[0055] Note that the feature amount for each user may be stored in the user feature storage unit 56. In this case, at the timing such as when the imaging device 11 is activated, the fingerprint is read to perform user login, and the determination is made using the feature amount prepared for the logged-in user.
[0056] The user feature storage unit 56 is configured by an internal memory. The user feature storage unit 56 stores the feature amount of the user's voice signal. The user feature storage unit 56 may be provided in a device external to the imaging device 11, such as a server device on the cloud.
[0057] Note that the determination by the user feature determination unit 55 may be made based on an image obtained by imaging the user, rather than being made based on the voice signal. In this case, the user feature storage unit 56 stores the feature amount of the image obtained by imaging the user's state when speaking in a normal manner. The user feature determination unit 55 determines whether the way the user speaks when the voice command is input is different from the normal way of speaking, based on the image obtained by imaging the user's state when the voice command is input. Note that the state of the user when the voice command is input is imaged by, for example, an in-camera mounted on the imaging device 11.
[0058] Also, the determination by the user feature determination unit 55 may be made based on the sensor data detected by the wearable sensor worn by the user. In this case, the user feature storage unit 56 stores the feature amount of the sensor data detected by the wearable sensor when speaking in a normal manner. The user feature determination unit 55 determines whether the way the user speaks is different from the normal way of speaking, based on the sensor data detected when the voice command is input.
[0059] The parameter value storage unit 57 stores the set value of the parameter set by the user feature determination unit 55.
[0060] The voice command execution unit 58 reads the set value of the parameter from the parameter value storage unit 57. Based on the analysis result supplied from the voice command semantic analysis unit 54, the voice command execution unit 58 executes processing corresponding to the voice command input by the user using the parameter read from the parameter value storage unit 57.
[0061] For example, when a voice command representing adjusting the color tone of an image is input, the voice command execution unit 58 causes the signal processing unit 34 to perform image processing for adjusting the color tone of the image using the parameter set by the user feature determination unit 55.
[0062] The imaging unit 33 is composed of an image sensor or the like. The imaging unit 33 converts the received light into an electrical signal and captures an image. The image captured by the imaging unit 33 is output to the signal processing unit 34.
[0063] The signal processing unit 34 performs various signal processes on the image supplied from the imaging unit 33 according to the control by the voice command execution unit 58. In the signal processing unit 34, various image processes such as noise reduction, correction processing, demosaicing, and processing for adjusting the appearance of the image are performed. The image on which the image process has been performed is supplied to the image data storage unit 35.
[0064] The image data storage unit 35 is composed of a DRAM (Dynamic Random Access Memory), an SRAM (Static Random Access Memory), or the like. The image data storage unit 35 temporarily stores the image supplied from the signal processing unit 34. The image data storage unit 35 outputs the image to the recording unit 36 or the display unit 37 according to the operation by the user.
[0065] The recording unit 36 is composed of an internal memory and a memory card attached to the imaging device 11. The recording unit 36 records the images supplied from the image data storage unit 35. The recording unit 36 may be provided in an external device such as an external HDD (Hard Disk Drive) or a server device on the cloud.
[0066] The display unit 37 is composed of a liquid crystal monitor 21 and a viewfinder. The display unit 37 converts the images supplied from the image data storage unit 35 to an appropriate resolution and displays them.
[0067] <3. Operations of the Imaging Device> Here, the operations of the imaging device 11 having the above configuration will be described.
[0068] First, with reference to the flowchart of FIG. 5, the shooting process will be described. The shooting process in FIG. 5 starts, for example, when a power-on command by the user is input to the operation input unit 31. At this time, the image capture by the imaging unit 33 is started. A live view image is displayed on the display unit 37.
[0069] In step S11, the operation input unit 31 receives a camera operation by the user. For example, operations such as framing and camera settings are performed by the user.
[0070] In step S12, the voice command input unit 51 determines whether voice has been input by the user.
[0071] If it is determined in step S12 that voice has been input, in step S13, the imaging device 11 performs image processing according to the voice command. By the image processing according to the voice command, image processing corresponding to the voice command is performed. Details of the image processing according to the voice command will be described later with reference to the flowchart of FIG. 6.
[0072] On the other hand, if it is determined in step S12 that no voice command has been input, the process of step S13 is skipped.
[0073] In step S14, the operation input unit 31 determines whether the shooting button has been pressed.
[0074] If it is determined in step S14 that the shooting button has been pressed, in step S15, the recording unit 36 records an image. The image captured by the imaging unit 33 and subjected to predetermined image processing by the signal processing unit 34 is supplied from the image data storage unit 35 to the recording unit 36 and recorded.
[0075] On the other hand, if it is determined in step S14 that the shooting button has not been pressed, the process of step S15 is skipped.
[0076] In step S16, the operation input unit 31 determines whether it has received a power-off command from the user.
[0077] If it is determined in step S16 that it has not received a power-off command, the process returns to step S11 and subsequent processing is performed. If it is determined in step S16 that it has received a power-off command, the process ends.
[0078] Next, with reference to the flowchart of FIG. 6, the image processing by the voice command performed in step S13 of FIG. 5 will be described.
[0079] In step S31, the voice signal processing unit 52 performs voice signal processing on the voice signal representing the voice input by the user.
[0080] In step S32, the voice command recognition unit 53 determines whether a voice command has been input based on the voice signal subjected to voice signal processing.
[0081] For example, when the specific word, which is a word for specifying a voice command, is included in the voice signal, the voice command recognition unit 53 determines that a voice command has been input. Also, when voice is input by the user while a predetermined button is pressed, the voice command recognition unit 53 determines that a voice command has been input.
[0082] When it is determined in step S32 that a voice command has been input, in step S33, the voice command processing unit 32 performs semantic analysis processing of the voice command. By the semantic analysis processing of the voice command, parameters for executing processing according to the voice command are determined. Details of the semantic analysis processing of the voice command will be described later with reference to the flowchart of FIG. 7.
[0083] In step S34, the signal processing unit 34 performs image processing using the parameters determined by the semantic analysis processing in step S33. After the image subjected to the image processing is stored in the image data storage unit 35, the process returns to step S13 of FIG. 5, and subsequent processing is performed.
[0084] Similarly, when it is determined in step S32 that no voice command has been input, the process returns to step S13 of FIG. 5, and subsequent processing is performed.
[0085] Next, with reference to the flowchart of FIG. 7, the semantic analysis processing of the voice command performed in step S33 of FIG. 6 will be described.
[0086] In step S41, the voice command semantic analysis unit 54 determines whether the voice command input by the user includes an ambiguous designation word.
[0087] When it is determined in step S41 that the voice command includes an ambiguous designation word, in step S42, the user feature determination unit 55 reads out the feature amount of the reference voice signal from the user feature storage unit 56. Also, the user feature determination unit 55 analyzes the voice signal representing the voice input by the user and extracts the feature amount.
[0088] In step S43, the user feature determination unit 55 compares the feature amount of the voice signal representing the voice input by the user with the feature amount of the reference voice signal, and detects the user state based on the difference therebetween.
[0089] In step S44, the user feature determination unit 55 determines whether the user's way of speaking is different from the normal way of speaking based on the determination result in step S43.
[0090] For example, when the user is angry, it is determined that the user's way of speaking is different from the normal way of speaking. Based on other user states such as when the user is speaking fast or when the user is depressed and has negative feelings, it may be determined whether the user's way of speaking is different from the normal way of speaking.
[0091] If it is determined in step S44 that the user's way of speaking when inputting the voice command is the same as the normal way of speaking, in step S45, the user feature determination unit 55 sets the parameters as usual. Specifically, the user feature determination unit 55 adjusts the current setting value by the amount of adjustment preset for the ambiguous designation word, and sets the parameters. For example, when the ambiguous designation word "more" is included in the voice command, the user feature determination unit 55 adjusts the current setting value by +1 and sets the parameters.
[0092] On the other hand, if it is determined in step S44 that the user's way of speaking when inputting the voice command is different from the normal way of speaking, in step S46, the user feature determination unit 55 sets the parameters to be larger than usual. Specifically, the user feature determination unit 55 adjusts the current setting value by an amount of adjustment larger than the amount of adjustment preset for the ambiguous designation word, and sets the parameters. For example, when the ambiguous designation word "more" is included in the voice command, the user feature determination unit 55 adjusts the current setting value by +100 and sets the parameters.
[0093] Note that the adjustment amount of the parameter may be changed according to the difference between the user's way of speaking when the voice command is input and the reference way of speaking.
[0094] In step S47, the user feature determination unit 55 determines the set value of the parameter and stores it in the parameter value storage unit 57.
[0095] In step S48, the user feature determination unit 55 stores the feature amount of the voice signal representing the voice input by the user in the user feature storage unit 56.
[0096] After the feature amount of the voice signal is stored in the user feature storage unit 56, or when it is determined in step S41 that the voice command does not include an ambiguous designation word, the process proceeds to step S49. When the voice command does not include an ambiguous designation word, parameter setting according to the user's way of speaking is not performed.
[0097] In step S49, the voice command execution unit 58 reads out the set value of the parameter from the parameter value storage unit 57 and sets the voice command together with the set value of the parameter in the signal processing unit 34.
[0098] Thereafter, the process returns to step S33 in FIG. 6, and subsequent processing is performed. In the signal processing unit 34, image processing corresponding to the voice command is performed using the parameter set by the voice command execution unit 58.
[0099] Note that after the semantic analysis process in FIG. 7 is performed once, when the voice command for adjusting the same parameter is input again by the user, the adjustment amount at the time of parameter setting may be adjusted. The re-input of the voice command for adjusting the same parameter is performed, for example, when the user is not satisfied with the parameter set according to the previously input voice command.
[0100] In this case, the adjustment amount used in step S45 or step S46 is adjusted to be, for example, a larger adjustment amount. By adjusting the adjustment amount of the parameter, the imaging device 11 is so to speak personalized according to the user's feeling.
[0101] As described above, when the voice input by the user contains ambiguous words, the parameter is adjusted according to the user's way of speaking, and the process corresponding to the voice command is performed. The user can operate the imaging device 11 by voice including natural expressions using ambiguous words such as "more" and "extremely".
[0102] <4. Regarding Other Embodiments> Although the case of performing image processing by voice including an ambiguous designation word has been mainly described, various controls of the device such as control related to imaging, control related to display, and control related to communication may be performed according to voice including an ambiguous designation word.
[0103] Although the operation by voice including an ambiguous designation word is assumed to be performed in the camera, the present technology can be applied to the processing in any device.
[0104] FIG. 8 is a block diagram showing a configuration example of an information processing apparatus 101 to which the present technology is applied.
[0105] The information processing apparatus 101 in FIG. 8 is, for example, a PC used for editing an image captured by a camera. Thus, the present technology is applicable not only to the processing of the live view image in the camera but also to the processing in an apparatus for editing an image stored in a predetermined recording unit.
[0106] In FIG. 8, the same components as those of the imaging device 11 in FIG. 4 are denoted by the same reference numerals. Redundant descriptions will be omitted as appropriate.
[0107] The configuration of the information processing apparatus 101 shown in FIG. 8 is the same as that of the imaging apparatus 11 described with reference to FIG. 4, except that a recording unit 111 and a processing data recording unit 112 are provided.
[0108] The recording unit 111 is composed of an internal memory or an external storage. Images captured by a camera such as the imaging apparatus 11 are recorded in the recording unit 111.
[0109] The signal processing unit 34 reads an image from the recording unit 111 and performs image processing related to image editing according to the control by the voice command execution unit 58. Operations related to image editing are performed by voice including ambiguous designated words. The image subjected to the image processing by the signal processing unit 34 is output to the image data storage unit 35.
[0110] The image data storage unit 35 temporarily stores the image supplied from the signal processing unit 34. The image data storage unit 35 supplies the image to the processing data recording unit 112 and the display unit 37 according to the operation by the user.
[0111] The processing data recording unit 112 is composed of an internal memory or an external storage. The processing data recording unit 112 records the image supplied from the image data storage unit 35.
[0112] The user can operate the information processing apparatus 101 by voice including natural expressions using ambiguous words such as "more" and "extremely", and cause image editing such as image processing to be performed.
[0113] <5. Regarding the computer> The above-described series of processes can be executed by hardware or by software. When the series of processes are executed by software, the program constituting the software is installed from a program recording medium into a computer in which the program is incorporated in dedicated hardware, or a general-purpose personal computer or the like.
[0114] FIG. 9 is a block diagram showing a configuration example of the hardware of a computer that executes the above-described series of processes by a program.
[0115] A CPU (Central Processing Unit) 301, a ROM (Read Only Memory) 302, and a RAM (Random Access Memory) 303 are interconnected by a bus 304.
[0116] Further connected to the bus 304 is an input / output interface 305. Connected to the input / output interface 305 are an input unit 306 composed of a keyboard, a mouse, etc., and an output unit 307 composed of a display, a speaker, etc. Also connected to the input / output interface 305 are a storage unit 308 composed of a hard disk, a non-volatile memory, etc., a communication unit 309 composed of a network interface, etc., and a drive 310 that drives a removable medium 311.
[0117] In the computer configured as described above, the CPU 301 loads and executes, for example, a program stored in the storage unit 308 into the RAM 303 via the input / output interface 305 and the bus 304, whereby the above-described series of processes are performed.
[0118] The program executed by the CPU 301 is recorded, for example, on the removable medium 311, or is provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting, and is installed in the storage unit 308.
[0119] Note that the program executed by the computer may be a program in which processing is performed in time series in accordance with the order described in this specification, or may be a program in which processing is performed in parallel or at a necessary timing such as when a call is made.
[0120] The effects described in this specification are merely examples and are not limiting, and there may be other effects.
[0121] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the gist of the present technology.
[0122] For example, the present technology can adopt a cloud computing configuration in which one function is shared and jointly processed by a plurality of devices via a network.
[0123] In addition, each step described in the above flowchart can be executed by one device or can be shared and executed by a plurality of devices.
[0124] Furthermore, when a plurality of processes are included in one step, the plurality of processes included in that one step can be executed by one device or can be shared and executed by a plurality of devices.
[0125] <Example of configuration combination> The present technology can also adopt the following configuration.
[0126] (1) When a voice command for instructing control of a device input by a user includes a predetermined word determined to have ambiguous control degree, a command processing unit that executes processing according to the voice command using a parameter corresponding to the way the user speaks when the voice command is input is provided. An information processing apparatus. (2) The command processing unit executes control according to the voice command using the parameter set based on the difference between the way the user speaks when the voice command is input and a reference way of speaking. The information processing apparatus according to (1) above. (3) When the way the user speaks when the voice command is input is different from the reference way of speaking, the command processing unit sets the parameter adjusted to be larger than the reference parameter. The information processing apparatus according to (2) above. (4) The information processing apparatus further includes a determination unit that determines whether the way the user speaks when inputting the voice command is different from a reference way of speaking. The information processing apparatus according to (3) above. (5) Based on a feature amount of the voice including at least any one of the speed, volume, and tone of the voice, the determination unit determines whether the way the user speaks when inputting the voice command is different from a reference way of speaking. The information processing apparatus according to (4) above. (6) Based on the emotion of the user when inputting the voice command, the determination unit determines whether the way the user speaks when inputting the voice command is different from a reference way of speaking. The information processing apparatus according to (4) above. (7) Based on the diction of the user when inputting the voice command, the determination unit determines whether the way the user speaks when inputting the voice command is different from a reference way of speaking. The information processing apparatus according to (4) above. (8) Based on an image obtained by imaging the user when inputting the voice command, the determination unit determines whether the way the user speaks when inputting the voice command is different from a reference way of speaking. The information processing apparatus according to (4) above. (9) Based on sensor data of a wearable sensor worn by the user when inputting the voice command, the determination unit determines whether the way the user speaks when inputting the voice command is different from a reference way of speaking. The information processing apparatus according to (4) above. (10) The voice command is a command related to image processing, The information processing apparatus further includes an image processing unit that performs image processing according to the voice command using the parameter. The information processing apparatus according to any one of (1) to (9) above. (11) The parameter is information representing at least any one of color, frame rate, amount of blur, and brightness. The information processing apparatus according to (10) above. (12) The information processing apparatus further includes an imaging unit that performs imaging, and the image processing unit performs the image processing on the image captured by the imaging unit. The information processing apparatus according to (10) or (11) above. (13) The image processing unit performs the image processing on the image read from a predetermined recording unit. The information processing apparatus according to (10) or (11) above. (14) When the information processing apparatus determines that a predetermined word indicating that the degree of control is ambiguous is included in a voice command for instructing control of a device input by a user, the information processing apparatus uses a parameter corresponding to the way the user speaks when the voice command is input, and executes processing according to the voice command. Information processing method. (15) A computer When a predetermined word indicating that the degree of control is ambiguous is included in a voice command for instructing control of a device input by a user, the computer uses a parameter corresponding to the way the user speaks when the voice command is input, and executes processing according to the voice command, and a program for causing the computer to function as such.
Description of Reference Numerals
[0127] 11 Imaging device, 31 Operation input unit, 32 Voice command input unit, 33 Imaging unit, 34 Signal processing unit, 35 Image data storage unit, 36 Recording unit, 37 Display unit, 51 Voice command input unit, 52 Voice signal processing unit, 53 Voice command recognition unit, 54 Voice command semantic analysis unit, 55 User feature determination unit, 56 User feature storage unit, 57 Parameter value storage unit, 58 Voice command execution unit, 101 Information processing device, 111 Recording unit, 112 Processing data recording unit
Claims
1. An analysis unit that performs semantic analysis of a voice command instructing control of a device input by a user, and determines whether the voice command includes a predetermined word with an ambiguous degree of control; When the predetermined word is included in the voice command, a command processing unit that executes processing according to the voice command using parameters according to the speaking style of the user when the voice command is input An information processing apparatus comprising:
2. The command processing unit executes control according to the voice command using the parameter set based on the difference between the speaking style of the user when the voice command is input and a reference speaking style The information processing apparatus according to claim 1.
3. When the speaking style of the user when the voice command is input is different from the reference speaking style, the command processing unit sets the parameter adjusted to be larger than the reference parameter The information processing apparatus according to claim 2.
4. The information processing apparatus according to claim 3, further comprising a determination unit that determines whether the speaking style of the user when the voice command is input is different from the reference speaking style The information processing apparatus according to claim 3.
5. The determination unit determines whether the speaking style of the user when the voice command is input is different from the reference speaking style based on a feature amount of voice including at least any one of voice speed, volume, and tone The information processing apparatus according to claim 4.
6. The determination unit determines whether the speaking style of the user when the voice command is input is different from the reference speaking style based on the emotion of the user when the voice command is input The information processing apparatus according to claim 4.
7. The determination unit determines whether the speaking style of the user when the voice command is input is different from the reference speaking style based on the diction of the user when the voice command is input The information processing apparatus according to claim 4.
8. The determination unit determines whether the speaking style of the user when the voice command is input is different from the reference speaking style based on an image obtained by imaging the user when the voice command is input The information processing apparatus according to claim 4.
9. The determination unit determines whether the way of speaking of the user when the voice command is input is different from a reference way of speaking based on sensor data of the wearable sensor worn by the user when the voice command is input. The information processing apparatus according to claim 4.
10. The voice command is a command related to image processing, The apparatus further includes an image processing unit that performs image processing according to the voice command using the parameter. The information processing apparatus according to claim 1.
11. The parameter is information representing at least any one of color, frame rate, amount of blur, and brightness. The information processing apparatus according to claim 10.
12. The apparatus further includes an imaging unit that performs imaging, The image processing unit performs the image processing on an image captured by the imaging unit. The information processing apparatus according to claim 10.
13. The image processing unit performs the image processing on an image read from a predetermined recording unit. The information processing apparatus according to claim 10.
14. An information processing apparatus, performs semantic analysis of a voice command for instructing control of a device input by a user, and determines whether the voice command includes a predetermined word with an ambiguous degree of control; when the voice command includes the predetermined word, executes processing according to the voice command using a parameter corresponding to the way of speaking of the user when the voice command is input. An information processing method including the above.
15. A computer, a parsing unit that performs semantic analysis of a voice command for instructing control of a device input by a user, and determines whether the voice command includes a predetermined word with an ambiguous degree of control; a command processing unit that executes processing according to the voice command using a parameter corresponding to the way of speaking of the user when the voice command is input when the voice command includes the predetermined word. A program for causing the computer to function as the above.
Citation Information
Patent Citations
Voice recognition device, voice recognition method and program
JP2014153663A
Voice response system
JP2018136500A
Information processing device, information processing method, and program
WO2019077897A1