On-vehicle device

The in-vehicle device expands the use of input content information by storing, playing back, and designating target portions of voice content, addressing the limited applicability of existing devices and improving functionality and accuracy across various scenarios.

JP2025098698APending Publication Date: 2025-07-02JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023215017
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-07-02

AI Technical Summary

Technical Problem

Existing in-vehicle devices limit the use of input content information to specific scenarios, such as destination input, restricting its applicability in a wide range of contexts.

Method used

An in-vehicle device equipped with a storage processing unit, playback unit, designation unit, and output processing unit to store, playback, and designate target portions of voice content for broader utilization across various scenarios.

Benefits of technology

Enables the wide use of input content information in multiple scenarios beyond destination input, enhancing the device's functionality and accuracy by allowing passengers to specify and utilize target portions of voice content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098698000001_ABST
    Figure 2025098698000001_ABST
Patent Text Reader

Abstract

To widely use information from contents input from information sources in many situations.SOLUTION: An on-vehicle device 1 mounted on a vehicle comprises a memory processing unit 101, a playback unit 102, a designation unit 104, and an output processing unit 105. The memory processing unit 101 stores voice VO contained in the contents of one or more information sources obtained on the vehicle in a memory unit 16. The playback unit 102 plays back a portion of the voice VO stored in the memory unit 16 from a time point of a predetermined input to a time point getting back for a predetermined period in the past based on an input of a predetermined instruction. The designation unit 104 designates a target portion TP of the voice VO played back by the playback unit 102. The output processing unit 105 outputs information of the target portion TP designated by the designation unit 104 to an information processing circuit 4, such as a dialogue agent.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an in-vehicle device.

Background Art

[0002] Patent Document 1 discloses a mobile navigation device. In this device, contents of a plurality of information sources are input in parallel or by switching, and language information is obtained from the input contents by speech recognition or image analysis. From the language information obtained from the contents, language information that matches keywords indicating place names, facility names, and store names is extracted, and the extracted language information is stored in a storage unit as destination candidates. From the destination candidates in the storage unit, candidates that accept user selection are selected, and the selected candidates are displayed on a display unit. The user can select a desired destination from the displayed candidates.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the device of Patent Document 1, keywords are registered in advance as classification data. The language information finally extracted from the contents of the information source is limited to the language information of destination candidates including languages that match the place names, facility names, and store names of the registered keywords. In the device of Patent Document 1, the scene where the information of the contents input from the information source can be used is limited to when the user inputs a destination to the device.

[0005] The present invention has been made in view of the above circumstances, and an object of the present invention is to enable wide use of the information of the contents input from the information source in a wide range of scenes.

Means for Solving the Problems

[0006] In order to achieve the above object, an in-vehicle device according to one aspect of the present invention is an in-vehicle device mounted on a vehicle, and includes a storage processing unit, a playback unit, a designation unit, and an output processing unit. The storage processing unit stores, in a storage unit, the voice included in the content of one or two or more information sources acquired on the vehicle. The playback unit plays back, based on the input of a predetermined instruction, a portion of the voice stored in the storage unit from the time point of the input of the predetermined instruction to a past time point that is a predetermined period back. The designation unit designates a target portion in the voice played back by the playback unit. The output processing unit outputs information on the target portion designated by the designation unit to an information processing circuit.

Effects of the Invention

[0007] According to the present invention, the information of the content input from the information source can be widely used in many scenarios.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The same or equivalent parts or components are denoted by the same reference numerals throughout the drawings.

[0010] The embodiments shown below illustrate devices and the like for embodying the technical idea of this invention. The technical idea of this invention does not specify the arrangement of each component part as follows.

[0011] FIG. 1 is an explanatory diagram showing a basic configuration of an in-vehicle device 1 according to an embodiment of the present invention. The in-vehicle device 1 of the present embodiment shown in FIG. 1 is mounted on a vehicle (not shown). The in-vehicle device 1 can be configured by, for example, a display audio device to which an external device can be connected. In the present embodiment, the case where the in-vehicle device 1 is configured by a display audio device will be described. The in-vehicle device 1 includes a controller 10, a touch panel 11, a speaker 12, a microphone 13, a television tuner 14, a radio tuner 15, and a storage unit 16. The storage unit 16 stores audio included in the content of one or more information sources acquired on the vehicle. The one or more information sources acquired on the vehicle include information sources that can output the audio of the content from the speaker 12. The content that can output audio from the speaker 12 will be described later. The storage unit 16 can be configured by, for example, an eMMC (embedded Multi Media Card), an SSD (Solid State Drive), or an HDD (Hard Disk Drive). The controller 10 will be described later.

[0012] The touch panel 11 is used as a user interface, accepts input of information from the user, or outputs information to the user. The touch panel 11 can be configured by combining a display unit 111 such as a liquid crystal panel and a position input unit 112. In the display unit 111, images of various types of information provided to the user can be displayed. For the position input unit 112, for example, a touch pad can be used. The touch pad may also be referred to as a touch screen, touch panel, or contact screen, etc. The display audio device can perform operations corresponding to various functions according to touch operations on the touch pad, etc. The speaker 12 can output voices of various types of information provided to the user. The microphone 13 can collect voices uttered by the user, for example, for input in place of the touch operation of the touch pad of the position input unit 112.

[0013] The TV tuner 14 uses, for example, TV broadcasts as an information source, receives and demodulates the radio waves of TV broadcasts. The in-vehicle device 1 can output the images and voices of TV broadcasts of the radio waves received and demodulated by the TV tuner 14 from the display unit 111 and the speaker 12 respectively as the content of the information source. The radio tuner 15 uses, for example, radio broadcasts as an information source, receives and demodulates the radio waves of radio broadcasts. The in-vehicle device 1 can output the voices of radio broadcasts of the radio waves received and demodulated by the radio tuner 15 from the speaker 12 as the content of the information source.

[0014] The in-vehicle device 1 of this embodiment has a car audio function. The in-vehicle device 1 may further have an output function for car navigation information. The car navigation information can include, for example, the position information of the vehicle equipped with the in-vehicle device 1, information such as the route to the destination of the vehicle, etc. The car navigation information can be provided to the passengers of the vehicle, who are the users, by image and voice using, for example, the display unit 111 and the speaker 12. The car navigation information provided to the passengers by the output function of the car navigation information may be selected from a plurality of information based on, for example, a touch operation of the position input unit 112 or the voice of the passenger collected by the microphone 13.

[0015] The in-vehicle device 1 may itself have a configuration for acquiring car navigation information, or may be configured to be able to externally attach an external device (not shown) for acquiring car navigation information. This external device may be, for example, a portable terminal such as a smartphone (smartphone) or a tablet possessed by the passengers of the vehicle. The portable terminal can acquire car navigation information, for example, by starting the car navigation application software installed on the portable terminal. The in-vehicle device 1 can be connected to the portable terminal by wireless communication such as Bluetooth (registered trademark), Wi-Fi, etc. The in-vehicle device 1 connected to the portable terminal can output the image of the car navigation information acquired by the portable terminal on the display unit 111 and output the voice of the car navigation information acquired by the portable terminal from the speaker 12. The configuration for acquiring car navigation information corresponds to the information source acquired on the vehicle. The car navigation information by voice output from the speaker 12 provided to the passengers by the output function of the car navigation information corresponds to the content of the information source.

[0016] The car audio function of in-vehicle device 1 can be realized, for example, by using a TV tuner 14 and a radio tuner 15 built in the in-vehicle device 1. The car audio function of in-vehicle device 1 can also be realized, for example, by using an external device connected to the in-vehicle device 1. The external device includes an external device 2 and an external media player 3. The external device 2 and the external media player 3 can be connected to the in-vehicle device 1 by wire, for example, via a USB (Universal Serial Bus) interface. The external device 2 and the external media player 3 may be connected to the in-vehicle device 1 by wireless communication such as Bluetooth (registered trademark) or in-vehicle wireless LAN (Local Area Network).

[0017] The external device 2 may be, for example, a device having a built-in memory (not shown) that stores files including audio. The external device 2 having a built-in memory that stores files including audio can play back the files read from the memory independently in the external device 2. The external device 2 capable of playing back the files in the memory can be, for example, a digital media player. The digital media player may be an AV (audio-visual) player that plays back video files including images and audio, or an audio player that plays back music files by audio. The external device 2 capable of playing back the files in the memory may be, for example, a portable terminal held by a vehicle occupant that can play back video files or music files. The in-vehicle device 1 connected to the external device 2 capable of playing back the files in the memory can output the image of the video file played back by the external device 2 on the display unit 111 and output the audio of the video file or music file played back by the external device 2 from the speaker 12.

[0018] When the external device 2 is a smartphone that can be used as a mobile phone, a hands-free call with the other party of the call destination can be made by wirelessly connecting the smartphone to the in-vehicle device 1. The voice of the passenger who spoke during the hands-free call can be collected by the microphone 13 and input from the in-vehicle device 1 to the smartphone. The voice of the other party of the call destination who spoke during the hands-free call can be input from the smartphone to the in-vehicle device 1 and output from the speaker 12. During the hands-free call, the call corresponds to the information source, and the call content corresponds to the content. The call content includes at least the voice of the other party of the call destination output from the speaker 12. The call content may also include the voice of the passenger collected by the microphone 13.

[0019] The external media player 3 is a device that reads a file including voice from a media (not shown). The media (not shown) may be, for example, an optical disc such as a compact disc or a DVD (Digital Versatile Disc), a memory card, or a USB flash memory. The voice of the voice file read by the external media player 3 from the media (not shown) can be input from the external media player 3 to the in-vehicle device 1 and output from the speaker 12. The voice file output from the speaker 12 corresponds to the content, and the media from which the external media player 3 reads the voice file corresponds to the information source.

[0020] The controller 10 of the in-vehicle device 1 can be configured using, for example, a general-purpose microcontroller (not shown). The microcontroller includes a CPU (Central Processing Unit) and a memory. The memory includes a ROM (Read Only Memory) and a RAM (Random Access Memory). The microcontroller can virtually construct a plurality of information processing circuits by the CPU executing a program stored in the memory. When configuring the controller 10 using a microcontroller, the plurality of information processing circuits constructed by the microcontroller inside the controller 10 can be used to configure each of the units 101 to 105 described later.

[0021] In the present embodiment, an example is shown in which a plurality of information processing circuits constructed in the controller 10 are realized by software. Of course, it is also possible to prepare dedicated hardware for executing each of the information processes described below to configure the information processing circuit. Further, the plurality of information processing circuits may be configured by individual hardware. A touch panel 11, a speaker 12, a microphone 13, a television tuner 14, a radio tuner 15, and a storage unit 16 are connected to the controller 10.

[0022] The memory processing unit 101 stores the voice included in the content of one or more information sources acquired on the vehicle in the memory unit 16. The voice to be stored in the memory unit 16 can be, for example, the content of each information source of the above-described TV tuner 14, radio tuner 15, external device 2, and external media player 3 that can output voice from the speaker 12. The content that can output voice from the speaker 12 includes TV broadcasts, radio broadcasts, and car navigation information by voice. The content that can output voice from the speaker 12 includes video files and music files played by the external device 2, call content by smartphone, and voice files read from the media by the external media player 3. The content for storing voice in the memory unit 16 may be limited to the content that has actually output voice from the speaker 12, or may be all content that can output voice from the speaker 12 if selected by the occupant even if voice is not actually output from the speaker 12. When the memory processing unit 101 stores the voice of all content in the memory unit 16, the voice can be stored separately for each content.

[0023] The playback unit 102 plays back the voice stored in the storage unit 16 based on a predetermined instruction input to the controller 10. The playback unit 102 plays back a portion of the voice stored in the storage unit 16 for a predetermined period. The predetermined period is a period from the time point when the predetermined instruction is input to the controller 10 back to a past time point for the predetermined period. The predetermined instruction may be given, for example, by the utterance of a predetermined phrase by a vehicle occupant. The predetermined phrase may be, for example, the line "Play it again". The predetermined instruction by voice of the utterance can be collected by, for example, the microphone 13, converted into text data by voice recognition, and input to the controller 10. For voice recognition, for example, a voice recognition circuit (not shown) mounted on the in-vehicle device 1 can be used. As the voice recognition circuit, a known circuit that converts voice data into text data using a voice recognition model can be used. The predetermined instruction may be given, for example, by a touch operation on the touch panel 11 by a vehicle occupant. The predetermined instruction by the touch operation can be detected by the position input unit 112 and input to the controller 10. The predetermined instruction by the vehicle occupant may be detected by operating a switch or the like other than the touch panel 11 and input to the controller 10.

[0024] As shown in FIG. 2, for example, the playback unit 102 may output the voice VO of the storage unit 16 played back based on a predetermined instruction of the occupant as it is from the speaker 12 of the in-vehicle device 1. The voice VO of the storage unit 16 played back by the playback unit 102 may be converted into a character string by voice recognition by the display processing unit 103 in FIG. 1 and displayed on the display unit 111. For voice recognition of the played-back voice VO, for example, the above-described voice recognition circuit can be used.

[0025] The specifying unit 104 specifies the words that are the target part in the voice VO of the storage unit 16 reproduced by the reproducing unit 102. The words to be specified may be a single word or a clause formed by connecting a plurality of words such as a natural sentence. The target part specified by the specifying unit 104 can be used as an input to the information processing circuit 4 connected to the in-vehicle device 1. The information processing circuit 4 may be, for example, an interactive agent or a circuit constituting an output function of car navigation information.

[0026] For example, the specifying unit 104 can detect an operation by a vehicle occupant to specify a target part and, based on the detected operation, specify the target part of the voice VO to be input to the information processing circuit 4. As shown in FIG. 3, the specification of the target part TP can be performed, for example, by the speech of the occupant instructing the start point SP and the end point EP of the target part TP during the output of the voice VO from the speaker 12. The specifying unit 104 can identify the start point SP and the end point EP of the target part TP in the voice VO and specify the target part TP based on the timings of the speech of the start point SP and the end point EP by the occupant collected by the microphone 13 during the reproduction of the voice VO in the storage unit 16.

[0027] As shown in FIG. 4, the specification of the target part TP may be performed, for example, by a touch operation of the occupant instructing the start point SP and the end point EP of the target part TP. FIG. 4 shows a case where an occupant who has visually recognized the character string CH of the voice VO displayed on the display unit 111 by the display processing unit 103 indicates the target part TP by a touch operation on the display unit 111 tracing the target part TP with a finger in the character string CH. The touch operation on the display unit 111 can be detected by the position input unit 112. The specifying unit 104 can specify the target part TP based on the start point SP and the end point EP of the touch operation detected by the position input unit 112. As shown in FIG. 5, the display processing unit 103 may, for example, detect a touch operation from the start point SP to the end point EP by the position input unit 112 and cause the target part TP specified by the specifying unit 104 to be reversely displayed TPi on the display unit 111.

[0028] The output processing unit 105 outputs the information of the target part TP specified by the specifying unit 104 to the information processing circuit 4. The information of the target part TP may be, for example, the voice VO of the target part TP specified by the specifying unit 104, or the text data of the target part TP converted from the voice VO by voice recognition.

[0029] Hereinafter, an example of the procedure of the process executed by the controller 10 of the in-vehicle device 1 of the present embodiment will be described with reference to the flowchart of FIG. 6.

[0030] The controller 10 checks whether the content including the voice VO has been input from the information source (step S1). If the content including the voice VO has not been input (NO in step S1), step S1 is repeated. If the content including the voice VO has been input (YES in step S1), the storage processing unit 101 starts recording the voice VO of the input content (step S3). The controller 10 checks whether the reproduction of the voice VO stored in the storage unit 16 has been instructed by the input of a predetermined instruction (step S5).

[0031] If the reproduction of the voice VO has not been instructed (NO in step S5), the process proceeds to step S13 described later. If the reproduction of the voice VO has been instructed (YES in step S5), the reproduction unit 102 reproduces a predetermined period portion of the voice VO stored in the storage unit 16 (step S7). The voice VO reproduced by the reproduction unit 102 may be output as it is from the speaker 12 by the controller 10, or may be converted into a character string CH by the display processing unit 103 and displayed on the display unit 111.

[0032] The controller 10 acquires an input for specifying the target part TP by the specifying unit 104 (step S9). The output processing unit 105 outputs the voice VO or the text data converted from the voice as the information of the target part TP specified by the specifying unit 104 to the information processing circuit 4 (step S11). Then, the process proceeds to step S13.

[0033] In step S13, the controller 10 checks whether the input of the content including the voice VO continues. If the input of the content including the voice VO continues (YES in step S13), it returns to step S5. If the content including the voice VO is no longer input (NO in step S13), the memory processing unit 101 ends the recording of the voice VO of the content (step S15) and ends the series of processes.

[0034] In the in-vehicle device 1 of the present embodiment, when the content including the voice VO is input from the information source to the controller 10, the voice VO included in the content of the information source is stored in the storage unit 16. When the reproduction of the voice VO is instructed by the input of a predetermined instruction by the speech or operation of the occupant, a portion of the voice VO stored in the storage unit 16 for a predetermined period is reproduced. When the target portion TP in the reproduced voice VO is specified by the speech or operation of the occupant, the voice VO or text data of the target portion TP is output to the information processing circuit 4 as the information of the specified target portion TP.

[0035] In the in-vehicle device 1 of the present embodiment, for example, when a passenger gives an instruction to play the voice VO, a portion of the voice VO recorded and stored in the storage unit 16 from the time point of the play instruction to a past time point retroactively for a predetermined period is played. The passenger can give an instruction to play the voice VO, for example, when the passenger notices that there is a word of concern in the voice VO output from the speaker 12 and being listened to. The passenger designates, as the target portion TP of the word of concern, the words of an arbitrary portion from among the played voice VOs. When designating the target portion TP, the passenger designates from the start point SP to the end point EP of the target portion TP. The target portion TP can be designated not only as a word but also as a clause. In the in-vehicle device 1 of the present embodiment, after converting the voice VO into text data by voice recognition, the degree of freedom of the words that can be designated as the target portion TP can be improved compared to the target portions in word units extracted by logic from the character string CH. According to the present embodiment, the information on the target portion TP obtained from the content input from the information source can be widely used in many scenarios, not only in scenarios such as extracting words of places, facilities, etc. that are candidates for the destination of the vehicle.

[0036] Since the designation unit 104 designates the target portion TP using the voice VO played from the storage unit 16, it is possible to provide more accurate information to the information processing circuit 4 than when the passenger repeats the word of concern relying on their own memory and inputs it to the information processing circuit 4.

[0037] Note that the memory processing unit 101 may sequentially erase the old voice VOs in the memory unit 16 by overwriting them with, for example, the voice VOs of the content newly acquired by the in-vehicle device 1 from the information source. The old voice VOs sequentially erased from the memory unit 16 may be, for example, the voices after a predetermined period has passed since the memory processing unit 101 stored them in the memory unit 16. Even if the memory processing unit 101 sequentially erases the old voice VOs in the memory unit 16, the playback unit 102 can playback the voice VOs in the past for a predetermined period from the memory unit 16 from the playback instruction. The memory unit 16 may continue to store the voice VOs acquired by the in-vehicle device 1 from the information source within the range of the storage capacity of the memory unit 16 even after a predetermined period has elapsed. Even after a certain amount of time has passed since the output from the speaker 12, the passenger can specify, for example, the words of concern as the target part TP from the voice VOs for a predetermined period in the memory unit 16 that can be played back by the playback unit 102.

[0038] In this embodiment, it is assumed that the specifying unit 104 specifies the start point SP and the end point EP of the target part TP by an operation of specifying the target part TP of the passenger. However, for example, only the end point EP of the target part TP may be specified, and the start point SP of the target part TP may be automatically specified at the position of the word break that can be determined logically in the voice VO at a time point before the end point EP in the time series. In this case, the automatically specified start point SP can be, for example, the point where at least one noun is included between the start point SP and the end point EP.

[0039] In addition, the playback unit 102 may play back, as the voice VO in the storage unit 16, the synthesized voice generated by voice synthesis processing based on, for example, the text data converted by voice recognition. When the playback unit 102 plays back the voice VO in the storage unit 16 with the synthesized voice, the in-vehicle device 1 may include a voice synthesis circuit that generates the synthesized voice of the voice VO played back by the playback unit 102. When the playback unit 102 plays back the synthesized voice of the voice VO, the information of the target portion TP input by the output processing unit 105 to the information processing circuit 4 remains the content of the synthesized voice. When the occupant designates the target portion TP from the voice VO of the synthesized voice played back by the playback unit 102, by confirming the content of the voice VO of the synthesized voice, it is possible to confirm in advance that the information of the target portion TP is input to the information processing circuit 4 as the content intended by the occupant.

[0040] The in-vehicle device 1 is not limited to a display audio device, and may have a function of acquiring content including voice from an information source regardless of whether it has a car navigation function itself, and may be mounted on a vehicle.

Explanation of Signs

[0041] 1 In-vehicle device 2 External device 3 External media player 4 Information processing circuit 10 Controller 14 TV tuner 15 Radio tuner 16 Storage unit 101 Storage processing unit 102 Playback unit 103 Display processing unit 104 Designation unit 105 Output processing unit 111 Display unit CH Character string EP End point SP Start point TP Target portion VO Voice

Claims

1. An in-vehicle device mounted on a vehicle, comprising: a storage processing unit that stores in a storage unit voice included in content of one or more information sources acquired on the vehicle; a playback unit that, based on an input of a predetermined instruction, plays back a portion of the voice stored in the storage unit from a past time point up to a predetermined period back from the time point of input of the predetermined instruction; a specifying unit that specifies a target portion in the voice played back by the playback unit; an output processing unit that outputs information on the target portion specified by the specifying unit to an information processing circuit; The in-vehicle device comprising the above components.

2. The in-vehicle device according to claim 1, wherein the specifying unit detects an operation by an occupant of the vehicle who has listened to the voice played back by the playback unit to specify the target portion, and specifies the target portion.

3. The in-vehicle device according to claim 1, further comprising a display processing unit that causes a string obtained by voice recognition of the voice played back by the playback unit to be displayed on a display unit, wherein the specifying unit detects an operation by an occupant of the vehicle who has visually recognized the string displayed on the display unit to specify the target portion, and specifies the target portion.

4. The in-vehicle device according to any one of claims 1 to 3, wherein the specifying unit specifies at least a start point and an end point of the target portion.

5. The in-vehicle device according to any one of claims 1 to 3, wherein the playback unit plays back the voice stored in the storage unit by a synthesized voice obtained by voice synthesis processing.

Citation Information

Patent Citations

  • Moving-body navigation device

    WO2016185540A1