Voice image display method and system, electronic device, and vehicle
By using a sound source localization algorithm to select a suitable voice image display device (screen or starry sky ceiling), the problem that passengers not in the first row cannot see the voice image is solved, improving the voice interaction experience for users in all positions in the vehicle.
Patent Information
- Application Number
- CN202310484826.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-04-28
AI Technical Summary
When using voice interaction in the car, passengers who are not in the first row cannot accurately and conveniently see the voice feedback on the central control screen, resulting in a poor user experience.
The target vehicle's location is determined using a sound source localization algorithm. A suitable voice image display device (screen display or starry sky ceiling) is selected, and the target voice image is displayed on that device, including on the screen display and in the starry sky ceiling area.
It improves the voice interaction effect for users in different positions inside the vehicle, enhancing the user experience.
Smart Images

Figure CN116339670B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent control of automobiles, and in particular to a voice image display method and system, an electronic device, and a vehicle. BACKGROUND
[0002] There are various voice images on the vehicle side, which can interact with users through voice images combined with various voice service scenarios, so that the voice interaction process in the vehicle is more interesting.
[0003] At present, the voice image display in the vehicle is mainly realized through the center control screen. However, if the passengers who perform voice interaction are not the first-row passengers, these passengers cannot accurately and conveniently see the feedback voice image through the center control screen to obtain the voice interaction feedback result, which will cause voice interaction obstacles and seriously affect the user experience. SUMMARY
[0004] Therefore, the present application aims to provide a voice image display method and system, an electronic device, and a vehicle to realize feedback of voice images using different voice image display devices according to the positions of users who issue instructions, improve the voice interaction effect of users at different positions, and further improve the user experience.
[0005] To achieve the above purpose, the present application provides a voice image display method, which comprises the following steps:
[0006] receiving target voice information, determining a target in-vehicle position corresponding to the target voice information and a target voice image corresponding to the target voice information according to the target voice information;
[0007] determining a voice image display device according to the target in-vehicle position; wherein the voice image display device comprises a display screen or a starry sky roof;
[0008] if the voice image display device is a display screen, controlling the display screen to display the target voice image;
[0009] if the voice image display device is a starry sky roof, determining a target starry sky roof area corresponding to the target in-vehicle position according to the target in-vehicle position, and controlling the starry sky roof to display the target voice image in the target starry sky roof area.
[0010] To achieve the above purpose, the present application further provides a voice image display system, which comprises a vehicle machine, a display screen, and a starry sky roof, and the starry sky roof comprises a starry sky roof control module and a starry sky roof display module, wherein:
[0011] The car machine is configured to receive target voice information, determine a target in-vehicle position corresponding to the target voice information and a target voice image corresponding to the target voice information according to the target voice information, determine a voice image display device according to the target in-vehicle position, control the display screen to display the target voice image if the voice image display device is a display screen, and control the starry sky roof to display the target voice image in a target starry sky roof area corresponding to the target in-vehicle position if the voice image display device is a starry sky roof.
[0012] The display screen is configured to display the target voice image.
[0013] The starry sky roof control module is configured to control the starry sky roof to display the target voice image in the target starry sky roof area of the starry sky roof display module.
[0014] To achieve the above object, the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and is characterized in that the processor implements the voice image display method provided by any one of the embodiments of the present application when executing the program.
[0015] To achieve the above object, the present application further provides a vehicle comprising the voice image display system provided by any one of the embodiments of the present application or the electronic device provided by any one of the embodiments of the present application.
[0016] As can be seen from the above, the voice image display method provided by the present application determines a target in-vehicle position by sound source positioning on the received target voice information, so as to determine a voice image display device, judges whether to use a display screen or a starry sky roof to feed back the target voice information, so as to improve the adaptability of the voice image display effect to users, facilitate users to obtain feedback information, and identify and process the target voice information to obtain a target voice image, and then display the target voice image on the display screen if the voice image display device is a display screen, and analyze a target starry sky roof area corresponding to the target in-vehicle position if the voice image display device is a starry sky roof, so as to display the target voice image in the target starry sky roof area, so that users can more conveniently obtain feedback of the target voice information, the voice image display device is used to feed back the voice image according to the position of the user who issues the instruction, the voice interaction effect of users at different positions is improved, and the user experience is improved. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the application or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related art description. Obviously, the drawings in the following description only constitute the embodiments of the application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0018] Figure 1 A flow chart of a voice image display method provided by an embodiment of the application;
[0019] Figure 2 A flow chart of another voice image display method provided by an embodiment of the application;
[0020] Figure 3 A structural schematic diagram of a voice image display system provided by an embodiment of the application;
[0021] Figure 4 A structural schematic diagram of another voice image display system provided by an embodiment of the application;
[0022] Figure 5 A hardware structural schematic diagram of an electronic device provided by an embodiment of the application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the application more clear, the following will further describe the application in detail with specific embodiments and with reference to the drawings.
[0024] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the application should be understood as the common meanings understood by those skilled in the art to which the embodiments of the application belong. The terms "first", "second" and similar terms used in the embodiments of the application do not represent any order, number or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar terms mean that the elements or objects before the terms cover the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms do not mean physical or mechanical connection, but can include electrical connection, whether direct or indirect. The terms "upper", "lower", "left", "right" and the like only represent relative positional relationships, and when the absolute positions of the described objects change, the relative positional relationships may also change accordingly.
[0025] Figure 1A flowchart of a voice image display method provided for an embodiment of the present application is shown in the figure. The method is set in a car machine and mainly applicable to controlling different voice image display devices to display voice images corresponding to voice instructions according to the positions of users sending the voice instructions. The method can be executed by a voice image display system which can be configured in an electronic device. As shown in the figure, the method can specifically include the following steps: Figure 1
[0026] S110, receiving target voice information, determining a target in-vehicle position corresponding to the target voice information and a target voice image corresponding to the target voice information according to the target voice information.
[0027] The target voice information is information spoken by a user in the vehicle during voice interaction, which can be captured by a sound collecting device in the vehicle. The target in-vehicle position is the position of the user sending the target voice information, which can be a specific position or a position range, for example, the left side of the second row, the middle of the third row, etc. The target voice image is a voice image fed back after responding to the target voice information and corresponding processing.
[0028] Specifically, the sound collecting device in the vehicle captures the target voice information sent by the user in the vehicle and forwards the target voice information to the car machine. After receiving the target voice information, the car machine analyzes the target voice information, determines the target in-vehicle position sending the target voice information, and determines the target instruction corresponding to the target voice information. Further, the target instruction is executed, and according to the execution result of the target instruction, a voice image corresponding to the execution result is determined as the target voice image.
[0029] It can be understood that the voice image corresponding to various execution results can be determined according to the pre-configured corresponding relationship between the execution results and the voice images.
[0030] S120, determining a voice image display device according to the target in-vehicle position.
[0031] The voice image display device is a device for displaying the target voice image. The voice image display device includes a display screen or a starry sky roof. The display screen can be a display screen of a center console, etc. The starry sky roof is set on the top of the vehicle and is composed of multiple lamp holders, controllers, etc., and supports multiple color settings. At present, the starry sky roof can display some meteor and constellation effects, so that the user has a better visual experience.
[0032] Specifically, according to the target in-vehicle position, it can be further judged that the display content of which display device is in the field of view of the user corresponding to the target in-vehicle position and can be viewed by the user corresponding to the target in-vehicle position, and then the determined display device is taken as the voice image display device for subsequent display of the target voice image to feed back the target voice information.
[0033] On the basis of the above examples, the voice image display device can be determined according to the target in-vehicle position in the following manner:
[0034] If the target in-vehicle position corresponds to the first-row position, the voice image display device is determined to be the display screen.
[0035] If the target in-vehicle position corresponds to a non-first-row position, the voice image display device is determined to be the star ceiling in the case where the opening state of the star ceiling is open, and the voice image display device is determined to be the display screen in the case where the opening state of the star ceiling is closed.
[0036] The first-row position is a row of positions arranged close to the front windshield in the vehicle, and the non-first-row position is at least one row of positions arranged away from the front windshield in the vehicle. The first-row position can be understood as a row of positions where the driver and the front passenger are located, and the non-first-row position can be at least one row of positions behind the first-row position in the vehicle, such as the second-row position, the third-row position, etc. The opening state is used to describe the switch of the star ceiling to determine whether the star ceiling can be used to display the target voice image.
[0037] Specifically, if the target in-vehicle position belongs to the first-row position, it indicates that the user located at the target in-vehicle position can see the display screen, and therefore the display screen can be taken as the voice image display device to display the target voice image required to be displayed subsequently through the display screen. If the target in-vehicle position belongs to the non-first-row position, it indicates that the user located at the target in-vehicle position cannot directly see the display screen, but the user located at the target in-vehicle position can see the star ceiling by looking up, and therefore the star ceiling can be taken as the voice image display device to display the target voice image required to be displayed subsequently through the star ceiling. In addition, the car machine can receive the opening state of the star ceiling through the display screen, and if the opening state of the star ceiling is open, it indicates that the star ceiling is available, and therefore the voice image display device is determined to be the star ceiling; if the opening state of the star ceiling is closed, it indicates that the star ceiling is unavailable, and therefore the voice image display device is determined to be the display screen.
[0038] S130, if the voice image display device is the display screen, the display screen is controlled to display the target voice image.
[0039] Specifically, if the voice image display device is a display screen, the vehicle machine can send the continuous animation video frames corresponding to the target voice image to the display screen to display the target voice image to the user corresponding to the target in-vehicle location through the display screen, and feed back the target voice information.
[0040] In S140, if the voice image display device is a starry sky roof, a target starry sky roof area corresponding to the target in-vehicle location is determined according to the target in-vehicle location, and the starry sky roof is controlled to display the target voice image in the target starry sky roof area.
[0041] The target starry sky roof area is a starry sky roof area within the field of view of the user of the target in-vehicle location.
[0042] Specifically, since the starry sky roof covers a large area of the roof, if the entire starry sky roof is used to display the target voice image, the display content will exceed the field of view of the user, and may also affect the line of sight of other users in the vehicle. If the target voice image is displayed in a fixed area of the starry sky roof, the fixed area cannot be adaptively matched with the target in-vehicle location, and the intelligence is poor. Therefore, according to the target in-vehicle location, a region corresponding to the target in-vehicle location is determined as the target starry sky roof area. The vehicle machine can convert the continuous animation video frames corresponding to the target voice image into continuous starry sky roof control signals, and send the starry sky roof control signals to the starry sky roof, and also send the target starry sky roof area to the starry sky roof, so that the starry sky roof displays the target voice image in the target starry sky roof area according to the starry sky roof control signals.
[0043] The voice image display method provided in this embodiment determines the target in-vehicle location by performing sound source positioning on the received target voice information, so as to determine the voice image display device, judge whether to use a display screen or a starry sky roof to feed back the target voice information, improve the adaptability of the voice image display effect to the user, facilitate the user to obtain feedback information, and perform identification and processing on the target voice information to obtain the target voice image. Then, in the case where the voice image display device is a display screen, the target voice image is displayed through the display screen, and in the case where the voice image display device is a starry sky roof, a target starry sky roof area corresponding to the target in-vehicle location is analyzed to display the target voice image in the target starry sky roof area, so that the user can more conveniently obtain the feedback of the target voice information. The voice image display device is used to feed back the voice image according to the position of the user who issues the instruction, the voice interaction effect of each position user is improved, and the user experience is improved.
[0044] Figure 2A flowchart of another voice image display method provided by the embodiments of the present application is shown in FIG. 10. Based on the above embodiments, the manner of determining the target in-vehicle position, the target voice image, and the target starry sky top area is exemplarily illustrated. The explanations of the same or corresponding terms as those in the above embodiments are not repeated here. As shown in FIG. 10, the method can specifically include the following steps. Figure 2
[0045] S210, receiving target voice information.
[0046] S220, determining a target in-vehicle position corresponding to the target voice information based on a preset sound source positioning algorithm.
[0047] The preset sound source positioning algorithm can be a time difference of arrival algorithm or other algorithm for obtaining a sound source position.
[0048] Specifically, when the target voice information is received, the target voice information can be processed based on the preset sound source positioning algorithm to calculate the position of the target voice information, i.e., the target in-vehicle position corresponding to the target voice information.
[0049] Based on the above examples, the target in-vehicle position corresponding to the target voice information can be determined based on the preset sound source positioning algorithm by the following method:
[0050] The target in-vehicle position is determined according to the sound receiving time of each sound receiving device receiving the target voice information and the installation position of each sound receiving device.
[0051] The sound receiving devices are a plurality of microphones installed in the vehicle, which can be in the form of a microphone array. The sound receiving time can be the time when each sound receiving device receives the target voice information, and the installation position can be the position of each sound receiving device in the vehicle.
[0052] Specifically, when each sound receiving device receives the target voice information, the target voice information and the sound receiving time can be sent to the vehicle machine together. The vehicle machine can determine the difference between each two sound receiving times according to the received sound receiving times of each sound receiving device, and determine the sound source position, i.e., the target in-vehicle position, according to the calculated difference and the pre-recorded installation position of each sound receiving device.
[0053] S230, determining a target operation corresponding to the target voice information based on a preset voice recognition algorithm, executing the target operation, and determining a target voice image corresponding to the operation result based on the operation result of the target operation.
[0054] The preset voice recognition algorithm can be a recognition algorithm for converting voice information into instruction information, including voice understanding, conversion and the like. The target operation is an operation to be performed in response to the instruction corresponding to the target voice information. The operation result can include operation success or operation failure, and if the operation is successful, it can also include relevant information obtained by executing the target operation.
[0055] Specifically, the target voice information is subjected to voice recognition processing based on the preset voice recognition algorithm to determine the instruction carried by the target voice information. After obtaining the instruction, the car machine determines and executes the target operation corresponding to the instruction, and after the execution of the target operation ends, the operation result of the target operation can be obtained. Further, the target voice image corresponding to the operation result is obtained for subsequent feedback of the operation result.
[0056] For example, if the operation result of the target operation is wake-up success, the target voice image can be a virtual character appearing; if the operation result of the target operation is close success, the target voice image can be a virtual character waving goodbye; if the operation result of the target operation is that the weather is raining, the target voice image can be a virtual character with raindrops falling around, etc.
[0057] On the basis of the above examples, in addition to using the voice image for feedback after executing the target operation, the voice playing module as the end of the link in the voice recognition process can report and feed back the operation result of the execution, state judgment, interaction, etc. of the target operation, which can be specifically:
[0058] According to the operation result of the target operation, determine the to-be-reported information, and control the voice playing module to report the to-be-reported information.
[0059] The to-be-reported information can be a voice signal form of the operation result of the target operation. The voice playing module can be a loudspeaker, a sound box or the like, which is a device for recognizing and playing voice signals.
[0060] Specifically, according to the operation result of the target operation, the operation result can be converted into a voice signal that can be recognized by the voice playing module, i.e. the to-be-reported information, and the to-be-reported information is sent to the voice playing module, so that the voice playing module plays the to-be-reported information, and the operation result of the target voice information is fed back in the form of voice playing.
[0061] S240, determining a voice image display device according to the target in-vehicle position.
[0062] S250, if the voice image display device is a display screen, controlling the display screen to display the target voice image.
[0063] S260, if the voice image display device is a starry sky roof, determining a starry sky roof projection position corresponding to the target in-vehicle position according to the target in-vehicle position, determining a target starry sky roof region according to the starry sky roof projection position and a preset display area, and controlling the starry sky roof to display the target voice image in the target starry sky roof region.
[0064] The starry sky roof projection position is a perpendicular projection position of the target in-vehicle position on the starry sky roof in the vehicle, which can be understood as the intersection position of the target in-vehicle position in the direction directly above the starry sky roof. The preset display area is the area of the region on the starry sky roof for displaying the target voice image, which can be an area determined according to the field of view of the user in the vehicle to the starry sky roof.
[0065] Specifically, if the voice image display device is a starry sky roof, a perpendicular projection of the target in-vehicle position to the starry sky roof can be performed, and the projection result is the starry sky roof projection position. Then, the position of the target starry sky roof region can be located according to the position of the preset starry sky roof projection position in the target starry sky roof region, and the target starry sky roof region can be expanded from the starry sky roof projection position according to the preset display area, so as to facilitate subsequent display of the target voice image in the target starry sky roof region.
[0066] It can be understood that the reason for first determining the position of the preset starry sky roof projection position in the target starry sky roof region is that the starry sky roof projection position can not be the center of the target starry sky roof region. Since the user's line of sight is forward, the starry sky roof projection position can be set as the edge position of the target starry sky roof region in the vehicle tail direction, so that the user can see the target voice image by slightly looking up.
[0067] On the basis of the above example, the starry sky roof can be controlled to display the target voice image in the target starry sky roof region in the following manner:
[0068] determining a video frame sequence corresponding to the target voice image, performing image processing and conversion operations on the video frame sequence to obtain a starry sky roof control signal;
[0069] sending the starry sky roof control signal and the target starry sky roof region to the starry sky roof, so that the starry sky roof displays the target voice image in the target starry sky roof region.
[0070] The video frame sequence can be a continuous video frame of an animation corresponding to the target voice image. The starry sky top control signal is a signal that can be recognized by the starry sky top controller and forwarded to the starry sky top display. The image processing and conversion operations include sorting processing operations, contour extraction processing operations, color extraction processing operations, and control signal conversion operations. The sorting processing operation is an operation of sorting each video frame in the video frame sequence according to the time sequence. The contour extraction processing operation is an operation of performing contour extraction on each video frame in the video frame sequence. The color extraction processing operation is an operation of performing color extraction on each video frame in the video frame sequence. The control signal conversion operation is to convert each information of the image processing into a control signal that can be recognized by the starry sky top. The starry sky top control signal is used to control the starry sky top to light up as required.
[0071] Specifically, after determining the target voice image, the video frame sequence corresponding to the target voice image can be retrieved from the database of the car machine. The video frame sequence is subjected to image processing and conversion operations to sort, contour extract, and color extract each video frame, and finally convert the video frame sequence into a control signal that can be recognized by the starry sky top, i.e., a starry sky top control signal. Then, the starry sky top control signal and the target starry sky top area are sent to the starry sky top. The starry sky top control module receives the starry sky top control signal and transmits it to the starry sky top display module to display the target voice image in the target starry sky top area on the starry sky top display module.
[0072] The voice image display method provided in this embodiment determines a target in-vehicle position corresponding to the target voice information based on a preset sound source positioning algorithm, improves the accuracy of the position judgment of the user, determines a target operation corresponding to the target voice information based on a preset voice recognition algorithm, executes the target operation, and determines a target voice image corresponding to the operation result based on the operation result of the target operation, to more accurately respond to the target voice information and obtain the target voice image that needs to be fed back subsequently. In addition, a starry sky top projection position corresponding to the target in-vehicle position is determined according to the target in-vehicle position, and a target starry sky top area is determined according to the starry sky top projection position and the preset display area, so that the target voice image displayed by the starry sky top is more easily received by the user located at the target in-vehicle position, the voice interaction effect is improved, and the adaptability of the voice image display to the user is enhanced.
[0073] It should be noted that the method of the embodiments of the present application can be executed by a single device, such as a computer or a server. The method of the embodiments of the present application can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present application, and the multiple devices can interact with each other to complete the method.
[0074] It is to be noted that some embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.
[0075] Based on the same inventive concept, the present application also provides a voice image display system corresponding to any of the above-mentioned embodiment methods. Figure 3 A structural schematic diagram of a voice image display system provided by an embodiment of the present application is shown in Figure 3 The voice image display system comprises a car machine 310, a display screen 320, and a starry sky roof 330, the starry sky roof 330 comprises a starry sky roof control module 331 and a starry sky roof display module 332, wherein:
[0076] The car machine 310 is configured to receive target voice information, determine a target in-vehicle position corresponding to the target voice information and a target voice image corresponding to the target voice information according to the target voice information, determine a voice image display device according to the target in-vehicle position, control the display screen 320 to display the target voice image if the voice image display device is the display screen 320, and control the starry sky roof 330 to display the target voice image in a target starry sky roof area corresponding to the target in-vehicle position if the voice image display device is the starry sky roof 330.
[0077] The display screen 320 is configured to display the target voice image.
[0078] The starry sky roof control module 331 is configured to control the starry sky roof display module 332 to display the target voice image in a target starry sky roof area.
[0079] In the above embodiment, the car machine 310 comprises a voice image display device determination module configured to determine the voice image display device as the display screen 320 if the target in-vehicle position corresponds to a first-row position, determine the voice image display device as the starry sky roof 330 if the target in-vehicle position corresponds to a non-first-row position and the opening state of the starry sky roof is open, and determine the voice image display device as the display screen 320 if the opening state of the starry sky roof 330 is closed, wherein the first-row position is a row of positions arranged close to a front windshield in the vehicle, and the non-first-row position is at least one row of positions arranged away from the front windshield in the vehicle.
[0080] In the above embodiment, optionally, the car machine 310 comprises a target in-vehicle position and target voice image determination module, configured to determine a target in-vehicle position corresponding to the target voice information based on a preset sound source positioning algorithm; determine a target operation corresponding to the target voice information based on a preset voice recognition algorithm, execute the target operation, and determine a target voice image corresponding to an operation result of the target operation based on the operation result.
[0081] In the above embodiment, optionally, the target in-vehicle position and target voice image determination module is further configured to determine a target in-vehicle position according to a sound receiving time of each sound receiving device receiving the target voice information and an installation position of each sound receiving device.
[0082] In the above embodiment, optionally, after the target operation is executed, the car machine 310 further comprises a feedback broadcast module, configured to determine to-be-broadcast information according to the operation result of the target operation, and control a voice playing module to broadcast the to-be-broadcast information.
[0083] In the above embodiment, optionally, the car machine 310 comprises a target starry sky roof area determination module, configured to determine a starry sky roof projection position corresponding to the target in-vehicle position according to the target in-vehicle position, and determine a target starry sky roof area according to the starry sky roof projection position and a preset display area.
[0084] In the above embodiment, optionally, the car machine 310 comprises a starry sky roof display control module, configured to determine a video frame sequence corresponding to the target voice image, perform image processing and conversion operations on the video frame sequence to obtain a starry sky roof control signal, wherein the image processing and conversion operations comprise sorting processing operations, contour extraction processing operations, color extraction processing operations and control signal conversion operations, and send the starry sky roof control signal and the target starry sky roof area to the starry sky roof, so that the starry sky roof 330 displays the target voice image in the target starry sky roof area.
[0085] The car machine in the voice image display system of the above embodiment is used to implement the corresponding voice image display method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be described here.
[0086] Based on the same inventive concept, corresponding to any of the above method embodiments, the present application further provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the voice image display method of any of the above embodiments when executing the program.
[0087] Figure 4 Another structure schematic diagram of a voice image display system provided by an embodiment of the present application is shown in FIG. 4. The voice image display system includes a car machine 410, a display screen 420, a starry sky roof 430, and a sound module 440. The car machine 410 includes a voice recognition module 411, a broadcast module 412, and a voice image unit 413. The starry sky roof 430 includes a starry sky roof control module 431 and a starry sky roof display module 432. Figure 4 The car machine 410 is an information entertainment terminal in a vehicle-mounted system, and internally integrates the voice recognition module 411, the broadcast module 412, and the voice image unit 413. Meanwhile, the car machine 410 performs data transmission with the display screen 420, and the display screen 420 implements target voice image display for the driver and the front passenger (first-row position) and switch setting for starry sky roof voice image display for the second-row, the third-row, and the like (non-first-row position).
[0088] The voice recognition module 411 includes voice characteristics such as wake-up, recognition, sound source positioning, and semantic understanding, and can complete overall voice interaction, i.e., voice recognition and sound source positioning, through the voice recognition module 411. The broadcast module 412 is the end of the link in the voice recognition process, and performs broadcast feedback for the execution of the final instruction, state judgment, and interaction. The voice image unit 413 is a voice image built in the car machine 410, and displays corresponding actions in combination with the voice recognition process to improve the overall voice interaction experience through visual feedback.
[0089] It should be noted that in combination with a microphone array, the function of sound source positioning can accurately identify the position (target in-vehicle position) of the user who issues the voice control command (target voice information), such as the second-row left side, the second-row right side, the third-row left side, the third-row right side, and the like. For different positions, the display screen 420 or the starry sky roof 430 is used to display the target voice image, and further, when the starry sky roof 430 is used to display, the target voice image is displayed according to the corresponding position of the starry sky roof 430.
[0090] The starry sky roof control module 431 acquires the switch state of the starry sky roof voice image display for the second-row, the third-row, and the like through data transmission with the car machine 410. When in the on state, the starry sky roof 430 needs to synchronously display the target voice image to the second-row, the third-row, and the like according to the target voice image (when the car machine 410 changes the voice image, the voice image displayed by the starry sky roof 430 at the second-row, the third-row, and the like also needs to be changed synchronously); when in the off state, the starry sky roof 430 maintains the previous set state, and the second-row, the third-row, and the like do not synchronously display the target voice image.
[0091]
[0092] After the starry sky top control module 431 obtains the target voice image of the vehicle machine 410, the vehicle machine 410 needs to synchronize all sequence frame files (the files contain all scenes displayed in voice interaction) of the target voice image to the starry sky top control module 431, and after the starry sky top control module 431 receives the sequence frame files, the starry sky top control module 431 performs the actions of picture sorting, graphic contour extraction, and graphic color extraction on each frame picture, and after the above actions are completed, the starry sky top display module 432 is synchronized in the form of a control signal to perform final effect display.
[0093] It can be understood that the actions of picture sorting, graphic contour extraction, and graphic color extraction on each frame picture and the conversion into the form of a control signal can also be implemented by the vehicle machine 410, in which case the purpose of the starry sky top control module 431 is signal transmission and feedback.
[0094] The sound module 440 performs audio data transmission with the vehicle machine 410, and the sound module 410 completes sound output.
[0095] The starry sky top display module 432 is a medium for final image display, and is controlled by the starry sky top control module 431. When the starry sky top control module 431 obtains the target voice image that needs to be synchronized with the second row and the third row, the starry sky top display module 432 is controlled to display in the form of a sequence frame animation.
[0096] The voice image display system of the above embodiment is used to implement the corresponding voice image display method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.
[0097] Figure 5 A more specific electronic device hardware structure schematic diagram provided by the embodiment is shown, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 for communication within the device.
[0098] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit, central processor), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the embodiments of the present specification.
[0099] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided in the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0100] The input / output interface 1030 is configured to connect an input / output module to realize information input and output. The input / output module can be configured in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0101] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0102] The bus 1050 includes a channel to transmit information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.
[0103] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only include the components necessary to implement the solutions of the embodiments of the present specification, and does not have to include all the components shown in the figure.
[0104] The electronic device of the above embodiments is used to implement the corresponding voice image display method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here again.
[0105] Based on the same inventive concept, the present application also provides a vehicle, wherein the vehicle comprises the voice image display system according to the above embodiments or the electronic device according to the above embodiments.
[0106] Based on the same inventive concept, the present application also provides a non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the voice image display method according to any of the above embodiments.
[0107] The computer readable medium of the embodiments can include permanent and non-permanent, removable and non-removable media, which can be realized by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0108] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the voice image display method according to any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not described here.
[0109] Those skilled in the art should understand that the above discussion of any of the embodiments is only exemplary and is not intended to imply that the scope (including claims) of the present application is limited to these examples; the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of the different aspects of the embodiments of the present application as described above. In order to be brief, they are not provided in detail.
[0110] Additionally, to simplify the description and discussion, and so as not to obscure the embodiments of the application being presented, the well-known functions or constructions of integrated circuit (IC) chips and other components can or can not be shown in the figures and will be omitted as not to unnecessarily obscure the embodiments of the application being presented. Moreover, the devices can be shown in block diagram form in order to avoid unnecessary obscurity of the present embodiments, and this also acknowledges the fact that the details in regard to the implementation of such block diagram devices are highly dependent on the platform within which the present embodiments are to be implemented (i.e., such details should be well within the purview of one of ordinary skill in the art). Where specific details are set forth in order to describe an illustrative embodiment of the application, it will be apparent to one of ordinary skill in the art that the embodiment of the application can be practiced without these specific details. In other instances, detailed descriptions of well-known methods, devices, and materials can be omitted so as not to obscure the description of the present embodiments of the application. It is intended that the specific embodiments disclosed herein are presented by way of example only and that the present application is not limited by the embodiments presented herein.
[0111] Although the present application has been described in connection with certain specific embodiments thereof, many modifications, changes, variations and substitutions will be apparent to those of ordinary skill in the art. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.
[0112] It is therefore intended that the present application cover all such modifications, changes, variations and substitutions that fall within the broad scope of the appended claims. Accordingly, any one or more of the features, functions, structures, or other aspects of the embodiments described herein can be combined in any suitable manner to form additional embodiments, which are also within the scope of the present application. Thus, various additional embodiments of the present application are also contemplated. Therefore, the foregoing description is not intended to be exhaustive or to limit the application to the precise forms disclosed.
Claims
1. A voice image display method, characterized by, The method comprises the following steps: receiving target voice information, determining a target in-vehicle position corresponding to the target voice information and a target voice image corresponding to the target voice information according to the target voice information; determining a voice image display device according to the target in-vehicle position; wherein the voice image display device comprises a display screen or a starry sky roof; if the voice image display device is the display screen, controlling the display screen to display the target voice image; if the voice image display device is the starry sky roof, performing orthographic projection on the starry sky roof according to the target in-vehicle position to determine a starry sky roof projection position corresponding to the target in-vehicle position; expanding the starry sky roof projection position to a target starry sky roof area according to a preset display area, and controlling the starry sky roof to display the target voice image in the target starry sky roof area.
2. The method of claim 1, wherein, The step of determining the voice image display device according to the target in-vehicle position comprises the following steps: if the target in-vehicle position corresponds to a first row position, determining that the voice image display device is the display screen; if the target in-vehicle position corresponds to a non-first row position, determining that the voice image display device is the starry sky roof when the starry sky roof is in an open state, and determining that the voice image display device is the display screen when the starry sky roof is in a closed state; wherein the first row position is a row position close to a front windshield in the vehicle, and the non-first row position is at least one row position away from the front windshield in the vehicle.
3. The method of claim 1, wherein, The step of determining the target in-vehicle position corresponding to the target voice information and the target voice image corresponding to the target voice information according to the target voice information comprises the following steps: determining the target in-vehicle position corresponding to the target voice information based on a preset sound source positioning algorithm; determining a target operation corresponding to the target voice information based on a preset voice recognition algorithm, executing the target operation, and determining a target voice image corresponding to an operation result of the target operation based on the operation result.
4. The method of claim 3, wherein, The step of determining the target in-vehicle position corresponding to the target voice information based on the preset sound source positioning algorithm comprises the following steps: determining the target in-vehicle position according to sound receiving times of each sound receiving device receiving the target voice information and installation positions of each sound receiving device.
5. The method of claim 3, wherein, After executing the target operation, the method further comprises the following steps: determining to-be-announced information according to the operation result of the target operation, and controlling a voice playing module to announce the to-be-announced information.
6. The method of claim 1, wherein, The step of controlling the starry sky roof to display the target voice image in the target starry sky roof area comprises the following steps: determining a video frame sequence corresponding to the target voice image, performing image processing and conversion operations on the video frame sequence to obtain a starry sky roof control signal; wherein the image processing and conversion operations comprise sorting processing operations, contour extraction processing operations, color extraction processing operations and control signal conversion operations; sending the starry sky roof control signal and the target starry sky roof area to the starry sky roof, so that the starry sky roof displays the target voice image in the target starry sky roof area.
7. A voice image display system characterized by comprising: The method comprises the following steps: A vehicle machine, a display screen, and a starry sky roof, the starry sky roof comprising a starry sky roof control module and a starry sky roof display module, wherein: The vehicle machine is configured to receive target voice information, determine a target in-vehicle location corresponding to the target voice information and a target voice image corresponding to the target voice information according to the target voice information, determine a voice image display device according to the target in-vehicle location, control the display screen to display the target voice image if the voice image display device is the display screen, and control the starry sky roof to perform orthographic projection according to the target in-vehicle location if the voice image display device is the starry sky roof, determine a starry sky roof projection location corresponding to the target in-vehicle location, expand the starry sky roof projection location to a target starry sky roof area according to a preset display area, and control the starry sky roof to display the target voice image in the target starry sky roof area. The display screen is configured to display the target voice image. The starry sky roof control module is configured to display the target voice image in the target starry sky roof area of the starry sky roof display module.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the voice image display method according to any one of claims 1 to 6 when executing the program.
9. A vehicle characterized by comprising: The vehicle comprises the voice image display system according to claim 7 or the electronic device according to claim 8.
Citation Information
Patent Citations
Vehicle-mounted rear-row projection display system
CN111487842A
In-vehicle voice interaction method and system
CN113782020A
Starry sky top display method and device, vehicle and storage medium
CN115009160A
Vehicle-mounted service display method, vehicle and program product
CN115826802A