Method and device for generating vehicle-mounted content, vehicle and product

By obtaining vehicle travel information and driving scenarios, and dynamically adjusting the voice content and images of the in-vehicle entertainment system, the problem of boring user experience in the existing technology is solved, more intelligent in-vehicle content generation is achieved, and users' driving or riding experience is improved.

CN120386903APending Publication Date: 2025-07-29MOBILITY ASIA SMART TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410102853.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing in-vehicle entertainment system lacks intelligence during long driving or riding, and cannot dynamically adjust voice content and images according to user needs and vehicle status, resulting in boring and boring user experience.

Method used

By obtaining vehicle travel information, voice content and images that meet user needs are generated, dynamically adjust it according to the vehicle driving scene, optimize content generation based on user input and historical data, and use sensors and servers to make real-time adjustments.

Benefits of technology

It improves the difference and dynamics of voice content and images, provides a more attractive travel experience, and enhances the user's driving or riding fun and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386903A_ABST
    Figure CN120386903A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and equipment for generating vehicle-mounted content, a vehicle and a product. The method may include obtaining travel information of a vehicle. The method may also include generating voice content and / or an image corresponding to the voice content based on the travel information. Further, the method may include adjusting at least a portion of the voice content and / or the image according to a driving scene of the vehicle. According to the method and the device, the vehicle-mounted content with high attraction can be generated according to the travel element, and better travel experience is provided for the user. Besides, the voice content and the image can be dynamically adjusted when the vehicle runs, so that the difference and dynamic nature of the voice content and / or the image can be effectively improved, and the travel experience of the user is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method, device, vehicle, and product for generating in-vehicle content. Background Art

[0002] As the range of vehicle activities becomes wider, the driving time will increase as the journey becomes longer. When users spend a long time driving or riding in a vehicle, they will feel bored. In order to improve the driving experience of users, vehicles can use some in-vehicle devices to provide rich entertainment activities for users, making the driving or riding time more interesting and enjoyable. For example, audio content such as music, videos, news, novels, and stories can be played through the in-vehicle devices in the vehicle. However, the current methods still have certain problems and are not intelligent enough. Summary of the Invention

[0003] According to an exemplary embodiment of the present disclosure, a technical solution for generating in-vehicle content is provided. Voice content and / or images that better meet travel needs can be generated based on travel information, thereby providing an extremely attractive travel experience for users. In addition, during the driving process of the vehicle, the voice content and / or images can be dynamically adjusted according to the driving scenario, thereby effectively enhancing the difference and dynamics of the voice content and / or images, and further improving the travel experience.

[0004] In a first aspect of the present disclosure, a method for generating in-vehicle content is provided. The method may include obtaining travel information of the vehicle. The method may further include generating voice content and / or an image corresponding to the voice content based on the travel information. And the method may further include adjusting at least a part of the voice content and / or the image according to the driving scenario of the vehicle. Implementing the method provided in the first aspect, the in-vehicle content generated according to the travel information better meets the travel needs and preferences of users. In addition, the in-vehicle content can be dynamically adjusted, thereby effectively enhancing the difference and dynamics of the voice content and images, and further optimizing the travel experience.

[0005] In a second aspect of the present disclosure, an electronic device for generating in-vehicle content is provided. The electronic device includes: a processor, and a memory coupled to the processor. The memory has instructions stored therein, and when the instructions are executed by the electronic device, the electronic device executes the method provided in the first aspect of the present disclosure.

[0006] In a third aspect of the present disclosure, a vehicle is provided. The vehicle includes the electronic device provided in the second aspect.

[0007] In a fourth aspect of the present disclosure, there is provided a computer program product tangibly stored on a computer-readable medium and including computer-executable instructions that, when executed, cause a computer to perform the method according to the first aspect of the present disclosure.

[0008] As can be seen from the above description, according to the solutions of the embodiments of the present disclosure, the voice content and images generated based on the travel information of the vehicle are more in line with the travel state of the vehicle, more in line with the preferences of the user, and the information contained in the voice content and images is also richer. Moreover, the method can also dynamically adjust the voice content or images according to the driving scenario of the vehicle, further optimizing the driving experience of the user.

[0009] It should be understood that the summary is provided to introduce a selection of concepts in a simplified form, which will be further described in the detailed description below. The summary is not intended to identify the key features or main features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0011] Figure 1 shows an example environment in which the devices and / or methods according to the embodiments of the present disclosure may be implemented;

[0012] Figure 2 shows a schematic diagram of the system architecture of a vehicle according to some embodiments of the present disclosure;

[0013] Figure 3 shows a schematic diagram for generating in-vehicle content according to some embodiments of the present disclosure;

[0014] Figure 4 shows a flowchart of a method for generating in-vehicle content according to some embodiments of the present disclosure;

[0015] Figure 5 shows a schematic diagram for adjusting the generated images according to the driving scenario according to some embodiments of the present disclosure;

[0016] Figure 6 shows a schematic diagram for panoramic segmentation of an image to determine a background region according to some embodiments of the present disclosure;

[0017] Figure 7 shows a schematic diagram for generating and adjusting in-vehicle content according to some embodiments of the present disclosure; and

[0018] Figure 8 FIG. shows a schematic structural diagram of a device that can be used to implement an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0020] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0021] As the activity range of vehicles becomes wider and wider, the driving time will also increase as the journey becomes longer. When users drive or ride in a vehicle for a long time, they will feel bored. In order to improve the driving experience of users, vehicles can use some in-vehicle devices to provide rich entertainment activities for users, making the driving or riding time of users more interesting and enjoyable. For example, audio content such as music, videos, news, novels, stories, etc. can be played through the in-vehicle devices in the vehicle. When playing the audio content, in order to enrich the driving experience of users, images associated with the voice content can be displayed simultaneously using other in-vehicle devices in the vehicle.

[0022] However, the current method usually plays the audio of fixed content downloaded or generated in advance when playing voice content such as novels or stories, and the pictures associated with the voice content are also fixed, without considering the differences of different users and different vehicles during actual driving, so to a certain extent, it cannot attract users for a long time and cannot completely solve the problem that users feel bored when driving or riding in a vehicle.

[0023] To at least address problems similar to the above, the present disclosure proposes a method, device, vehicle, medium, and product for generating in-vehicle content. The method can generate voice content and an image corresponding to the voice content based on the travel information of the vehicle. The voice content and the image generated in this way are more in line with the user's own needs. In addition, the voice content and the image generated based on the travel information of the vehicle are more in line with the travel state of the vehicle, and the information contained in the voice content and the image is also richer. Moreover, the method can also dynamically adjust the voice content or the image according to the driving scenario of the vehicle, further improving the user's driving experience and making the generation and adjustment of the voice content and the image more intelligent.

[0024] The following refers to Figures 1 to 8 to illustrate the basic principles and several exemplary implementations of the present disclosure. It should be understood that these exemplary embodiments are given only to enable those skilled in the art to better understand and then implement the embodiments of the present disclosure, rather than limiting the scope of the present disclosure in any way.

[0025] Figure 1 FIG. shows a schematic diagram of an exemplary environment 100 in which the device and / or method according to an embodiment of the present disclosure can be implemented. As Figure 1 shown, in some embodiments, the exemplary environment 100 includes a vehicle 102 and a server 104. The vehicle 102 can be communicatively connected to the server 104 via a network to send the collected data to the electronic device 104, and the server 104 generates or adjusts the in-vehicle content based on the collected data. Among them, the data collected by the vehicle 102 can include vehicle data and environmental data of the vehicle's surrounding environment. According to an embodiment of the present disclosure, the vehicle 102 refers to any type of motorized or non-motorized vehicle that can carry people and / or objects and is movable. As Figure 1 shown, the vehicle 102 is illustrated as a car. It should be understood that although the vehicle 102 is Figure 1 illustrated as a car in

[0026] In some embodiments, server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0027] Combined with the above Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is depicted. It should be understood that the environment 100 is merely illustrative and does not limit the scope of the present disclosure. The environment 100 may include Figure 1 More components are shown, and various components in environment 100 may also be implemented in different ways.

[0028] Figure 2 Schematic diagram of the system architecture of a vehicle according to some embodiments of the present disclosure is shown. Figure 2 As shown, the vehicle 102 includes multiple subsystems, such as a sensor system 202, an onboard device system 204, a user interface 206, etc. A subsystem may include multiple components or devices, and each subsystem is connected to the components or devices therein via wired or wireless connections.

[0029] In some embodiments, the sensor system 202 may include a plurality of sensors disposed at various locations in the vehicle (eg, in front, on the left, or at the rear of the vehicle) for sensing information about the environment surrounding the vehicle 102. Figure 2 As shown, the sensor system 202 may include but is not limited to a global positioning system 208 (the positioning system may be a GPS system, a BeiDou system, or other positioning systems), an inertial measurement unit (IMU) 210, a radar sensor 212, a visual sensor 214, and the like. The sensor system 202 may also include sensors of the internal systems of the detected vehicle (e.g., an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). The visual sensor 214 may include a camera or a video camera (e.g., a monocular camera, a multi-camera, a depth camera, etc.). In some embodiments of the present disclosure, the data collected by the sensor system 202 may include information about the vehicle's route, speed, acceleration, and the environment surrounding the vehicle (e.g., road conditions, pedestrians, road signs, buildings, obstacles, etc.) during driving.

[0030] like Figure 2As shown, in some embodiments, the in-vehicle device system 204 may include devices such as an air conditioner 216, ambient lights 218, aromatherapy 220, a sound system 222, seats 224, etc. The user interface 206 may be used to provide information to or receive information from the vehicle's users. Optionally, the user interface 206 may include one or more input / output devices, such as a wireless communication system, a microphone, a speaker, a touch display screen, etc. The user interface 206 can receive input information entered by the user in various ways. The input information can be various information (such as "Play fairy tales", "Tell a story about a rabbit and a tortoise", "Navigate to ◇◇ building", etc.) entered by the user through voice input, touch input, etc.).

[0031] Figure 3 A schematic diagram for generating in-vehicle content according to some embodiments of the present disclosure is shown. As Figure 3 shown, the input of the server 104 is the travel information 302 of the vehicle. Among them, the travel information 302 of the vehicle can be collected by the vehicle using a sensor system (such as Figure 2 the sensor system 202 shown in), or obtained from an electronic device connected to the vehicle. The travel information 302 of the vehicle may include any one or more of weather, travel time period, driving route, representative buildings on the driving route to be passed through in advance, and traffic information on the front section of the road. Among them, the driving route may include the distance, width, number of lanes on the section, and type of lanes of multiple sections (straight sections, left-turn sections, right-turn sections) that make up the driving route.

[0032] In block 304, the server 104 may generate corresponding voice content 306 and / or an image 308 corresponding to the voice content according to the travel information 302. For example, in one example, when the travel information 302 is "raining, 8 pm, the representative building is the ancient city wall", the generated voice content 306 may be "On a pitch-black night, it was pouring with rain. Suddenly, a beam of light illuminated the night sky. The leader of the martial arts alliance in a black bamboo hat and the female knight in a red robe stood on the city wall...". The generated image 308 is associated with the generated voice content 306.

[0033] It should be noted that after generating the voice content 306 and the image 308, the voice content 306 and the image 308 can be sent to the vehicle, and the in-vehicle devices of the vehicle can display the image 308 associated with the voice content while playing the voice content 306, so as to provide rich entertainment methods for users and optimize the user's travel experience.

[0034] It can be understood that, in order to provide users with a more diverse entertainment experience, images associated with the voice content can be presented while playing the voice content, thereby enhancing the richness of information presentation. It should be noted that, different from the way video software plays videos (such as movies and TV dramas), in some embodiments of the present disclosure, images associated with the exciting plot of the voice content can be generated at the exciting plot of the voice content, while no images are generated at the ordinary plot of the voice content. For example, images associated with the duel plot of a martial arts story can be displayed when playing the duel plot.

[0035] In this way, users do not need to constantly watch the images and can instantaneously view the associated images at the exciting plot of the voice content, so that the driver can focus on driving to ensure driving safety. In addition, it can also protect the eyesight of passengers sitting in the back row or the co-pilot of the vehicle.

[0036] In block 312, the server 104 can dynamically adjust the voice content and / or the image corresponding to the voice content according to the driving scenario 310 of the vehicle, so as to effectively improve the dynamics and difference of the voice content or the image. Among them, the driving scenario 310 of the vehicle can be collected by the vehicle and transmitted to the server 104 through the network. For example, during the driving process, the vehicle can collect driving information through the sensor system. The driving information can include driving state information related to the vehicle itself, such as the position of the vehicle, the driving speed of the vehicle, the driving direction of the vehicle, and the driving time period of the vehicle. The driving information can also include environmental information related to the surrounding environment of the vehicle, such as the road conditions in front of the vehicle during driving (such as traffic congestion, whether there are obstacles on the road), the sounds in the surrounding environment of the vehicle, and the objects in the surrounding environment of the vehicle. The objects in the environment can include, but are not limited to, pedestrians, other vehicles, buildings, traffic signs, lane lines, etc.

[0037] In this way, the embodiments of the present disclosure can generate voice content and / or images that better meet the travel needs according to the travel information, thereby providing users with an extremely attractive travel experience. In addition, during the driving process of the vehicle, the voice content and / or the image can be dynamically adjusted according to the driving scenario, so as to effectively improve the difference and dynamics of the voice content and the image, and further optimize the travel experience.

[0038] The following combines Figure 4 to describe the flowchart of the method 400 for generating in-vehicle content according to an embodiment of the present disclosure. It should be noted that the method 400 according to an embodiment of the present disclosure can be implemented, for example, at Figure 1 the server 104 shown in. Additionally or alternatively, it can also be implemented at Figure 1 the vehicle 102 shown in, for example, at the controller of the vehicle.

[0039] At block 402, method 400 may obtain travel information of a vehicle. Among them, the travel information of the vehicle may include, but is not limited to, any one or more of weather (such as sunny, cloudy, rainy, snowy) information, driving route, departure time period (such as morning, noon, evening), driving duration (such as 30 minutes, 1 hour, 2 hours, 2 hours 30 minutes), temperature (such as 20°, 30°, -5°). Among them, the driving route includes the departure place, the destination, and each intermediate place passed through in the interval from the departure place to the destination. In some embodiments, weather information may be obtained through temperature sensors, humidity sensors, etc. in the vehicle, or weather information may be queried through a meteorological authoritative website connected to the vehicle's in-vehicle system.

[0040] At block 404, method 400 may generate speech content and / or an image corresponding to the speech content based on the travel information. Among them, the speech content is the content that can be expressed in language and is created. For example, it can be a story, a poem, music, a script, a novel, etc. Among them, the type of the story can be a fairy tale, a martial arts story, a fable story, a horror story, etc. For example, a language model may be used to generate speech content and / or an image corresponding to the speech content according to the input information and the travel information. Further, a prompt word may be generated according to the travel information. The language model may process the prompt word to generate the corresponding speech content and / or an image corresponding to the speech content. Among them, the prompt word refers to the text input to the language model, which is used to prompt or guide the language model to give an output result that meets the expectations. The language model is a model that uses deep learning technology to learn the laws and knowledge of language based on a large amount of text data, so as to be able to generate natural and fluent content. In some examples, the language model may be a conversational language model. The conversational language model may adjust the speech content and / or the image according to the user's multi-round input information or the driving scenarios collected multiple times.

[0041] At block 406, method 400 may adjust at least a part of the speech content and / or the image according to the driving scenario of the vehicle. Among them, the driving scenario is a comprehensive reflection of the environment and driving behavior within a certain time and space range, and the driving scenario may be used to characterize the driving state of the vehicle during driving and the environmental state of the surrounding environment during driving. Among them, the driving state may include, but is not limited to, the driving speed, driving direction, and current positioning information of the vehicle. In other embodiments, the driving state of the vehicle may further include the state of the vehicle's electrical equipment, such as whether the lighting device is turned on, the rotation angle and rotation frequency of the steering wheel, whether the windshield wiper is turned on, whether the sunroof is turned on, etc.

[0042] In some embodiments of the present disclosure, the environmental state of the surrounding environment refers to the objects contained in the surrounding environment and the states and behaviors of the objects. Among them, the objects contained in the surrounding environment can be pedestrians, buildings (such as office buildings, convenience stores, libraries, etc.), other vehicles, animals (such as pet cats, pet dogs, etc.), lane lines, traffic signs, obstacles, etc. The environmental state of the surrounding environment can also include meteorological conditions (such as weather conditions, environmental temperature, environmental humidity, etc.). In some examples, the state or behavior of a pedestrian refers to the behavioral characteristics (such as walking, running, brisk walking, etc.) and appearance characteristics of the object. In other embodiments, the driving scenario can also include the traffic congestion degree of the road along which the vehicle travels during driving and the sounds around the vehicle (such as sound types, sound volumes, etc.).

[0043] In some embodiments of the present disclosure, images or video streams of the surrounding environment can be collected through the sensor system of the vehicle. For example, a vision sensor deployed at the front of the vehicle can collect an image of the environment in front of the current position of the vehicle. After collecting the image, image recognition can be performed on the image to determine the environmental state of the surrounding environment. In some embodiments, an image can be collected at preset time intervals. The preset time interval can be set to 1 minute, 2 minutes, 5 minutes, etc. according to the actual driving situation of the vehicle, and no limitation is made here.

[0044] In some embodiments of the present disclosure, there are many types of forms for adjusting the voice content according to the driving scenario. In some embodiments, the rhythm of the voice content can be adjusted according to the driving speed of the vehicle. For example, when the vehicle is driving fast, the rhythm of the voice content can be changed to a fast rhythm (for example, the rhythm of the voice content can be changed to a fast rhythm by shortening the length of sentences, deleting some paragraph scenarios, etc.). In some embodiments, the content length of the voice content can be adjusted according to the traffic congestion situation. For example, when the road ahead is severely congested, the content length of the voice content can be appropriately increased. In other embodiments, the scenario of the voice content can be adjusted according to the environmental state of the surrounding environment. For example, when the generated voice content is "The white clouds in the sky turned into fairies wearing colorful clouds in the morning light and smiled as they came to the busy market...", and the environmental state of the surrounding environment is "A flower exchange meeting is in progress", the voice content can be adjusted to "The white clouds in the sky turned into fairies wearing colorful clouds in the morning light and smiled as they came to the bustling flower street..." according to the environmental state of the surrounding environment.

[0045] It should be noted that the overall context of the voice content will not be adjusted when adjusting the voice content. For example, in the case of the fairy tale "Snow White" as the voice content, the general plot of the "Snow White" story does not need to be adjusted, but the story scene, plot, etc. can be adjusted. Additionally or alternatively, when generating the voice content, a general plot of the story can be generated, and the specific plot and scene of the story can be supplemented according to the driving scenario. In this way, the integrity of the voice content can be ensured, and the situation where the voice content deviates from the normal track due to unexpected situations can be avoided, thereby preventing the impact of reverse teaching.

[0046] In some embodiments of the present disclosure, there can be various types of adjusting the form of the image according to the driving scenario. In some embodiments, the state of the characters in the image can be adjusted according to the state of the people in the surrounding environment. For example, the clothing of character A in the image can be adjusted according to the clothing of pedestrian A in the surrounding environment. In some embodiments, the buildings included in the background area of the image can be adjusted according to the buildings in the surrounding environment. For example, when the vehicle passes by a garden during driving, the buildings included in the background area of the image can be changed to a garden.

[0047] Figure 5 A schematic diagram of an image generated by adjusting according to the driving scenario according to some embodiments of the present disclosure is shown. As Figure 5 shown, according to the plot included in the voice content, the building in the pre-generated image 502 is a self-built house 504 in the countryside. When the driving scenario of the vehicle is an urban scenario and the buildings passed by are high-rise buildings, the self-built house 504 in the pre-generated image 502 can be adjusted to a high-rise building 508 to generate an adjusted image 506.

[0048] In some other examples, the image can also be adjusted according to the weather state at the current position of the vehicle. For example, when the weather at the current position of the vehicle is "pouring rain", raindrops can be added to the corresponding generated image. It can be understood that each of the embodiments described herein is only to help those skilled in the art better understand the idea of the present disclosure and is not intended to limit the scope of the present disclosure in any way.

[0049] In this way, the embodiments of the present disclosure can generate voice content and / or images that better meet the travel needs according to the user's input information and travel information, thereby providing an extremely attractive travel experience for the user. Additionally, during the driving process of the vehicle, the voice content and / or images can be dynamically adjusted according to the driving scenario, so as to effectively improve the difference and dynamics of the voice content and images, and further optimize the travel experience.

[0050] In some embodiments, in order to make the generated voice content and images more in line with the user's preferences, the corresponding voice content and images can be generated by combining the user's input information and the vehicle's travel information. Among them, the types of input information include sound, text, image or video, or any combination of the above types. The input information can be information input by the user according to their own needs through voice input, keyboard input, gesture input, etc. For example, in some examples, the input information can be "Please tell a short story that makes children understand humility", "Write a seven-character quatrain with kindness as the theme", "Tell me a more thrilling ghost story", etc.

[0051] In some embodiments, in order to make the voice content and images generated before the trip more in line with the user's preferences and closer to the actual travel needs, the voice content and / or images can be generated by combining the vehicle's historical driving data. Among them, the historical driving data refers to the driving data generated by the vehicle during previous driving processes. For example, it can include but is not limited to historical driving speed, historical driving route associated with the travel information, representative buildings on the historical driving route, historical surrounding environment associated with the travel information, historical road condition information, historical driving occasions (city, rural road, airport, wharf), etc.

[0052] For example, the navigation route included in the travel information is Point A - Point B - Point C. Determine the driving itinerary (Point A - Point B - Point C) with the same route as this navigation route from the historical driving data. Then, determine the representative buildings and the environmental status of the surrounding environment along this driving itinerary. Based on this environmental status and the building, the generated voice content can be determined. For example, in an exemplary embodiment, if the buildings passed by during the driving itinerary from Point A to Point B to Point C are all Chinese garden-style, the age background included in the voice content can be set to "ancient times", and the form of the buildings included in the background of the generated image can also be set to Chinese garden-style.

[0053] In other embodiments, the voice content can also be generated by combining the historical input information of the user (such as the driver, passenger, etc.). Among them, the historical input information can be the multi-round dialogue information between the user and the language model. The historical input information can represent the user's preferences and experience feelings. Therefore, the voice content generated based on the historical input information will be more in line with the user's preferences, thereby further enhancing the user's travel experience.

[0054] In some embodiments, after generating the voice content, multiple images can be generated according to the plot contained in the voice content. Herein, the plot contained in the voice content refers to the development process of the story. For example, the plot can be each event that occurs in the text contained in the voice content. For example, in the case where the voice content is the fairy tale of "Snow White", the plot can include "the princess fell into a deep sleep after eating the apple", "the seven dwarfs took care of the princess", "the prince and the princess lived happily ever after", etc. In some embodiments, after determining the plot of the voice content, the information contained in each plot can be processed for text information to determine the keywords corresponding to each plot. The keywords can be verbs, nouns, words that appear frequently, etc. For example, in the case where the plot is "the princess fell into a deep sleep after eating the apple", the keywords can be "princess", "eating the apple", "falling into a deep sleep". After that, images corresponding to the plot can be generated according to the keywords to vividly tell the fairy tale of "Snow White" to the user and enhance the user's travel experience.

[0055] In some embodiments, the number of generated images cannot be too large to prevent the user from paying too much attention to the images, thereby damaging the user's eyesight or reducing the driving safety. The number of generated images cannot be too small, otherwise it will reduce the user's travel experience. On this basis, the number of generated images can be determined according to the travel information or input information. In some examples, when the expected travel duration is long, a larger number of images can be generated; when the expected travel duration is short, a smaller number of images can be generated. In some examples, if parents take their children on a trip and do not want the children to watch too many images and affect their eyesight, a smaller number of images can be generated. Additionally or alternatively, in other embodiments, the number of generated images can also be determined according to the number of plots or chapters of the voice content.

[0056] Furthermore, in other embodiments, since the vehicle may encounter many uncontrollable factors during driving (such as traffic congestion on the road ahead due to an accident), these may affect the travel duration and also affect the user's driving experience to a certain extent. Therefore, the number of images can be dynamically adjusted according to the driving scenario during the vehicle's driving process. In some examples, when the traffic congestion on the road ahead is relatively serious, n images can be added according to the plot of the voice content and the environmental state of the vehicle's surrounding environment. The number of n can be determined according to the traffic congestion situation. For example, when the traffic congestion on the road ahead is relatively serious, an indication for generating additional images can be given so that the language model can generate multiple additional images according to the indication.

[0057] In some embodiments, to further optimize the user's driving experience, the background of the image can be dynamically adjusted in combination with the real situation of the vehicle's surrounding environment. For example, images of the vehicle's surrounding environment can be collected through vision sensors (such as cameras) in the vehicle's sensor system. Images of the vehicle's surrounding environment can include static or moving images of the environment collected in various directions of the vehicle, and the various directions include the forward direction, backward direction, left direction, right direction, etc. of the vehicle. In some examples, the image can include a 360-degree surround view image collected using multiple cameras of the vehicle. In other examples, the image can also include a panoramic image stitched together from multiple images collected in different directions, which is not limited here.

[0058] In some embodiments, after collecting the images of the surrounding environment, semantic recognition can be performed on the images to determine the background area corresponding to the images. For example, the background area corresponding to the images can be determined by means of panoramic segmentation, semantic segmentation, or instance segmentation. Figure 6 A schematic diagram of panoramic segmentation of an image to determine the background area according to some embodiments of the present disclosure is shown. As Figure 6 shown, the image 602 can be input into a panoramic segmentation network (not shown in the figure), and the panoramic segmentation network outputs a panoramic segmentation image 604. As Figure 6 shown, the panoramic segmentation image 604 can include multiple image blocks. For example, image block 1 is the road surface, image block 2 belongs to the vehicle, image block 3 is the sky, image block 4 belongs to the building, image blocks 5 and 6 both belong to the trees, and image block 7 is a person. Among them, the panoramic segmentation network can be trained through an image sample set, and the image sample set can include target objects of common categories.

[0059] In some embodiments, after determining the background area corresponding to the images of the surrounding environment, the background of the image associated with the voice content can be adjusted based on the background area. For example, the background of the image of the surrounding environment can be input into a language model, and after being processed by the language model, an adjusted image is output. Further, the background of the generated image can be replaced with the background of the image of the surrounding environment, or the background of the image of the surrounding environment can be fused with the background of the generated image to determine the background of the adjusted image. Here, "associated" means the image associated with the voice content played at the current moment.

[0060] In this way, the background of the image can be dynamically adjusted in combination with the real situation of the vehicle's surrounding environment, thereby providing a more immersive entertainment experience for the user. That is to say, the generated virtual scene and the real-world scene are fused to provide a more real and immersive entertainment experience for the user.

[0061] In some embodiments, in order to make the generated image more attractive and further optimize the user's travel experience, the human features included in the image of the vehicle's surrounding environment can also be used to adjust the human features in the generated image. Herein, the human features refer to one or more of appearance features and behavioral features. Appearance features can include, but are not limited to, looks, body shape (such as tall, short, fat, thin), hairstyle (such as long hair, short hair, black hair, yellow hair), clothing (dressing style), emotion (such as happy, angry, sad, joyful), etc. Behavioral features can include, but are not limited to, walking state, body orientation, etc. For example, in one example, when the human features included in the collected image of the surrounding environment are "a woman wearing a red suit", an instruction for adjusting the image (such as "replace the woman's clothes in the generated Image A with a red suit") can be generated according to the human features, so that the language model can adjust the corresponding image according to the instruction, making the human features in the image meet the requirements of the instruction.

[0062] In some embodiments, the user can express their needs in various forms (such as images, sounds, texts, etc.). On this basis, the modalities of travel information and input information may be of various types. For example, the input information can be an image, a speech sequence, and a text composed of multiple words. Herein, the modality refers to some ways of expressing or perceiving things, and each source or form of information can be called a modality. Thus, the prompt words generated based on travel information or input information are also of multiple modalities. The prompt words of multiple modalities are input into the pre-trained model, and the pre-trained model outputs the corresponding speech content and / or the image corresponding to the speech content. Herein, the pre-trained model refers to a pre-trained language model. For example, the initial language model can be trained in a self-supervised manner, and the trained speech model can be fine-tuned with manually labeled data to determine the pre-trained model. Herein, fine-tuning refers to further adjusting and training the language model with manually labeled data in a specific application scenario to make the language model more suitable for a specific application scenario.

[0063] Since the prompt words of different modalities can characterize the user's needs from multiple dimensions, the speech content and the image generated in the above manner can meet the user's needs in all aspects, further optimizing the user's travel experience.

[0064] In some embodiments, the plot of the voice content and / or the image can also be adjusted according to whether a detection event exists in the driving scenario. For example, the driving images during the vehicle driving process can be collected through the vehicle's sensor system. Then, image recognition is performed on the driving images to determine whether a detection event exists in the driving images. In some embodiments, the user can pre-construct a preset detection event library according to actual needs. The preset detection event library can include multiple preset detection events, such as running a red light event, large event (such as flag-raising ceremony, opening ceremony, etc.). In some embodiments, the driving images and the preset detection events can be compared and recognized through an image recognition algorithm to determine the detection event from multiple preset detection events. For example, an image recognition algorithm based on machine learning can be used to recognize the detection event.

[0065] In some embodiments, after determining the detection event included in the driving scenario, the plot of the voice content and / or the image corresponding to the plot can be adjusted according to the detection event. For example, in an example, when the plot of the voice content to be played at the next moment is "a man and his group of good friends go to participate in a patriotic preaching activity" and the detection event is "flag-raising ceremony", the voice content can be adjusted to "a man and his group of good friends go to participate in the flag-raising ceremony".

[0066] In some embodiments, since the surrounding environment of the vehicle is constantly changing during driving and the user's perception of the real world is also constantly changing, in order to ensure the user's travel experience, the working parameters of the in-vehicle device for playing the voice content can be adjusted according to the driving scenario, so that the playing effect corresponds to the environmental state of the surrounding environment. The in-vehicle device for playing the voice content can be a speaker, an audio system, etc. For example, when the surrounding environment is noisy, the volume of the in-vehicle device for playing the voice content can be turned up; while when the surrounding environment is quiet with little sound, the volume of the in-vehicle device for playing the voice content can be turned down.

[0067] It can be understood that in other embodiments, the working parameters of the in-vehicle device for playing the voice content can also be adjusted according to the plot, plot, etc. included in the voice content. For example, when the corresponding plot of the voice content being played in the vehicle is relatively soothing, the volume of the in-vehicle device for playing the voice content can be turned down; while when the voice content being played in the vehicle is the climax plot of the story (such as the duel plot in a martial arts story), the volume of the in-vehicle device for playing the voice content can be turned up. Additionally or alternatively, the working parameters of the in-vehicle device can also be adjusted according to the user's feedback information.

[0068] In some embodiments, in order to provide users with an immersive travel experience, the in-vehicle devices of the vehicle can be combined to create a good atmosphere for users. Among them, the in-vehicle devices can include one or more of ambient lights, headlights, air conditioners, seats, and fragrances. For example, the in-vehicle control system can determine the content being played and the image being displayed at the current moment based on the voice content and image collected in real time. According to the content being played and the image being displayed, the corresponding in-vehicle device is determined, and the working parameters of the in-vehicle device are adjusted so that the working state of the in-vehicle device corresponds to the content being played and the image being displayed, in order to provide users with a richer travel experience. Among them, when the in-vehicle device is an air conditioner, the working parameters of the air conditioner can include temperature, wind direction, wind speed, etc.; when the in-vehicle device is an ambient light, the working parameters of the ambient light can include color, brightness, whether to flash, flash frequency, etc.

[0069] For example, in one example, the color of the ambient light can be adjusted according to the color of the picture included in the image (for example, the color of the atmosphere is adjusted to the color of the picture). In other embodiments, the working parameters of the air conditioner or seat can be adjusted according to the scene corresponding to the voice content. For example, when the scene is a "snowing scene", the working mode of the air conditioner can be adjusted to the "cooling mode", and the temperature of the air conditioner can be adjusted to "16 degrees". When the scene is a "racing scene", the seat can be adjusted to the "vibration mode", and the vibration frequency of the seat can be adjusted adaptively.

[0070] In this way, it is possible to use the in-vehicle devices in the vehicle to create an environment and atmosphere corresponding to the voice content and the picture for users, making users feel more immersive and improving the travel experience of users.

[0071] Figure 7 The figure shows a schematic diagram for generating and adjusting in-vehicle content according to some embodiments of the present disclosure. As Figure 7 shown, in block 702, the travel information of the vehicle can be extracted based on the information input by the user and the information collected by the vehicle. For example, the navigation route included in the travel information can be generated according to the destination input by the user and the current location information of the vehicle. In block 704, the style of the voice content to be generated can be determined according to the input information input by the user. For example, the user can input "Tell me a story of the suspense genre" according to their own preferences.

[0072] At block 706, prompt words can be generated based on travel information and style. Among them, the modalities of the prompt words can have multiple types, such as images, sounds, texts, etc. At block 708, based on the travel information and style, a language model can be used to generate a story outline in that style. And, at block 710, a language model can also be used to generate an image associated with the story outline, such as a cover image. Here, the story outline refers to a summary description of a brief plot and main characters, usually used to convey the core content and basic structure of the story. In some embodiments, corresponding images can be generated according to the content at important plot points. The number of images can be determined by the number of plots or the estimated driving duration, or can be set by the user according to actual personal needs. It can be understood that the language model is generally set in a cloud server connected to the vehicle.

[0073] At block 712, in order to dynamically adjust the story and the image, new prompt words can be generated according to the driving scenario of the vehicle. In the form of a multi-round dialogue, the language model is used to continuously adjust or optimize the content of the story and the image. It can be understood that at block 714 and block 716, according to the driving scenario at the current moment, the language model can be used to adjust the plot to be played at the next moment or the image to be displayed at the next moment. After adjusting to obtain the plot to be played at the next moment and the image to be displayed, the image and the plot can be sent to the vehicle. At block 718, the plot can be played by a first in-vehicle device (such as a stereo, a speaker, etc.) in the vehicle. At block 720, the image can be displayed by a second in-vehicle device (such as a display screen) in the vehicle or other terminal devices connected to the vehicle (such as a mobile phone, a personal computer, a tablet computer, etc.). It can be understood that playing the plot and displaying the image are carried out synchronously.

[0074] In this way, the embodiments of the present disclosure can generate voice content and / or images that better meet travel needs according to the user's input information and travel information, thereby providing an extremely attractive travel experience for the user. In addition, during the driving process of the vehicle, the voice content and / or images can also be dynamically adjusted according to the driving scenario, so as to effectively improve the difference and dynamics of the voice content and images, and further optimize the travel experience.

[0075] It should be understood that the implementation manners illustrated above are only illustrative. According to the actual application situation, Figures 1 to 7 the architecture or process illustrated in Figures 1 to 7 may have other different forms, and may also include more or fewer one or more functional modules and / or units. These modules and / or units can be partially or fully implemented as hardware modules, software modules, firmware modules, or any combination thereof. The embodiments of the present disclosure do not limit this.

[0076] It should be understood that the specific names and / or protocols of the various components in the systems described herein are merely for the purpose of assisting those skilled in the art in better understanding the ideas of the present disclosure and are not intended to limit the scope of the present disclosure in any way. Moreover, in other embodiments, more or better components may be included, or alternative components having the same or similar functions may be included.

[0077] Figure 8 FIG. shows a schematic structural diagram of an exemplary device 800 that can be used to implement some embodiments according to the present disclosure. The device 800 can be implemented as a server or a PC, etc., and the embodiments of the present disclosure do not limit the specific implementation type of the device 800. As Figure 8 shown, the device 800 includes a central processing unit (CPU) 801, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0078] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0079] The processing unit 801 can execute the various methods and / or processes described above, such as Figure 4 the method shown in. For example, in some embodiments, the method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU 801, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the CPU 801 can be configured to execute the method by any other appropriate means (such as by means of firmware, etc.).

[0080] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), Systems on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0081] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.

[0082] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0083] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0084] The computer program instructions for performing the operations of the present disclosure can be assembly instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions can be executed entirely on a user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.

[0085] These computer-readable program instructions can be provided to a processing unit of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, implement means for realizing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions which implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram. The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0086] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0088] In addition, although the operations are depicted in a particular order, it should be understood that such operations are required to be performed in the particular order shown or in a sequential order, or that all illustrated operations should be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation may also be implemented separately or in any suitable sub-combination in multiple implementations.

[0089] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

[0090] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for generating in-vehicle content, comprising: Obtaining travel information of a vehicle; Generating voice content and / or an image corresponding to the voice content based on the travel information; Adjusting at least a part of the voice content and / or the image according to the driving scenario of the vehicle.

2. The method according to claim 1, wherein generating voice content and / or an image corresponding to the voice content based on the travel information comprises: Obtaining input information of a user; And Generating voice content and / or an image corresponding to the voice content based on the input information and the travel information.

3. The method according to claim 1, wherein generating voice content and / or an image corresponding to the voice content based on the travel information comprises: Obtaining historical driving data of the vehicle based on the travel information; And Generating the voice content and / or an image corresponding to the voice content based on the historical driving data and the travel information.

4. The method according to claim 1, further comprising: Determining the number of generated images based on the travel information and / or the voice content; And Dynamically adjusting the number of generated images according to the driving scenario information.

5. The method according to claim 1, wherein generating voice content and / or an image corresponding to the voice content based on the travel information comprises: Generating the voice content based on the travel information; Obtaining a plurality of plots of the voice content and a plurality of keywords corresponding to the plurality of plots; And Generating the corresponding image based on the plurality of plots and the plurality of keywords.

6. The method according to claim 1, wherein adjusting at least a part of the voice content and / or the image based on the driving scenario information of the vehicle comprises: Obtaining an image of the surrounding environment of the vehicle collected during driving; Performing semantic recognition on the image to determine background features corresponding to the image; And Adjusting the background of at least a part of the images based on the background features.

7. The method according to claim 1, wherein adjusting at least a part of the voice content and / or the image based on the driving scenario information of the vehicle comprises: Obtaining an image of the surrounding environment of the vehicle collected during driving; Determining human features included in the image by performing image detection on the image; And Adjusting the human features of at least a part of the images based on the human features.

8. The method according to claim 1, wherein generating voice content and / or an image corresponding to the voice content based on the travel information comprises: Determining prompt words in multiple modalities based on the travel information; And Inputting the prompt words in multiple modalities into a pre-trained model, and outputting the corresponding voice content and / or an image corresponding to the voice content by the pre-trained model.

9. The method according to claim 1, wherein adjusting at least a part of the voice content and / or the image based on the driving scenario comprises: Determining detection events included in the driving scenario; and Adjust at least a part of the voice content and / or the image according to the detection event.

10. The method according to claim 2, wherein generating the voice content and / or the image corresponding to the voice content based on the input information and the travel information comprises: Determine the type of the voice story to be generated included in the input information; and Generate the voice content corresponding to the type and / or the image corresponding to the voice content.

11. The method according to claim 1, further comprising: Determine a first working parameter of a first in-vehicle device for playing the voice content in the vehicle based on the voice content and the driving scenario; Adjust the working parameter of the first in-vehicle device to the first working parameter.

12. The method according to claim 1, further comprising: Determine a second in-vehicle device for assisting in playing the voice content and displaying the image in the vehicle and a second working parameter corresponding to the second in-vehicle device based on the voice content, wherein the second in-vehicle device includes any one of an ambient light, a seat, an air conditioner, and a speaker; Adjust the working parameter of the second in-vehicle device to the second working parameter so that the working state of the second in-vehicle device corresponds to the voice content.

13. An electronic device, comprising: A processor, and A memory coupled to the processor and storing instructions that, when executed by the processor, cause the device to perform the method according to claims 1 to 12.

14. A vehicle, comprising the electronic device according to claim 13.

15. A computer program product, the computer program product being tangibly stored on a non-volatile computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the steps of the method according to any one of claims 1 to 12.