Voice broadcasting method and device for visually impaired passengers and storage medium

By acquiring vehicle environmental data and using a large language model to generate voice broadcast information, the problem of information loss for visually impaired passengers during their ride is solved, achieving logically coherent environmental introductions and dynamic broadcasts, thus improving the riding experience.

CN120895016APending Publication Date: 2025-11-04CHENGDU GREAT WALL MOTOR R&D CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510884202.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Visually impaired passengers are unable to obtain information about the vehicle's status and surrounding environment during their journey, leading to feelings of unease and anxiety.

Method used

By acquiring data on the environment of the target vehicle, a large language model is used to generate logically coherent and scene-appropriate voice broadcast information to introduce the vehicle environment to visually impaired passengers. The broadcast duration and priority are dynamically adjusted to ensure the timeliness and importance of the information.

Benefits of technology

It significantly enhances the sense of control and travel experience for visually impaired passengers, alleviates anxiety caused by information gaps, and provides detailed, real-time environmental descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895016A_ABST
    Figure CN120895016A_ABST
Patent Text Reader

Abstract

The invention discloses a voice broadcasting method and device for visually impaired passengers and a storage medium, and relates to the technical field of vehicle-mounted voice broadcasting. The scheme comprises the steps of obtaining environment data of an environment where a target vehicle is located; and generating voice broadcast information according to the environment data by using a large language model so as to introduce the environment of the target vehicle to visually impaired passengers taking the target vehicle. Therefore, the environment data is integrated into the voice broadcast information, so that the visually impaired passenger can know the riding environment based on the voice broadcast content, the visual information gap is filled, the travel control feeling is effectively improved, the unstable emotion caused by information loss is relieved, and the riding experience of the visually impaired passenger is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicle voice broadcast, in particular to a voice broadcast method for visually impaired passengers, a device and a storage medium. BACKGROUND

[0002] Normal users often pass the time and relieve their emotions by watching the outside scene when taking a vehicle. However, visually impaired passengers cannot perceive the external environment due to visual impairment, which makes it difficult for them to obtain information about the driving state of the vehicle, the road condition and the changes in the surrounding environment during the journey. The lack of such information leads to unstable emotions of visually impaired passengers when taking a vehicle, and may even cause anxiety and tension. Therefore, it is particularly important to solve the problem that visually impaired passengers cannot obtain information about the vehicle and the external environment when taking a vehicle. SUMMARY

[0003] Therefore, the present application is committed to providing a voice broadcast method for visually impaired passengers, a device and a storage medium, which can introduce the environment where the vehicle is located to visually impaired passengers, help them better understand the surrounding situation, relieve their unstable emotions and improve the experience of taking a vehicle.

[0004] According to a first aspect of the present application, a voice broadcast method for visually impaired passengers is provided, comprising: obtaining environment data of an environment where a target vehicle is located; using a large language model to generate voice broadcast information according to the environment data; and the voice broadcast information is used to introduce the environment where the target vehicle is located to a visually impaired passenger who takes the target vehicle. Thus, the environment data can be converted into voice broadcast content that is logically coherent and scene-adapted for visually impaired passengers, significantly improving the information richness and pertinence of voice broadcast. Compared with the fragmented and instructive information in the prior art, the present application enriches the voice broadcast content through the language processing capability of the large language model, so that the visually impaired passengers can construct a complete spatial environment picture based on the voice broadcast content, fill the gap of visual information, effectively improve the sense of journey control, and thus relieve the unstable emotions caused by the lack of information, significantly improve the experience of taking a vehicle for visually impaired passengers.

[0005] Optionally, the method further comprises: determining, according to a driving speed of the target vehicle, broadcast duration information of the voice broadcast information; the driving speed is negatively correlated with the broadcast duration of the voice broadcast information; and the generating, by the large language model, of the voice broadcast information according to the environmental data comprises: generating, by the large language model, the voice broadcast information according to the environmental data and the broadcast duration information. In this way, the broadcast duration information of the voice broadcast information is determined according to the driving speed of the target vehicle; as the driving speed increases, the broadcast duration information of the voice broadcast information is correspondingly shortened, the broadcast duration of the voice broadcast information is dynamically adjusted, the voice broadcast content generated by the large language model is dynamically adapted to the driving speed of the vehicle in terms of time length, it is ensured that the voice broadcast content can be broadcast when the vehicle passes through the relevant environmental area, the problem of lagging of the broadcast information is effectively avoided, the visually impaired passenger can obtain accurate and relevant environmental information in a timely manner, the sense of control over the journey is enhanced, and the overall travel experience is improved.

[0006] Optionally, the environmental data is of at least two types; the determining, according to the driving speed of the target vehicle, of the broadcast duration information of the voice broadcast information comprises: determining total duration information of the voice broadcast information according to the driving speed; and assigning segment duration information of a broadcast segment generated according to each type of environmental data according to the driving speed and a priority corresponding to each type of environmental data; and the generating, by the large language model, of the voice broadcast information according to the environmental data and the broadcast duration information comprises: generating, by the large language model, the voice broadcast information according to the environmental data, the total duration information and the segment duration information. In this way, the voice broadcast content is accurately matched with the driving scene of the vehicle: on the one hand, the total duration is dynamically adjusted according to the driving speed, it is ensured that the broadcast content is completely output within a time window when the vehicle passes through the target environmental area, and the problem of lagging or redundancy of information caused by changes in driving speed is avoided. On the other hand, the segment duration assignment strategy based on the priority reasonably controls the duration of each broadcast segment, a longer segment duration is assigned to high-priority environmental data to support detailed description, and the duration of low-priority data is compressed or the expression is simplified, the broadcast content is appropriately detailed and highlighted, and it is ensured that the visually impaired passenger can receive key environmental information according to the priority.

[0007] Optionally, the method further comprises: determining the broadcast duration information of the voice broadcast information according to the priority corresponding to the environmental data; the higher the priority corresponding to the environmental data is, the shorter the broadcast duration of the voice broadcast information is; and the generating, by the large language model, of the voice broadcast information according to the environmental data comprises: generating, by the large language model, the voice broadcast information according to the environmental data and the broadcast duration information. In this way, the broadcast duration information of the voice broadcast information is determined according to the priority corresponding to the environmental data; the high-priority environmental data shortens the broadcast duration to avoid interference of redundant information, ensuring that the voice broadcast content can timely remind the visually impaired passenger in an emergency; and the low-priority data appropriately prolongs the broadcast duration, enabling the visually impaired passenger to comprehensively obtain environmental information. In this way, the voice broadcast information generated by the large language model is more reasonable in duration allocation, ensuring that important information is not overlooked and giving consideration to the comprehensiveness of information, thereby improving user experience.

[0008] Optionally, the obtaining of the environmental data of the environment in which the target vehicle is located comprises at least one of the following: obtaining interest point data of interest points on a to-be-traveled path of the target vehicle; obtaining information data related to a location of the target vehicle; and obtaining environmental perception data collected by an environmental perception device of the target vehicle. In this way, the large language model integrates multi-dimensional environmental information such as interest point data, information data, and environmental perception data, and the voice broadcast information generated accordingly can comprehensively reflect the environment in which the vehicle is located, can provide detailed and real-time environmental description for the visually impaired passenger, effectively makes up for the lack of visual information, helps the visually impaired passenger to better understand the surrounding environment, alleviates the anxiety of the visually impaired passenger during the ride, and improves the ride experience.

[0009] Optionally, the obtaining of the interest point data of interest points on a to-be-traveled path of the target vehicle comprises: obtaining a to-be-traveled path set of the target vehicle; querying interest point attributes of a first interest point on each to-be-traveled path in the to-be-traveled path set; the interest point attributes at least include a category attribute of the interest point; filtering a second interest point that satisfies a preset priority condition from the first interest point according to a preset mapping relationship; the preset mapping relationship is used to represent a corresponding relationship between the category attribute of the interest point and the priority; and generating the interest point data according to the interest point attributes of the second interest point. In this way, the interest points with a high proportion of low broadcast value are filtered out through the category attribute, and high-priority interest points are preferentially retained, so that the voice broadcast information generated accordingly focuses on interest points with higher broadcast value, thereby preferentially allocating limited voice broadcast time to information of higher value for the visually impaired passenger, helping the visually impaired passenger to quickly capture key environmental information within a limited attention range, and improving the content quality of voice broadcast.

[0010] Optionally, the generating the interest point data according to the interest point attribute of the second interest point comprises: pre-generating an interest point sentence for describing the second interest point according to the interest point attribute of the second interest point by using the large language model; filtering a target interest point from the second interest point according to the position information of the second interest point; the target interest point comprises a second interest point with a distance from the target vehicle on the to-be-traveled path not exceeding a preset distance threshold; and determining the interest point sentence of the target interest point as the interest point data. In this way, the pre-generated interest point sentence can be directly used after the vehicle approaches to generate the voice broadcast information. In this way, the interest point sentence generation process is decoupled from the vehicle movement process by pre-generating the interest point sentence, effectively avoiding the delay problem of dynamically generating the interest point sentence during the vehicle travel process, and helping to improve the real-time performance of the voice broadcast.

[0011] Optionally, before the obtaining the environment data of the environment where the target vehicle is located, the method further comprises: determining whether there is a visually impaired passenger in the target vehicle; and the obtaining the environment data of the environment where the target vehicle is located comprises: in a case where it is determined that there is a visually impaired passenger in the target vehicle, obtaining the environment data of the environment where the target vehicle is located. In this way, in a case where it is determined that there is a visually impaired passenger in the target vehicle, the barrier-free mode is automatically started, and then the environment data of the environment where the target vehicle is located is obtained, and the voice broadcast information is generated according to the environment data by using the large language model. Compared with the mode triggered by the passenger, the inconvenience of manual operation of the visually impaired passenger is avoided, and the delay problem of function activation caused by complex operation is reduced.

[0012] Optionally, the determining whether there is a visually impaired passenger in the target vehicle comprises: obtaining in-vehicle environment perception data in the target vehicle; determining whether there is a personal item of a visually impaired passenger in the target vehicle according to the in-vehicle environment perception data; in a case where there is a personal item of a visually impaired passenger in the target vehicle, determining whether the limb action of the passenger in the target vehicle belongs to the action mode of the visually impaired passenger according to a preset action recognition model and the in-vehicle environment perception data; and determining whether there is a visually impaired passenger in the target vehicle according to whether the limb action of the passenger in the target vehicle belongs to the typical action mode of the visually impaired passenger. In this way, a logical verification is formed by a double screening mechanism: the target range is first narrowed by the personal item detection, and then the atypical behavior interference is excluded by the action mode recognition, effectively solving the limitations of a single detection means, effectively improving the accuracy and reliability of identifying the visually impaired passenger, reducing the misjudgment, and ensuring that the voice broadcast function can timely and accurately serve the visually impaired passenger.

[0013] According to a second aspect of the present application, a voice broadcast device for a visually impaired passenger is provided, comprising:

[0014] An acquisition module is configured to acquire environmental data of an environment in which a target vehicle is located.

[0015] A generation module is configured to generate, by using a large language model, voice broadcast information according to the environmental data, the voice broadcast information being used to introduce the environment in which the target vehicle is located to a visually impaired passenger who is on the target vehicle.

[0016] According to a third aspect of the present application, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; and the processor is configured to execute the method according to any one of the above embodiments.

[0017] According to a fourth aspect of the present application, a computer readable storage medium is provided, the storage medium storing a computer program, the computer program being configured to execute the method according to any one of the above embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 Fig. 1 shows a schematic diagram of an implementation environment provided by an embodiment of the present application.

[0019] Figure 2 Fig. 2 shows a flowchart of a voice broadcast method for a visually impaired passenger provided by an embodiment of the present application.

[0020] Figure 3 Fig. 3 shows a block diagram of a voice broadcast device for a visually impaired passenger provided by an embodiment of the present application.

[0021] Figure 4 Fig. 4 shows a structural block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0022] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0023] Summary of the application

[0024] The inventor has found that visually impaired passengers face a significant information loss problem when taking a vehicle, which leads to a lack of effective perception of the vehicle driving state and the surrounding environment. Due to the lack of visual information, the passengers lose the sense of control of the journey, and further produce an unstable emotion.

[0025] The existing vehicle-mounted navigation system has a voice broadcast function, but its original design is to provide navigation guidance to the driver. The voice prompts output by the system are mostly instructional information such as "enter [town / city name]" "turn left" "change lane". Such information essentially serves the decision-making assistance needs of visually intact drivers, lacks detailed descriptions of the surrounding environment, and does not provide passengers with environmental information such as scenic spots, historical sites, landmarks, and local hot news. The information presentation is non-targeted and fragmented, making it difficult for visually impaired passengers to establish a complete understanding of the travel environment through existing voice broadcasts, and the information acquisition effect is poor.

[0026] To solve the above problems, the embodiment of the present application obtains environmental data of the environment where the target vehicle is located; and then uses a large language model to generate voice broadcast information according to the environmental data to introduce the environment where the target vehicle is located to visually impaired passengers riding the target vehicle.

[0027] In this way, the information such as point of interest data, real-time information dynamics, and environmental perception data contained in the environmental data can be converted into logically coherent and scene-adapted voice broadcast content, significantly improving the information richness and targeting of voice broadcast. Compared with the fragmented instructional information in the prior art, the present solution integrates multiple sources of environmental data into a scene-based description for visually impaired passengers through the semantic integration capability of the large language model, enabling visually impaired passengers to understand the travel environment based on the voice broadcast content, filling the gap in visual information, effectively improving the sense of control during the trip, and thereby alleviating the unstable emotions caused by information loss, significantly improving the travel experience of visually impaired passengers.

[0028] After introducing the basic principles of the present application, various non-limiting embodiments of the present application will be specifically introduced with reference to the accompanying drawings.

[0029] Exemplary system

[0030] Figure 1 An implementation environment provided by an embodiment of the present application is shown. The implementation environment includes a computing device 110 at a target vehicle and a server 120. The computing device 110 can also be connected to the server 120 through a communication network. Optionally, the communication network is a wired network or a wireless network.

[0031] The computing device 110 can be a general-purpose computer or a computer device composed of a dedicated integrated circuit, and the like, and the embodiments of the present application do not limit this. For example, the computing device 110 can be a vehicle-mounted computing device, or a mobile terminal device such as a mobile phone or a tablet computer, or a personal computer (PC). Those skilled in the art can know that the number of the above computing devices 110 can be one or more, and the types thereof can be the same or different. For example, the above computing devices 110 can be one, or the above computing devices 110 can be tens or hundreds, or more. The number and type of the computing device 110 are not limited in the embodiments of the present application.

[0032] The server 120 is a server, or is composed of several servers, or is a virtualization platform, or is a cloud computing service center.

[0033] In some optional embodiments, the computing device 110 at the target vehicle can obtain environmental data of the environment where the target vehicle is located, and upload the environmental data to the server 120. The computing device 110 can obtain the position information and navigation information of the target vehicle, and the environmental perception information collected by the visual sensor, millimeter wave radar and the like.

[0034] The server 120 can be deployed with a large language model to generate voice broadcast information according to the received environmental data, and send the generated voice broadcast information back to the computing device 110 at the target vehicle, so as to introduce the environment to the visually impaired passenger riding in the target vehicle.

[0035] Since the operation of the large language model needs to consume a large amount of computing resources, and the hardware performance of the vehicle-mounted computing device and the mobile terminal is relatively limited, it is generally deployed in the cloud (i.e. at the server 120); in some cases, the large language model can also be directly deployed at the computing device 110, for example, a light-weight large language model optimized by knowledge distillation is deployed at the vehicle-mounted computing device.

[0036] Exemplary method

[0037] Figure 2 FIG. 1 is a flowchart of a voice broadcast method for a visually impaired passenger provided by an embodiment of the present application. Figure 2 The method is executed by the computing device 110 at the target vehicle or the server 120, but the embodiments of the present application are not limited thereto.

[0038] As shown in FIG. 2, the method includes the following contents: Figure 2

[0039] Step S210: Obtain environmental data of the environment where the target vehicle is located.

[0040] ​In the embodiments of the present application, the target vehicle can be various types of passenger cars or commercial vehicles, including but not limited to family cars, buses, taxis, etc., and can also be an intelligent car with automatic driving function, without specific limitation.

[0041] In the embodiments of the present application, the environment data can be all obtainable information related to the environment where the target vehicle is located, without specific limitation on its source and type. Specifically, the environment data can include point of interest (POI) data of POIs near the target vehicle, which covers various types of location information related to passenger travel, such as restaurants, gas stations, shopping malls, hospitals, scenic spots, etc.; can also include network data related to the location where the target vehicle is located, such as local weather conditions, news information, traffic conditions, etc. dynamic information; in addition, it can also be based on real-time environment perception data collected by the environment perception device of the target vehicle, such as image data captured by the vehicle-mounted camera, radar data detected by the vehicle-mounted radar.

[0042] Step S220: generating voice broadcast information according to the environment data by using a large language model; the voice broadcast information is used to introduce the environment where the target vehicle is located to the visually impaired passenger riding in the target vehicle.

[0043] In the embodiments of the present application, the large language model is used to process and analyze the input environment data to generate information suitable for voice broadcast, and the large language model can include a large language model, a multi-modal model, etc.

[0044] In the embodiments of the present application, the prompt words input to the large language model can include instructions (Instructions), roles (Roles), contexts (Contexts), input data (Input Data), output indicators (Output Indicators), etc. Among them, the instructions are used to specify the specific generation task that the model needs to perform, and can also contain role setting instructions and broadcast style limitations, etc.; the context is used to limit the associated scene of the environment data and the vehicle state, and the context is used to provide background information to narrow the understanding range of the model, such as "gender, age, etc. of the visually impaired passenger"; the input data is used to provide the environment data of the environment where the target vehicle is located, and the output indicator can be used to constrain the text structure and broadcast duration of the voice broadcast. The output indicator is used to standardize the output format, including the broadcast duration, the format constraint, the style limitation, etc.

[0045] In the embodiments of the present application, the voice broadcast information is used to introduce the environment in which the target vehicle is located to the visually impaired passenger riding in the target vehicle. The specific form can include text data or audio data. The text data can only include pure text content, and can also include a Speech Synthesis Markup Language (SSML) to facilitate subsequent voice synthesis and broadcast.

[0046] In some cases, the voice broadcast information can also be used to introduce the vehicle state and travel information to the visually impaired passenger, and provide safety warnings such as sudden braking, sharp turns, etc. The method further comprises: obtaining vehicle information of the target vehicle, which can include vehicle state information and vehicle travel information; using the large language model to generate voice broadcast information according to the vehicle information. The vehicle state information can include driving speed, steering wheel angle, accelerator pedal state, brake pedal state, longitudinal acceleration, lateral acceleration, and other vehicle state information; the vehicle travel information can include the path to be traveled, the remaining distance, the remaining time, etc.

[0047] In the embodiments of the present application, environment data of the environment in which the target vehicle is located is obtained; and then the environment data is input into the large language model to generate voice broadcast information according to the environment data, so as to introduce the environment in which the target vehicle is located to the visually impaired passenger riding in the target vehicle. In this way, the environment data can be converted into voice broadcast content that is logical, coherent, and scene-adapted for the visually impaired passenger, significantly improving the information richness and pertinence of voice broadcast. Compared with the fragmented and instructive information in the prior art, the present solution enriches the voice broadcast content through the language processing capability of the large language model, so that the visually impaired passenger can construct a complete spatial environment picture based on the voice broadcast content, fill the gap in visual information, effectively improve the travel control feeling, provide more temperature and practicality for the visually impaired group, and then alleviate the unstable emotions caused by information loss, significantly improve the travel experience of the visually impaired passenger.

[0048] Based on the method in Figure 2 The embodiments of the present specification also provide some specific implementation schemes of the method, which are described below.

[0049] In actual application, the answer generated by the large language model is usually lengthy, and when the vehicle is traveling at a high speed, the voice broadcast often cannot be completed before the corresponding environment area is exceeded, resulting in information lag and insufficient real-time performance.

[0050] Based on this, the embodiments of the present specification also provide some specific implementation schemes to improve the real-time performance of voice broadcast, which are implemented as follows:

[0051] Optionally, the method further comprises:

[0052] determine the broadcast duration information of the voice broadcast information according to the driving speed of the target vehicle; the driving speed is negatively correlated with the broadcast duration of the voice broadcast information;

[0053] The voice broadcast information is generated according to the environment data by using the large language model, and the voice broadcast information includes:

[0054] The voice broadcast information is generated according to the environment data and the broadcast duration information by using the large language model.

[0055] In the embodiments of the present application, the broadcast duration information is used to constrain the time length or the corresponding text word number of the voice broadcast content. The broadcast duration information can be a text character length threshold and an upper limit of audio stream duration.

[0056] In the embodiments of the present application, the driving speed is negatively correlated with the broadcast duration of the voice broadcast information, that is, the faster the driving speed, the shorter the broadcast duration. In some cases, the driving speed can be inversely proportional to the broadcast duration of the voice broadcast information.

[0057] For example, when the driving speed is greater than 60 km / h, the broadcast duration is limited to no more than 15 seconds to ensure the rapid transmission of key information in a high-speed scenario; when the driving speed is between 30-60 km / h, the broadcast duration can be controlled between 20-30 seconds to balance the information integrity and real-time performance; when the driving speed is less than 30 km / h, the broadcast duration is allowed to be more than 40 seconds to support more detailed environment description.

[0058] In the embodiments of the present application, the broadcast duration information of the voice broadcast information is determined according to the driving speed of the target vehicle; as the driving speed increases, the broadcast duration information of the voice broadcast information is shortened. Thus, the broadcast duration of the voice broadcast information is dynamically adjusted, so that the voice broadcast content generated by the large language model is dynamically adapted to the driving speed of the vehicle in terms of time length, ensuring that the voice broadcast content can be broadcast when the vehicle passes through the relevant environment area, effectively avoiding the problem of lagging broadcast information, improving the real-time performance of the voice broadcast information, enabling the visually impaired passenger to obtain the relevant environment information in time and accurately, and thus enhancing the sense of control over the journey and improving the overall travel experience.

[0059] In actual application, the sources and types of the environment data are diverse, such as point of interest data obtained from a map service, network connection data related to the location of the target vehicle obtained through vehicle network technology, and environment perception data obtained from a vehicle-mounted perception device. The information density and importance of these environment data are often different. If all environment data are indiscriminately converted into voice broadcast information, it is easy to cause the broadcast content to lack focus and be lengthy and tedious.

[0060] Based on this, the embodiments of the present specification also provide some specific implementations to reasonably allocate the time length of the broadcast segments generated according to various environmental data, which are implemented as follows:

[0061] Optionally, the types of the environmental data are at least two;

[0062] The determination of the broadcast time length information of the voice broadcast information according to the driving speed of the target vehicle comprises:

[0063] The determination of the total time length information of the voice broadcast information according to the driving speed;

[0064] The allocation of the segment time length information of the broadcast segments generated according to various environmental data according to the driving speed and the priority of various environmental data;

[0065] The generation of the voice broadcast information according to the environmental data and the broadcast time length information by using the large language model comprises:

[0066] The generation of the voice broadcast information according to the environmental data, the total time length information and the segment time length information by using the large language model.

[0067] In the embodiments of the present application, the total time length information is used to determine the overall time length of voice broadcast, ensuring that the broadcast content can be completed within the appropriate time.

[0068] In the embodiments of the present application, the priority is used to represent the broadcast value of the corresponding environmental data, which can be set according to the dimensions of acquisition speed, information density, importance, etc. The priorities of environmental data from different sources can be different, and the priorities of environmental data from the same source can also be different. Taking the point of interest data as an example, there are many types of points of interest and differences in the priorities of different types.

[0069] Under normal conditions, the point of interest data obtained from the map service has a fast acquisition speed and a high information density, and the priority can be set to be high; the hot news, weather forecast and other network data obtained through vehicle network technology have a fast acquisition speed, but the information density is poor and generally not very important, and the priority can be set to be general; the environmental perception data obtained from the vehicle-mounted perception device cannot be obtained in advance and the processing process is time-consuming, and is generally used to generate detailed explanation content, which has a slow acquisition speed and weak importance, and the priority can be set to be the lowest.

[0070] It should be noted that the above priority division is not absolute: collision risk, intense driving (such as sudden acceleration / sudden braking), abnormal weather and other emergency related data, although these information belong to networked data or environmental perception data, but because it is directly related to the safety of the vehicle, the priority can be divided into the highest level, and the interrupt mechanism can be triggered to insert the emergency warning content in real time; The dynamic data such as sudden changes in people flow and temporary activities that may affect the safety of the trip can also be set to a higher priority to ensure that key information is delivered in priority when the environment changes.

[0071] In the embodiment of the present application, the broadcast segment can refer to a voice broadcast part generated according to various types of environmental data, such as a scenic spot description segment generated based on point of interest data, a news segment generated based on networked data.

[0072] In the embodiment of the present application, the segment duration information refers to the time length of the broadcast segment, and the time length allocated to each broadcast segment is determined according to the driving speed and the priority of the corresponding environmental data.

[0073] It should be noted that the length of the segment duration information corresponding to some environmental data can be set to zero, that is, it is allowed to delete environmental data with lower priority. The segment duration information of the broadcast segment generated according to various environmental data can be allocated according to the driving speed and the priority of various environmental data, which can include: selecting target environmental data from environmental data according to the driving speed and the priority of various environmental data; according to the driving speed and the priority of the target environmental data, the segment duration information of the broadcast segment generated according to the target environmental data is allocated.

[0074] In the embodiment of the present application, when the segment duration of the information segment generated based on the information data is greater than zero, the information data is obtained. When the segment duration of the environment description segment generated based on the environmental perception data is greater than zero, the environmental perception data is obtained.

[0075] In the embodiment of the present application, the voice broadcast information can be generated by carrying the broadcast duration information and the segment duration information in the prompt word according to the environmental data, the total duration information and the segment duration information using the large language model.

[0076] In the embodiments of the present application, the total duration information of the voice broadcast information is determined according to the driving speed; the segment duration information of the broadcast segment generated according to various environmental data is allocated according to the priority corresponding to the driving speed and the various environmental data; and then the voice broadcast information is generated by using the large language model according to the environmental data, the total duration information and the segment duration information. In this way, the voice broadcast content is accurately matched with the vehicle driving scene: on the one hand, the total duration is dynamically adjusted according to the driving speed, ensuring that the broadcast content is completely output within the time window when the vehicle passes through the target environmental area, avoiding information lag or redundancy caused by changes in driving speed. On the other hand, the segment duration allocation strategy based on priority reasonably controls the duration of each broadcast segment, allocates longer segment duration to high-priority environmental data to support detailed description, and compresses the duration or simplifies the expression of low-priority data, so as to achieve appropriate detail and highlight of the broadcast content, and ensure that the visually impaired passengers can receive key environmental information according to priority.

[0077] Optionally, the method further comprises:

[0078] determining the broadcast duration information of the voice broadcast information according to the priority corresponding to the environmental data; the higher the priority corresponding to the environmental data is, the shorter the broadcast duration of the voice broadcast information is;

[0079] generating the voice broadcast information according to the environmental data by using the large language model, comprising:

[0080] generating the voice broadcast information according to the environmental data and the broadcast duration information by using the large language model.

[0081] In the embodiments of the present application, the priority is used to represent the broadcast value of the corresponding environmental data, which can be set according to the importance of the environmental data.

[0082] In the embodiments of the present application, the environmental data related to emergency events such as collision risk, intense driving (such as sudden acceleration / sudden braking), and abnormal weather has the highest priority (for example, emergency level), and can trigger a broadcast interruption mechanism to insert real-time emergency warning content; the dynamic data such as sudden flow of people and temporary activities that may affect the safety of the journey can also be set to a higher priority, higher than the data used for environmental introduction such as static knowledge layer, dynamic data layer and real-time environment layer, to ensure that key information is transmitted first when the environment changes.

[0083] In the embodiments of the present application, the broadcast duration information of the voice broadcast information is determined according to the priority corresponding to the environmental data; due to the importance of high-priority environmental data, shortening the broadcast duration can avoid redundant information interference and ensure that the voice broadcast content can timely remind the visually impaired passengers in an emergency; the broadcast duration of low-priority data is appropriately extended, so that the visually impaired passengers can comprehensively obtain environmental information. In this way, the voice broadcast information generated by the large language model is more reasonable in duration allocation, which can ensure that important information is not ignored and also take into account the comprehensiveness of the information, thereby improving the user experience.

[0084] Optionally, the environmental data of the environment where the target vehicle is located is obtained, including at least one of the following:

[0085] The interest point data of the interest point on the to-be-traveled path of the target vehicle is obtained.

[0086] The information data related to the location of the target vehicle is obtained.

[0087] The environmental perception data collected by the environmental perception device based on the target vehicle is obtained.

[0088] In the embodiments of the present application, the interest point (POI) can refer to a place in geographical space with a specific function or meaning; further, the interest point can refer to a geographical entity object with voice broadcast value. The interest point can include any category, or can be a designated category of place or facility. Preferably, the interest point can include scenic names, historical relics, various scenic spots, landmark buildings, city landmark sculptures, cultural squares, historical blocks, and other geographical entity objects with cultural value, which are not limited in particular.

[0089] In the embodiments of the present application, the interest point data can be data describing the interest point, which can include interest point attributes directly obtained from a map service, such as name, address, category, and picture, and other voice convertible attributes.

[0090] In the embodiments of the present application, the interest point data can also be a natural language sentence generated according to the interest point attributes, which integrates the name, address, and function attributes into coherent introduction content by using a large language model.

[0091] In the embodiments of the present application, the information data is network data related to the position of the vehicle obtained through the vehicle network function, including but not limited to weather information, news information, traffic conditions, etc. Specifically, the weather information includes, for example, real-time temperature, humidity, wind speed, precipitation probability, extreme weather warning, etc.; the news information includes, for example, hot events around the current position of the vehicle, public safety notifications, livelihood service information, etc.; and the traffic conditions include, for example, road congestion level, accident location, temporary traffic control measures, lane closure information, etc. real-time dynamic data related to the driving route of the vehicle.

[0092] In the embodiments of the present application, the environment perception device is used to perceive the environment around the vehicle in real time, mainly including sensors such as cameras, ultrasonic radars, laser radars, etc.

[0093] In the embodiments of the present application, the environment perception data is the data collected by the environment perception device, including image or video data collected by the camera, distance data collected by the ultrasonic radar, point cloud data collected by the laser radar, etc. Generally, the time interval or frequency of the data frame of the environment perception data input to the large language model can be determined according to the driving speed.

[0094] In the embodiments of the present application, the information data related to the position of the target vehicle can be obtained, including: obtaining the information data when the speed is not more than a first speed threshold, or when the segment duration of the information segment generated based on the information data is greater than zero.

[0095] In the embodiments of the present application, the environment perception data collected by the environment perception device based on the target vehicle can be obtained, including: obtaining the environment perception data when the speed is not more than a second speed threshold, or when the segment duration of the environment description segment generated based on the environment perception data is greater than zero.

[0096] In the embodiments of the present application, the large language model is used to integrate multi-dimensional environmental information such as point of interest data, information data and environment perception data, and the voice broadcast information generated based on this can comprehensively reflect the environment in which the vehicle is located, can provide detailed and real-time environment description for visually impaired passengers, effectively make up for the lack of visual information, help visually impaired passengers better understand the surrounding environment, alleviate their anxiety during the ride, and improve the ride experience.

[0097] Optionally, the point of interest data of the point of interest on the to-be-traveled path of the target vehicle is obtained, including:

[0098] A set of to-be-traveled paths of the target vehicle is obtained.

[0099] query a point of interest attribute of a first point of interest on each to-be-traveled path in the set of to-be-traveled paths; the point of interest attribute at least includes a category attribute of the point of interest;

[0100] obtain a second point of interest that meets a preset priority condition from the first point of interest according to a preset mapping relationship; the preset mapping relationship is used to represent a corresponding relationship between the category attribute of the point of interest and the priority;

[0101] obtain the point of interest data according to the point of interest attribute of the second point of interest.

[0102] In the embodiment of the application, the set of to-be-traveled paths can include a to-be-traveled path of the target vehicle; specifically, the set of to-be-traveled paths can be a to-be-traveled path sequence of navigation planning, can also be a road currently traveled by the target vehicle, and can further include a possible path to be traveled, for example, a path to be traveled at a next intersection of a current path.

[0103] In the embodiment of the application, the first point of interest can be a point of interest on a to-be-traveled path.

[0104] In the embodiment of the application, the preset mapping relationship is used to represent a corresponding relationship between the category attribute of the point of interest and the priority. For example, scenic spots such as scenic spots and historical sites are mapped to high priority, landmark buildings such as shopping malls, office buildings, and schools are mapped to medium priority, and daily life facilities such as restaurants, convenience stores, and newspaper stands are mapped to the lowest priority.

[0105] In the embodiment of the application, the preset priority condition is used to filter points of interest that meet the priority requirement, and can be set as a category priority threshold matching condition or a dynamically adjusted sorting condition.

[0106] In the embodiment of the application, the second point of interest can include a point of interest that meets a preset priority condition filtered from the first point of interest; specifically, the second point of interest can be a first point of interest with a priority higher than a preset priority, or a preset number of first points of interest with the highest priority.

[0107] In the embodiment of the application, obtaining the point of interest data according to the point of interest attribute of the second point of interest can include: determining a second point of interest and an attribute related to voice broadcast as the point of interest data; or using a large language model to process the point of interest attribute of the second point of interest to obtain a natural language sentence for describing the second point of interest.

[0108] In the embodiment of the present application, the interest point attributes of the first interest points on each to-be-traveled path in the set of to-be-traveled paths are queried; then, according to a preset mapping relationship, a second interest point satisfying a preset priority condition is filtered from the first interest points; the preset mapping relationship is used to represent a corresponding relationship between a category attribute of the interest point and a priority; and the interest point data is obtained according to the interest point attributes of the second interest point. In this way, by filtering out interest points with a high proportion of low broadcast value according to the category attribute, high-priority interest points are preferentially retained, and the voice broadcast information generated accordingly focuses on interest points with higher broadcast value, so that the limited voice broadcast time is preferentially allocated to information of higher value for the visually impaired passengers, helping the visually impaired passengers to quickly capture key environmental information within a limited attention range and improving the content quality of voice broadcast.

[0109] Optionally, the interest point attributes at least include a category attribute of the interest point.

[0110] The interest point data is obtained according to the interest point attributes of the second interest point, including:

[0111] The interest point data is obtained according to the interest point attributes of the second interest point, including:

[0112] According to the position information of the second interest point, a target interest point is filtered from the second interest point; the target interest point includes a second interest point on the to-be-traveled path and between the target vehicle, and the distance does not exceed a preset distance threshold.

[0113] The interest point data is obtained according to the interest point attributes of the second interest point, including:

[0114] In the embodiment of the present application, the interest point data can be a natural language sentence for describing the second interest point.

[0115] In the embodiment of the present application, the second interest point can include all interest points on the to-be-traveled paths of the target vehicle that satisfy the preset priority condition.

[0116] In the embodiment of the present application, the target interest point can include a second interest point on the to-be-traveled path and between the target vehicle, and the distance does not exceed a preset distance threshold, that is, an interest point that the target vehicle will pass through soon.

[0117] Generally, the generation speed of a large language model is slow, which results in a long time consumption and poor real-time performance in generating an interest point sentence according to interest point attributes when the target vehicle is about to pass through an interest point. The interest point attributes are relatively stable information and do not change within a short period of time, and can be regarded as static information.

[0118] In the embodiment of the present application, after the to-be-traveled path is determined, the large language model is used to pre-generate a point-of-interest sentence for describing the second point of interest according to the point-of-interest attribute of the second point of interest. The pre-generated point-of-interest sentence can be directly used to generate voice broadcast information when the vehicle approaches. In this way, by pre-generating the point-of-interest sentence, the point-of-interest sentence generation process is decoupled from the vehicle movement process, effectively avoiding the delay problem of dynamically generating the point-of-interest sentence during vehicle travel, and helping to improve the real-time performance of voice broadcast.

[0119] Optionally, before the environment data of the environment where the target vehicle is located is acquired, the method further includes:

[0120] determining whether there is a visually impaired passenger in the target vehicle;

[0121] acquiring the environment data of the environment where the target vehicle is located includes:

[0122] In the case where it is determined that there is a visually impaired passenger in the target vehicle, the environment data of the environment where the target vehicle is located is acquired.

[0123] In the embodiment of the present application, it can be determined whether there is a visually impaired passenger in the target vehicle after the door of the target vehicle is closed, or when the target vehicle is about to move or start.

[0124] In the embodiment of the present application, the determination of whether there is a visually impaired passenger in the target vehicle can have various implementation manners, for example:

[0125] According to the in-vehicle environment perception data in the target vehicle, it is determined whether there is a visually impaired passenger in the target vehicle. The in-vehicle environment perception data can be obtained by an occupant monitoring system (OMS) and can include image data collected by a camera and radar data collected by a radar.

[0126] According to the in-vehicle sound data, it is determined whether the speech of the passenger in the target vehicle contains a visually impaired help-seeking intention. According to whether the speech of the passenger in the target vehicle contains a visually impaired help-seeking intention, it is determined whether there is a visually impaired passenger in the target vehicle.

[0127] By matching a visually impaired passenger feature library through facial recognition technology, it is determined whether there is a registered visually impaired passenger in the target vehicle. The visually impaired passenger feature library is used to store the facial features of the registered visually impaired passengers.

[0128] According to the wireless connection state between the mobile terminal device carried by the visually impaired passenger and the vehicle-mounted system, it is determined whether there is a registered visually impaired passenger in the target vehicle.

[0129] According to whether the electronic tag of the pre-stored visually impaired user is detected, it is determined whether the registered visually impaired passenger exists in the target vehicle.

[0130] In some cases, the target vehicle can receive instructions for triggering or canceling voice broadcasting through a human machine interface (HMI). Specifically, the user can trigger or cancel voice broadcasting by clicking a button, issuing a specific voice instruction, making a preset gesture, etc.

[0131] In some cases, when it is determined that the visually impaired passenger exists in the target vehicle, a voice prompt can be given to ask whether to start the barrier-free mode. When the passenger confirms to start the barrier-free mode, the environment data of the environment where the target vehicle is located is obtained, and a large language model is used to generate voice broadcast information according to the environment data.

[0132] In the embodiments of the present application, when it is determined that the visually impaired passenger exists in the target vehicle, the barrier-free mode is automatically started, and then the environment data of the environment where the target vehicle is located is obtained, and a large language model is used to generate voice broadcast information according to the environment data. Compared with the mode triggered by the passenger, the inconvenience of manual operation of the visually impaired passenger is avoided, and the function activation delay problem caused by complex operation is reduced.

[0133] Optionally, the determining whether the visually impaired passenger exists in the target vehicle comprises:

[0134] Obtaining in-vehicle environment perception data in the target vehicle;

[0135] According to the in-vehicle environment perception data, it is determined whether the visually impaired passenger's personal belongings exist in the target vehicle;

[0136] When the visually impaired passenger's personal belongings exist in the target vehicle, according to a preset action recognition model and in-vehicle environment perception data, it is determined whether the passenger's limb action in the target vehicle belongs to the action mode of the visually impaired passenger;

[0137] According to whether the passenger's limb action in the target vehicle belongs to the typical action mode of the visually impaired passenger, it is determined whether the visually impaired passenger exists in the target vehicle.

[0138] In the embodiments of the present application, the in-vehicle environment perception data can include at least one of in-vehicle image data and in-vehicle radar data in the target vehicle. The in-vehicle environment perception data can be obtained through an occupant monitoring system (OMS), which is not limited in this regard.

[0139] In the embodiments of the present application, the personal belongings of the visually impaired passenger can include items that the visually impaired passenger generally carries and normal passengers generally do not carry, such as a guide cane, a guide dog, glasses, etc., without specific limitation.

[0140] In the embodiments of the present application, in the case that the in-vehicle environment perception data is image data, the determining whether the target vehicle has a visually impaired passenger according to the in-vehicle environment perception data can include processing the image data collected by the camera using a target detection model to identify whether the image has a visually impaired passenger's personal belongings. The target detection model can be an algorithm such as YOLO (You Only Look Once), SSD (Single Shot Multi Box Detector), etc.

[0141] In the embodiments of the present application, in the case that the in-vehicle environment perception data is radar data, the determining whether the target vehicle has a visually impaired passenger according to the in-vehicle environment perception data can include analyzing the point cloud data or echo signal collected by the radar to identify whether there is a target object that matches the structural features of the visually impaired passenger's personal belongings, and determining whether there is a visually impaired passenger's personal belongings based on the feature matching result.

[0142] In the embodiments of the present application, the preset action recognition model is used to identify passenger limb action features and match visually impaired passenger action patterns; for example, an LSTM (Long Short-Term Memory) network, a 3D pose estimation (3D Pose Estimation) model, etc.

[0143] In the embodiments of the present application, the action pattern of the visually impaired passenger is a unique action pattern of the visually impaired passenger, for example, whether the hands are continuously moving irregularly in a non-interaction area, whether the arms are kept stretched to detect obstacles, whether the head is frequently turned to assist auditory positioning, etc.

[0144] For example, in the case that the in-vehicle environment perception data is image data, the human body posture key points in the image sequence are extracted, and the preset action recognition model is input to analyze whether the action trajectory conforms to the visually impaired mode.

[0145] In the embodiments of the present application, if the limb action of the passenger in the target vehicle belongs to the typical action pattern of the visually impaired passenger, it is determined that the target vehicle has a visually impaired passenger. If the limb action of the passenger in the target vehicle does not belong to the typical action pattern of the visually impaired passenger, it is determined that the target vehicle does not have a visually impaired passenger.

[0146] In the embodiment of the present application, the presence of the personal belongings of the visually impaired passenger in the target vehicle is determined to perform preliminary screening; in the case that the personal belongings of the visually impaired passenger are present in the target vehicle, the limb movement of the passenger in the target vehicle is determined to belong to the movement mode of the visually impaired passenger according to the preset action recognition model and the in-vehicle environment perception data, so as to perform accurate screening.

[0147] In practical applications, the single use of the personal belongings detection may lead to misjudgment due to the similar objects such as crutches and pet dogs carried by ordinary passengers, and the pure reliance on the action recognition may also mistakenly match the occasional action of the ordinary passenger as the visually impaired action mode. In the embodiment of the present application, the target range is first narrowed by the personal belongings detection, and then the non-typical behavior interference is excluded by the action mode recognition, which effectively solves the limitations of the single detection means, effectively improves the accuracy and reliability of the identification of the visually impaired passenger, reduces the misjudgment, and ensures that the voice broadcast function can provide services for the visually impaired passenger in time and accurately.

[0148] Optionally, the determining whether the target vehicle contains a visually impaired passenger comprises:

[0149] Obtaining in-vehicle environment perception data in the target vehicle;

[0150] In the case that the target vehicle contains the personal belongings of the visually impaired passenger of the preset category, obtaining in-vehicle sound data in the target vehicle;

[0151] Determining whether the speech of the passenger in the target vehicle contains a visually impaired help-seeking intention according to the in-vehicle sound data;

[0152] Determining whether the target vehicle contains a visually impaired passenger according to whether the speech of the passenger in the target vehicle contains a visually impaired help-seeking intention.

[0153] In the embodiment of the present application, the determining whether the speech of the passenger in the target vehicle contains a visually impaired help-seeking intention according to the in-vehicle sound data can include: performing semantic analysis on the sound data collected by the in-vehicle microphone by using a natural language processing model, and identifying whether the speech of the passenger contains a feature expression related to the visually impaired, such as the keywords of "can't see", "help me", "navigation sound", "brighten", "where", "how to go", etc.

[0154] Based on the same technical concept, another voice broadcast method for a visually impaired passenger is also provided in the embodiment of the present application, which can include:

[0155] After the door of the target vehicle is closed, or in the case that the target vehicle is about to move or start, in-vehicle environment sensing data in the target vehicle is acquired; whether there is a personal item of a visually impaired passenger in the target vehicle is determined according to the in-vehicle environment sensing data; in the case that there is a personal item of a visually impaired passenger in the target vehicle, whether the limb movement of the passenger in the target vehicle belongs to the movement mode of a visually impaired passenger is determined according to a preset action recognition model and the in-vehicle environment sensing data; whether there is a visually impaired passenger in the target vehicle is determined according to whether the limb movement of the passenger in the target vehicle belongs to the typical movement mode of a visually impaired passenger. In some cases, whether the speech of the passenger in the target vehicle contains a visually impaired help-seeking intention can also be determined according to in-vehicle sound data; whether there is a visually impaired passenger in the target vehicle is determined according to whether the speech of the passenger in the target vehicle contains a visually impaired help-seeking intention.

[0156] In the case that it is determined that there is a visually impaired passenger in the target vehicle, it can be prompted by voice whether to start the barrier-free mode, and in the case that the passenger confirms to start, environment data of the environment where the target vehicle is located is acquired.

[0157] According to the driving speed, total time length information of the voice broadcast information is determined; according to the driving speed and the priority corresponding to each of the environment data, segment time length information of a broadcast segment generated according to each of the environment data is allocated.

[0158] Specifically, according to the driving speed, total time length information of the voice broadcast information is determined; according to the driving speed and the priority corresponding to each of the environment data, segment time length information of a broadcast segment generated according to each of the environment data is allocated. For example, when the driving speed is greater than 60 km / h, the broadcast time length is limited to not more than 15 seconds, mainly broadcasting information of the static knowledge layer and a small amount of information of the simplified dynamic data layer, and the voice broadcast information can be “1 kilometer ahead is a bridge, built in 1968”; when the vehicle speed is between 30-60 km / h, the broadcast time length is controlled to be 20-30 seconds, containing information of the static knowledge layer and the dynamic data layer, and the voice broadcast information can be “current traffic is heavy, there are 23 brick-wood structure small buildings on the right side 200 meters away”; and when the vehicle speed is less than 30 km / h, the broadcast time length can be extended to more than 40 seconds, covering detailed information of the static knowledge layer, the dynamic data layer and the real-time environment layer, and the voice broadcast information can include: expanding the description of the current real-time scene, building details, etc.

[0159] The static knowledge layer is used to provide basic information about the interest points on the path to be traveled by the target vehicle, such as the name, location, historical background, etc. of the interest points, which are usually obtained from a map service or a pre-stored database, for example, “The arch was built in the Wanli period, and the original plaque has been lost.” The dynamic data layer involves real-time or near-real-time information data related to the location of the target vehicle, such as weather conditions, traffic news, surrounding activities, etc. These data are usually obtained from cloud services through vehicle network functions, for example, “Social media shows that the arch base stone carving was discovered recently.” The real-time environment layer provides real-time information about the environment around the target vehicle based on data collected by the vehicle's environmental perception devices (such as cameras, radars, etc.), such as the dynamic changes of obstacles, pedestrians, vehicles, etc. in front of the vehicle, for example, “Rainwater forms a reflective path in the stone gully.”

[0160] The segment duration information of the broadcast segment generated according to various environmental data is used to determine whether to obtain information from the dynamic data layer or the real-time environment layer.

[0161] The environmental data of the environment where the target vehicle is located is obtained, including at least one of the following:

[0162] Interest point data of interest points on the path to be traveled by the target vehicle is obtained.

[0163] Information data related to the location of the target vehicle is obtained.

[0164] Environmental perception data collected based on the environmental perception devices of the target vehicle is obtained.

[0165] The interest point data of the interest points on the path to be traveled by the target vehicle is obtained by: obtaining a set of paths to be traveled by the target vehicle; querying the interest point attributes of the first interest points on each path to be traveled in the set of paths to be traveled; the interest point attributes at least include the category attribute of the interest point; according to a preset mapping relationship, a second interest point satisfying a preset priority condition is selected from the first interest point; the preset mapping relationship is used to represent the correspondence between the category attribute of the interest point and the priority; using the large language model, an interest point sentence describing the second interest point is generated in advance according to the interest point attribute of the second interest point; according to the location information of the second interest point, a target interest point is selected from the second interest point; the target interest point includes the second interest point whose distance from the target vehicle on the path to be traveled does not exceed a preset distance threshold; and the interest point sentence of the target interest point is determined as the interest point data.

[0166] The information data related to the location of the target vehicle is acquired, including: in a case where the speed does not exceed a first speed threshold, or in a case where a segment time length of an information segment generated based on the information data is greater than zero, the information data is acquired in advance.

[0167] The environment perception data collected by the environment perception device of the target vehicle is acquired, including: in a case where the speed does not exceed a second speed threshold, or in a case where a segment time length of an environment description segment generated based on the environment perception data is greater than zero, the environment perception data is acquired.

[0168] The voice broadcast information is generated by using the large language model according to the environment data, the total time length information and the segment time length information. The voice broadcast information is played by using a voice synthesis module of a vehicle-mounted computing device or a mobile terminal device.

[0169] Exemplary apparatus

[0170] The device embodiment of the present application can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0171] Figure 3 A block diagram of a voice broadcast device for a visually impaired passenger provided by an embodiment of the present application is shown. As shown in the figure, Figure 3 The device 300 includes:

[0172] The acquisition module 310 is configured to acquire environment data of an environment in which a target vehicle is located.

[0173] The generation module 320 is configured to generate voice broadcast information according to the environment data by using a large language model; the voice broadcast information is used to introduce an environment in which a target vehicle is located to a visually impaired passenger who rides the target vehicle.

[0174] Optionally, the device 300 further includes a time allocation module configured to determine broadcast time length information of the voice broadcast information according to a driving speed of the target vehicle; the driving speed is negatively related to the broadcast time length of the voice broadcast information.

[0175] The generation module 320 is configured to generate voice broadcast information according to the environment data and the broadcast time length information by using the large language model.

[0176] Optionally, the types of the environment data are at least two.

[0177] The time allocation module is configured to: determine total time length information of the voice broadcast information according to the driving speed; and allocate segment time length information of a broadcast segment generated according to various environment data according to the driving speed and priorities corresponding to the various environment data.

[0178] The generation module 320 is configured to generate voice broadcast information by using the large language model and according to the environment data, the total time length information, and the segment time length information.

[0179] Optionally, the acquisition module 310 includes at least one of the following:

[0180] An interest point unit is configured to acquire interest point data of an interest point on a to-be-traveled path of the target vehicle.

[0181] An information unit is configured to acquire information data related to a location of the target vehicle.

[0182] A perception unit is configured to acquire environment perception data collected by an environment perception device of the target vehicle.

[0183] Optionally, the interest point unit is configured to:

[0184] Acquire a set of to-be-traveled paths of the target vehicle.

[0185] Query an interest point attribute of a first interest point on each to-be-traveled path in the set of to-be-traveled paths; the interest point attribute at least includes a category attribute of the interest point.

[0186] According to a preset mapping relationship, a second interest point that meets a preset priority condition is filtered from the first interest point; the preset mapping relationship is used to represent a corresponding relationship between the category attribute of the interest point and the priority.

[0187] Generate the interest point data according to the interest point attribute of the second interest point.

[0188] Optionally, the generating the interest point data according to the interest point attribute of the second interest point includes:

[0189] Pre-generate an interest point sentence for describing the second interest point according to the interest point attribute of the second interest point by using the large language model;

[0190] According to location information of the second interest point, a target interest point is filtered from the second interest point; the target interest point includes a second interest point on the to-be-traveled path and having a distance from the target vehicle not exceeding a preset distance threshold.

[0191] Determine the interest point sentence of the target interest point as the interest point data.

[0192] Optionally, the apparatus 300 further comprises a judging module configured to judge whether there is a visually impaired passenger in the target vehicle.

[0193] The obtaining module 310 is configured to obtain environmental data of an environment in which the target vehicle is located, in a case where it is determined that there is a visually impaired passenger in the target vehicle.

[0194] Optionally, the judging module is configured to:

[0195] obtain in-vehicle environmental perception data in the target vehicle;

[0196] judge, according to the in-vehicle environmental perception data, whether there is a personal item of a visually impaired passenger in the target vehicle;

[0197] in a case where there is a personal item of a visually impaired passenger in the target vehicle, judge, according to a preset action recognition model and the in-vehicle environmental perception data, whether a limb action of a passenger in the target vehicle belongs to a motion pattern of a visually impaired passenger;

[0198] judge, according to whether the limb action of the passenger in the target vehicle belongs to the typical motion pattern of a visually impaired passenger, whether there is a visually impaired passenger in the target vehicle.

[0199] Exemplary electronic device

[0200] Hereinafter, an electronic device according to an embodiment of the present application will be described with reference to the accompanying drawings. Figure 4 FIG. 1 illustrates a block diagram of an electronic device according to an embodiment of the present application. Figure 4 FIG. 1 illustrates a block diagram of an electronic device according to an embodiment of the present application.

[0201] As shown in FIG. 4, the electronic device 400 includes one or more processors 410 and a memory 420. Figure 4 The processor 410 can have other forms of processing units of data processing capability and / or instruction execution capability, and can control other components in the electronic device 400 to perform desired functions.

[0202] The processor 410 can have other forms of processing units of data processing capability and / or instruction execution capability, and can control other components in the electronic device 400 to perform desired functions.

[0203] Specifically, the processor 410 can be a general purpose processor, such as a general purpose central processing unit (CPU), a microprocessor, or the like, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of program instructions for the present solution. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components. The processor 410 can also include a main processor, and can also include a baseband chip, a modem, and the like.

[0204] The memory 420 can include one or more computer program products, which can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM), cache memory, and / or the like. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, and the like. One or more computer program instructions can be stored on the computer-readable storage medium, and the processor 410 can run the program instructions to implement the voice broadcast method for visually impaired passengers and / or other desired functions of the various embodiments of the present application described above. Various contents such as category correspondence and the like can also be stored in the computer-readable storage medium.

[0205] In one example, the electronic device 400 can also include an input device 430 and an output device 440, which are interconnected by a bus system and / or other forms of connection mechanism (not shown).

[0206] In addition, the input device 430 can also receive data and information input by the user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer or a gravity sensor, and the like. The output device 440 can output various information to the outside. The output device 440 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.

[0207] Of course, in order to simplify, Figure 4 Only some of the components in the electronic device 400 related to the present application are shown in the figure, and components such as buses, input / output interfaces, and the like are omitted. In addition, the electronic device 400 can also include any other appropriate components according to specific application conditions.

[0208] Exemplary computer program product and computer readable storage medium

[0209] In addition to the methods and devices described above, embodiments of the present application can also be a computer program product including computer program instructions that, when run by a processor, cause the processor to perform the steps of the voice broadcast method for visually impaired passengers according to various embodiments of the present application described in the above "Exemplary Methods" section of the specification.

[0210] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0211] In addition, embodiments of the present application can also be a computer readable storage medium having stored thereon computer program instructions that, when run by a processor, cause the processor to perform the steps of the voice broadcast method for visually impaired passengers according to various embodiments of the present application described in the above "Exemplary Methods" section of the specification.

[0212] The computer readable storage medium can be any combination of one or more non-transitory media. The non-transitory medium can be a non-transitory signal medium or a non-transitory storage medium. The non-transitory storage medium can include, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the non-transitory storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0213] The basic principles of the present application are described above in combination with specific embodiments, but it should be noted that the advantages, advantages, effects, etc. mentioned in the present application are only examples and are not limiting, and these advantages, advantages, effects, etc. cannot be considered as the necessary possession of each embodiment of the present application. In addition, the above-mentioned specific details are only for the purpose of example and for the purpose of understanding, and are not limited to the above-mentioned specific details, and the above-mentioned details do not limit the present application to the above-mentioned specific details.

[0214] For simple description, each of the foregoing method embodiments is described as a combination of a series of actions, but those skilled in the art shall understand that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art shall understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0215] It should be noted that each of the embodiments in the present specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between each embodiment can be understood by referring to each other. For device embodiments, since they are basically similar to method embodiments, they are described more simply, and the relevant parts can be understood by referring to the part of the method embodiment.

[0216] The steps in the method of each embodiment of the present application can be adjusted, combined and reduced in sequence according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0217] The block diagrams of the devices, apparatuses, equipment, systems involved in the present application are only illustrative examples and are not intended to require or imply that the connection, arrangement, configuration must be as shown in the block diagram. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any way. Words such as "include", "contain", "have" and the like are open-ended words, which mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.

[0218] It should also be noted that in the devices, equipment and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombination shall be regarded as equivalent solutions of the present application.

[0219] The modules or sub-modules described as separate components can or can not be physically separated, and the components as modules or sub-modules can or can not be physical modules or sub-modules, i.e. they can be located in one place, or distributed on multiple network modules or sub-modules. Part or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the present embodiment.

[0220] In addition, each functional module or sub-module in each embodiment of the present application can be integrated in one processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The integrated module or sub-module can be realized in the form of hardware or in the form of a software functional module or sub-module.

[0221] Those skilled in the art will further appreciate that the functions or steps of the examples described herein can be implemented using electronic hardware, computer software, or any combination of the two. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their functional generalities. Whether such functions are implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0222] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software executed by a processor, or in a combination of the two. A software unit can reside in random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0223] Finally, it needs to be pointed out that, in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or sequence between the entities or operations. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0224] The above description of disclosed embodiments enables one of ordinary skill in the art to make and use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A voice announcement method for visually impaired passengers, characterized by, The method comprises the following steps: obtaining environment data of an environment in which a target vehicle is located; generating voice broadcast information according to the environment data by using a large language model; the voice broadcast information is used to introduce the environment in which the target vehicle is located to a visually impaired passenger who rides in the target vehicle.

2. The method of claim 1, wherein, The method further comprises: determining broadcast duration information of the voice broadcast information according to a driving speed of the target vehicle; the driving speed is negatively correlated with the broadcast duration of the voice broadcast information; the step of generating the voice broadcast information according to the environment data by using the large language model comprises: generating the voice broadcast information according to the environment data and the broadcast duration information by using the large language model.

3. The method of claim 2, wherein, The types of the environment data are at least two; the step of determining the broadcast duration information of the voice broadcast information according to the driving speed of the target vehicle comprises: determining total duration information of the voice broadcast information according to the driving speed; allocating segment duration information of a broadcast segment generated according to various environment data according to the driving speed and priority corresponding to the various environment data; the step of generating the voice broadcast information according to the environment data and the broadcast duration information by using the large language model comprises: generating the voice broadcast information according to the environment data, the total duration information and the segment duration information by using the large language model.

4. The method of claim 1, wherein, The method further comprises: determining the broadcast duration information of the voice broadcast information according to priority corresponding to the environment data; the higher the priority corresponding to the environment data is, the shorter the broadcast duration of the voice broadcast information is; the step of generating the voice broadcast information according to the environment data by using the large language model comprises: generating the voice broadcast information according to the environment data and the broadcast duration information by using the large language model.

5. The method of claim 1, wherein, The step of obtaining the environment data of the environment in which the target vehicle is located comprises at least one of the following steps: obtaining interest point data of an interest point on a to-be-traveled path of the target vehicle; obtaining information data related to a location of the target vehicle; obtaining environment perception data collected by an environment perception device based on the target vehicle.

6. The method of claim 5, wherein, The step of obtaining the interest point data of the interest point on the to-be-traveled path of the target vehicle comprises: obtaining a to-be-traveled path set of the target vehicle; querying an interest point attribute of a first interest point on each to-be-traveled path in the to-be-traveled path set; the interest point attribute at least comprises a category attribute of the interest point; obtaining a second interest point that satisfies a preset priority condition from the first interest point according to a preset mapping relationship; the preset mapping relationship is used to represent a corresponding relationship between the category attribute of the interest point and the priority; generating the interest point data according to the interest point attribute of the second interest point.

7. The method of claim 6, wherein, The step of generating the interest point data according to the interest point attribute of the second interest point comprises: pre-generating an interest point sentence used to describe the second interest point according to the interest point attribute of the second interest point by using the large language model; According to the position information of the second interest point, a target interest point is filtered from the second interest point; the target interest point includes a second interest point on the to-be-traveled path and between the target vehicle, and the distance between the target interest point and the target vehicle does not exceed a preset distance threshold; An interest point statement of the target interest point is determined as the interest point data.

8. The method of claim 1, wherein, Before the environment data of the environment where the target vehicle is located is acquired, the method further includes: Determining whether a visually impaired passenger exists in the target vehicle; The environment data of the environment where the target vehicle is located is acquired, including: In a case where it is determined that the visually impaired passenger exists in the target vehicle, the environment data of the environment where the target vehicle is located is acquired.

9. The method of claim 8, wherein, The determination of whether the visually impaired passenger exists in the target vehicle includes: Acquiring in-vehicle environment perception data in the target vehicle; According to the in-vehicle environment perception data, it is determined whether a personal item of the visually impaired passenger exists in the target vehicle; In a case where the personal item of the visually impaired passenger exists in the target vehicle, according to a preset action recognition model and the in-vehicle environment perception data, it is determined whether a limb action of a passenger in the target vehicle belongs to an action mode of the visually impaired passenger; According to whether the limb action of the passenger in the target vehicle belongs to the typical action mode of the visually impaired passenger, it is determined whether the visually impaired passenger exists in the target vehicle.

10. An electronic device, comprising: Comprise: A processor; A memory for storing instructions executable by the processor; The processor is configured to execute the method in any one of claims 1 to 8.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer program instructions, and the computer program instructions, when executed by the processor, cause the processor to execute the method in any one of claims 1 to 9. The computer readable storage medium stores computer program instructions, and the computer program instructions, when executed by the processor, cause the processor to execute the method in any one of claims 1 to 9.

Citation Information

Cited By

  • Vehicle intelligent agent service system based on large model

    CN121069858A

  • Ice and snow tourism creative tour guiding and abnormity early warning system based on behavior state monitoring

    CN121438540A