Method and device for generating a prompt message for an artificial intelligence

FR3149406B1Active Publication Date: 2026-07-17PSA AUTOMOBILES SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
PSA AUTOMOBILES SA
Filing Date
2023-06-01
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Users face difficulties in effectively expressing needs to artificial intelligence systems, particularly multimodal ones, leading to complex and resource-intensive interactions, especially in vehicles, due to the need for specialized knowledge and trial-and-error formulation of prompt messages.

Method used

A method and device that utilize voice recognition, lexical analysis, and data acquisition to generate structured prompt messages for artificial intelligence, enabling users to interact simply and efficiently by converting vocal inputs into relevant computer instructions, reducing complexity and resource requirements.

Benefits of technology

Simplifies user interaction with multimodal AI systems by translating vocal needs into structured prompt messages, allowing for efficient activation of vehicle functions with reduced computational resources and complexity.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to a method and device for generating a prompt message for an artificial intelligence to receive a list of computer instructions from said artificial intelligence. The method is implemented by a processor and comprises the following steps: - Acquisition (200), by a speech recognition system, of an audio signal; Generation (210), by said speech recognition system, of a text from said audio signal; - Segmentation (220), by a lexical analysis system, of said text; - Determination (230), from said segmented text, of a type of prompt message, said type of prompt message being recognized by said artificial intelligence; - Determination (250) of said prompt message from said segmented text and said type of prompt message, said prompt message being transmitted (260) to said artificial intelligence. Figure to be published for the abstract: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for generating an invitation message for an artificial intelligence Technical field of the invention

[0001] The invention is in the field of artificial intelligence. In particular, the invention relates to simplifying the use of artificial intelligence, particularly in a vehicle. State of the art

[0002] The term “vehicle” means any type of vehicle such as a motor vehicle, a moped, a motorcycle, a storage robot in a warehouse, etc.

[0003] "Artificial intelligence", or AI, means a machine, a computer device, designed from large quantities of data and capable of providing decisions or requests based on an expressed need. It is known that the design of these devices is based on learning from expressed needs and expected decisions or requests. The inputs and outputs of artificial intelligence are digital, interpretable by computer.

[0004] For example, an expressed need includes an image, and artificial intelligence is able to extract or recognize shapes. For example, an expressed need includes several measurements of a state of a vehicle (speed, position, etc.) and artificial intelligence is able to make a decision (creation of a signal to alert / brake / accelerate, etc.).

[0005] A "prompt message" (or "prompt" in English) is an expressed need that can be interpreted by an artificial intelligence. It is known that, to obtain good results from an artificial intelligence, the prompt message must be correctly formulated. This is even more true for multimodal intelligences.

[0006] By "type" of prompt message is meant a specifically structured and / or formatted prompt message capable of being well interpreted by an artificial intelligence, the artificial intelligence having been designed by learning to respond (provide a list of instructions / requests, provide one or more decisions, etc.) according to the type of prompt message.

[0007] The term "list of instructions" means a computer signal that can be interpreted by a computer device. For example, a list of instructions is a text to be displayed, a sound or voice message to be emitted, a decision such as activating or deactivating a device, a decision such as detecting a threshold crossing by a combination of variables, etc.

[0008] “Multimodal intelligence” means an approach to artificial intelligence which aims to enable the computing device to take into account different modes of communication such as speech, image, text, gestures, etc. as an expressed need.

[0009] Any user interacting with artificial intelligence encounters difficulties in expressing a need understood and adapted to artificial intelligence. Thus, the list of computer instructions provided as a result by artificial intelligence, which may be information, a decision, a (computer) request, etc., is often not relevant to the needs expressed (out of context for example).

[0010] Currently, to partially overcome this difficulty, artificial intelligence is becoming increasingly complex, for example by increasing the number of layers in a neural network, to take into account the diversity of how different users express a need. This also generates increasingly long learning times with more and more data. For example, version 3.5 of the artificial intelligence known as "ChatGPT®" is a monomodal artificial intelligence whose training data size contains 175 billion parameters to be trained. The training data set contains 300 billion words, or approximately 570 gigabytes of textual data. Version 4 of ChatGPT® has become a multimodal intelligence, it has required even more training data, and it has become much more complex.

[0011] Also, voice recognition systems are known that allow human speech to be converted into written text. These systems operate using signal processing algorithms that analyze the audio and identify the different components of language, such as phonemes, words, intonations, etc.

[0012] Some of these voice recognition systems are simpler. There is no conversion of human speech into written text. From the identification of intonations or a predetermined sound sequence, action requests such as "turn off" or "turn on" a car radio are recognized. These are simpler systems but they are limited to a very limited number of language component identifications. The expressed need is limited to a very limited number (about ten) of functions. Furthermore, these systems are not capable of taking into account different modes of communication. Summary of the invention

[0013] An object of the present invention is to remedy the aforementioned problem, in particular to simplify the use of artificial intelligence by any user, who is not an expert in artificial intelligence and without knowledge of how the artificial intelligence was designed. This also makes it possible to limit the complexity of artificial intelligence, and thus allows a use requiring less computing resources (computing power, memory, etc.). Also, this allows the use of multimodal artificial intelligences simply without being an expert in how to "prompt" (create the prompt message adapted to the artificial intelligence).

[0014] To this end, a first aspect of the invention relates to a method for generating an invitation message for an artificial intelligence to receive from said artificial intelligence a list of computer instructions, said method being implemented by a processor and comprising the steps of: - Acquisition, by a voice recognition system, of an audio signal; - Generation, by said voice recognition system, of a text from said audio signal; - Segmentation, by a lexical analysis system, of said text; - Determining, from said segmented text, a type of prompt message, said type of prompt message being recognized by said artificial intelligence; - Determining said prompt message from said segmented text and said prompt message type, said prompt message being transmitted to said artificial intelligence.

[0015] Thus, a need expressed vocally by any user is translated into a prompt message capable of being processed appropriately by an artificial intelligence. The prompt message is constructed in a suitable manner so that the artificial intelligence can interpret this expressed need in a relevant manner to provide a reliable list of computer instructions. The user needs only little effort to interact with the artificial intelligence. The user does not need to search, by trial and error, empirically or haphazardly, for the formulation that is suitable for his expressed need to be processed by the artificial intelligence.

[0016] Advantageously, the processor is embedded in a vehicle.

[0017] Thus, simplified and relevant interaction between a user and artificial intelligence in a vehicle is possible. Also, since the complexity of the intelligence is less, the artificial intelligence can be embedded in a computer device in the vehicle.

[0018] Advantageously, said method further comprises a step of receiving said list of instructions, said list of instructions being processed by said processor or by another processor on board the vehicle.

[0019] Thus, the simplified and relevant interaction between a user and artificial intelligence allows the activation of functions in a vehicle, without being very limited by the number of functions.

[0020] Advantageously, said artificial intelligence is a multimodal intelligence.

[0021] Thus, a need expressed vocally by any user is translated into a prompt message that can be processed appropriately by multimodal artificial intelligence.

[0022] Advantageously, the method further comprises the steps of: - Acquisition of first data determining a state of a vehicle and / or second data determining a state of a passenger of said vehicle and / or third data determining an environment close to said vehicle; - Determining said prompt message from said prompt message type, from said segmented text, and from said first, second and / or third data.

[0023] Thus, the multimodal artificial intelligence receives as input, via the adapted prompt message, data coming from different modes of communication.

[0024] Advantageously, the second or third data comprise at least one image acquired by an image acquisition system.

[0025] Thus, the prompt message can transmit images to a multimodal intelligence capable of processing images as well.

[0026] Advantageously, following the acquisition of said audio signal, the method further comprises a step of recognizing a predetermined sound sequence without generating a text in order to determine a triggering of the step of generating a text from said audio signal.

[0027] Thus, the following steps of the method are only activated following recognition of the predetermined sound sequence. This makes it possible to limit the need for computing resources.

[0028] A second aspect of the invention relates to a device comprising a memory associated with at least one processor configured to implement the method according to the first aspect of the invention.

[0029] The invention also relates to a vehicle comprising the device.

[0030] The invention also relates to a computer program comprising instructions which, when the program is executed by the device according to the second aspect of the invention, lead the latter to implement the method according to the first aspect of the invention. Brief description of the figures

[0031] Other characteristics and advantages of the invention will emerge from the description of the non-limiting embodiments of the invention below, with reference to the appended figures, in which:

[0032] [Fig.l] schematically illustrates a device, according to a particular example of embodiment of the present invention.

[0033] [Fig.2] schematically illustrates a method of generating a prompt message to the attention of an artificial intelligence, according to a particular exemplary embodiment of the present invention. Detailed description of the invention

[0034] The invention is described below in its non-limiting application to the case of a motor vehicle traveling on a road or on a traffic lane. Other applications such as a robot in a storage warehouse or a motorcycle on a country road are also conceivable.

[0035] [Fig. 1] represents an example of a device 101 included in the vehicle, in a network (“cloud”) or in a server. This device 101 can be used as a centralized device in charge of at least certain steps of the method described below with reference to [Fig. 2]. In one embodiment, it corresponds to an autonomous driving computer.

[0036] This device 101 can take the form of a box comprising printed circuits, any type of computer or even a mobile telephone (“smartphone”).

[0037] The device 101 comprises a random access memory 102 for storing instructions for the implementation by a processor 103 of at least one step of the method as described above. The device also comprises a mass memory 104 for storing data intended to be retained after the implementation of the method.

[0038] The device 101 may further comprise a digital signal processor (DSP) 105. This DSP 105 receives data to format, demodulate and amplify, in a manner known per se, this data.

[0039] The device 101 also comprises an input interface 106 for receiving the data implemented by the method according to the invention and an output interface 107 for transmitting the data implemented by the method according to the invention.

[0040] For example, the input interface 106 can receive the following data: position or geographical location of the vehicle, speed and / or acceleration of the vehicle, set or predetermined positions / speeds / accelerations, engine speed, position and / or travel of the clutch, brake and / or acceleration pedal, detection of other vehicles or objects, position or geographical location of the other vehicles or objects detected, speed and / or acceleration of the other vehicles or objects detected, energy level or capacity such as the level of a battery, the level of a fuel and / or oxidant tank, operating states of sensors, confidence index of data originating from or processed by sensors and / or devices similar to the device 101.For example, the sensors capable of providing data are: GPS associated or not with mapping, tachometers, accelerometers, RADAR, LIDAR, lasers, ultrasound, camera, microphone, audio and / or image acquisition devices, ammeter, voltmeter, ohmmeter, all-or-nothing sensors, etc.

[0041] Also, the input interface 106 can receive an audio signal, a sound sequence, a text generated by a voice recognition system, a segmented text, a type of prompt message, a list of instructions, images or videos of the interior or exterior of a vehicle, a signal representative of a request to trigger a step of the method implemented by the processor 103, ...

[0042] For example, the output interface 107 can transmit data similar to the data received by the input interface 106. The output interface 107 can transmit an audio signal, a sound sequence, a text generated by a voice recognition system, a segmented text, a type of prompt message, a list of instructions, images or videos of the interior or exterior of a vehicle, a prompt message, a signal representative of a request to trigger a step of the method implemented by the processor 103, ...

[0043] The transmission may be to other devices similar to the device 101. These devices may be present in the vehicle and / or outside the vehicle, for example in a network (“cloud”) or in a server.

[0044] A voice recognition system may be included in a device 101.

[0045] [Fig.2] schematically illustrates a method for generating an invitation message for an artificial intelligence, according to a particular exemplary embodiment of the present invention.

[0046] Step 200, AcqA, is a step of acquisition, by a voice recognition system, of an audio signal. Voice recognition systems are known, they generally comprise at least one microphone or one sound acquisition device.

[0047] An acquisition of the audio signal can be triggered by different means such as from a human-machine interface (for example, pressing a button, etc.), or from, automatically, a minimum sound level, etc. A duration of acquisition of the signal is variable and can depend on numerous parameters, including the identification of a sound level below a threshold.

[0048] The device 101 may comprise the voice recognition system or, via the input interface 106, receive the audio signal.

[0049] For example, the audio signal is a signal representative of a need expressed by an occupant of the vehicle. For example, this need can be “Hello dear Artificial Intelligence. The weather is nice today. I would like to listen to some music”, “Play me some music!”, “I would like to listen to my playlist”, “Let’s play some music!” ...

[0050] Step 210, GenTxt, is a step of generating, by said voice recognition system, a text from said audio signal. It is known that certain voice recognition systems are more advanced and include technology for converting human speech into written text. These systems are called “speech-to-text” in English. This technology works using algorithms of signal processing that analyzes the acquired audio signal and identifies the different components of language, such as phonemes, words, intonations, etc. Conventionally, the audio signal is processed by the speech recognition system to remove background noise and interference. This system analyzes the acoustic characteristics of speech to identify sounds and words that are spoken. A text is then generated from the results of the analysis. This system generates information, like a computer signal. This information can be interpreted by a 101 type device to extract the generated text.

[0051] The device 101 may comprise all or part of the voice recognition system. For example, the device 101, via its input interface 106, may receive the audio signal, may apply the signal processing adapted to provide, via its output interface 107, a signal comprising the generated text. In another example, the device 101 transmits to another device 101 the audio signal totally or partially processed by signal processing (for example after elimination / reduction of noise and interference). The other device 101 may be outside the vehicle such as in a network (“cloud”) or in a server. In this case the device may receive, via its input interface 106, the text generated by the other device 101.

[0052] For example, if an occupant says “Good morning dear Artificial Intelligence. The weather is nice today. I would like to listen to some music,” said voice recognition system may provide the text: “Good morning dear Artificial Intelligence. The weather is nice today. I would like to listen to some music.”

[0053] Step 220, SegTxt, is a step of segmenting, by a lexical analysis system, said text. Lexical analysis systems are known. For example, a lexical analysis system is also called a lexeme analyzer or “tokenizer” in English.

[0054] The lexical analysis system is configured to divide a text into smaller pieces, called lexemes, lexical morphemes or, in English, "tokens". The tokens can be words, sentences, paragraphs or even individual characters. This division allows automatic processing of language (words emitted in the vehicle by an occupant for example).

[0055] A lexical analysis system is a computer device. For example, the device 101 may comprise a lexical analysis system. In another example, the device 101 transmits to another device 101 the text generated in step 210. The other device 101 may be outside the vehicle such as in a network (“cloud”) or in a server. In this case the device 101 may receive, via its input interface 106, the lexemes generated by the other device 101.

[0056] For example, if said voice recognition system provides the text "Hello dear Artificial Intelligence. The weather is nice today. I would like to listen to some music”, at the end of step 220 the segmented text can be: “Hello”, “Artificial intelligence”, “nice weather”, “today”, “listen”, “music”. In another example, the tokinenizer to split certain words into phonemes, lexemes, morphenes or other, and “listen” can be split into “eco-” and “-ter”.

[0057] Step 230, DetMsgType, is a step of determining, from said segmented text, a type of prompt message, said type of prompt message being recognized by said artificial intelligence. Indeed, for the artificial intelligence to provide relevant output, for example a list of instructions, the artificial intelligence must receive a recognized prompt message. By recognized, we mean that the artificial intelligence has been trained with certain types of prompt message. This also makes the artificial intelligence less complex because the artificial intelligence will then need to be less able to take into account all the possible variations of a naturally spoken language. The expressed need will be structured.

[0058] For example, a type of prompt message may be structured to include only one portion configured to contain a task to be performed. In another example, a type of prompt message may be structured to include a first portion configured to contain data describing a situation or context, and a second portion configured to contain a task to be performed. In another example, a type of prompt message may be structured to include several first portions configured to contain data describing a situation or context, and a second portion configured to contain a task to be performed. A first of the first portions may contain data describing a state of the vehicle (position, speed, etc.) and tokens from step 220. A second of the first portions may contain data or images of the interior of the vehicle and tokens from step 220.A third of the first parts may contain images of the exterior of the vehicle and tokens from step 220. And a fourth of the first parts may contain tokens from step 220.

[0059] For example, a part configured to contain a task to be performed begins with a keyword significant of an action to be performed such as "Question:", "Task:", "Request:", "Action:", ... For example, a part configured to contain data describing a situation or the context begins with a keyword significant of a context explanation such as "Either:", "Context:", "Speed:", "Position:", "Interior image:", "Exterior image:", ... The artificial intelligence is configured to recognize the keywords.

[0060] Advantageously, the determination of a type of invite message is based on the tokens. This selection can be based on a matrix which links sets of tokens to a type of messaging. For example, if the tokens resulting from the segmentation step only comprise orders, the type of invite message determined will be a message structured prompt message comprising only one part configured to contain a task to be performed. In another example, if the tokens resulting from the segmentation step comprise commands and comprise words, lexemes, lexical morphemes, ..., related to a state of the vehicle or a vehicle component, the determined type of prompt message will be a structured prompt message comprising a first part configured to contain data describing the state of the vehicle or the component, and a second part configured to contain a task to be performed.

[0061] In another example, the determination of the type of prompt message is based on a neural network that links sets of tokens to a messaging type, the neural network having been previously configured for this determination by training in a preliminary step.

[0062] For example, if the segmented text is "Hello", "Artificial Intelligence", "nice weather", "today", "listen", "music"", in this step, "listen" and the action complement "music" will be recognized. Also, the absence of additional information (vehicle status, driver status, nearby environment) will be recognized as well. In this case, the prompt message type is the prompt message type structured to include only a part configured to contain a task to be performed.

[0063] In another embodiment, on the example of segmented text above with "listen" and the action complement "music" recognized, in the presence of a determination of a more extended message type, the type of prompt message is the type of structured prompt message comprising a first part configured to contain data describing a situation or the context, and a second part configured to contain a task to be executed. From the word "music", the car radio / infotainment system / multimedia device is recognized. By "infotainment" or "infotainment device" is meant a multimedia device, an infotainment device, a car radio device, ... Thus, the context linked to the recognized device, here for example on, off, low volume, radio channel, ..., will be used in the following steps.

[0064] Step 240, AcqX, is a step of acquiring first data determining a state of a vehicle and / or second data determining a state of a passenger of said vehicle and / or third data determining an environment close to said vehicle.

[0065] Advantageously, the second or third data comprise at least one image acquired by an image acquisition system. For example, if the expressed need concerns the immediate environment in front of the vehicle, an image acquired by a camera on board the top of the windshield is included in the prompt message.

[0066] Advantageously, the data to be acquired are determined from the tokens. From a Similar to determining the type of prompt message from tokens, the data to be acquired is determined from the tokens from a matrix, a previously configured neural network system, or from any other decision support system.

[0067] For example, if one of the tokens contains the word "music", the car radio / in-fotainement / Multimedia component is recognized. Data on the state of this component, such as on, off, low volume, radio channel, etc., are retrieved by the input interface 106.

[0068] Step 250, DetMsg, is a step of determining said prompt message from said segmented text and said prompt message type. Advantageously, said prompt message is determined from said prompt message type, said segmented text, said first, second and / or third data.

[0069] Advantageously, the artificial intelligence is a multi-modal artificial intelligence when the prompt message comprises several modes of communication, for example text and images, or text and data of a state of a vehicle component, ....

[0070] In the example where the expressed need is “Hello dear Artificial Intelligence. The weather is nice today. I would like to listen to music”, where the determined prompt message type is a prompt message type structured to include only a part configured to contain a task to be performed, the prompt message may be “Request: listen to music”.

[0071] In the example where the expressed need is “Hello dear Artificial Intelligence. The weather is nice today. I would like to listen to music”, where the type of prompt message determined is the type of structured prompt message comprising a first part configured to contain data describing a situation or the context, and a second part configured to contain a task to be executed, the prompt message may be “Either: infotainment device {state on or off], {Volume X%}. Request: listen to music”, where “{state on or off]”, “Volume X%]” are alphanumeric information received in step 250. The states or information “{state on or off]” and “Volume X%]” are non-limiting examples.

[0072] Advantageously, the learning of artificial intelligence is carried out from a list of predetermined organs and states, information, from predetermined keywords, from predetermined tasks, expected actions. This makes it possible to obtain artificial intelligence optimized to the needs and thus a less complex artificial intelligence, simpler to implement, requiring fewer computing resources, and therefore embeddable in a vehicle.

[0073] In step 260, Tx, said prompt message is transmitted to said artificial intelligence. Advantageously, said artificial intelligence is a multi-intelligence modal. In one mode of operation, the artificial intelligence is embedded in the vehicle. In another mode of operation, the artificial intelligence is outside the vehicle, such as in a network ("cloud") or in a server.

[0074] Step 270, Rx, is a step of receiving said list of instructions, said list of instructions being processed by said processor or by another processor embedded in the vehicle. Advantageously, from the list of instructions are determined a component of said vehicle and an action linked to said component. Then from this identification, the action request is transmitted to the component so that the component carries out the action.

[0075] For example, from the list of instructions, such as "turn on infotainment", the infotainment device and the related action "turn on" / "start up" are identified. The processor transmits a start-up command to the infotainment device.

[0076] Advantageously, following the acquisition of said audio signal, the method further comprises a step of recognizing a predetermined sound sequence without generating a text in order to determine a triggering of the step of generating a text from said audio signal.

[0077] The present invention is not limited to the embodiments described above as examples: it extends to other variants.

[0078] For example, the method has been described according to a sequence of steps. Certain steps can be carried out in parallel or according to another sequence.

[0079] Thus, an embodiment has been described above in which information is transmitted. These transmissions can be provided by any type of communication technique such as wired or optical communications, radio frequency communications, wave communications, etc.

Claims

Claims

1. Method for generating a prompt message for an artificial intelligence to receive from said artificial intelligence a list of computer instructions, said method being implemented by a processor (103) and comprising the steps of: - Acquisition (200), by a voice recognition system, of an audio signal; - Generation (210), by said voice recognition system, of a text from said audio signal; - Segmentation (220), by a lexical analysis system, of said text; - Determination (230), from said segmented text, of a type of prompt message, said type of prompt message being recognized by said artificial intelligence; - Determination (250) of said prompt message from said segmented text and said type of prompt message, said prompt message being transmitted (260) to said artificial intelligence.

2. The method of claim 1, wherein the processor (103) is embedded in a vehicle.

3. The method of claim 2, wherein said method further comprises a step of receiving (270) said list of instructions, said list of instructions being processed by said processor (103) or by another processor on board the vehicle.

4. Method according to one of the preceding claims, wherein said artificial intelligence is a multi-modal intelligence.

5. Method according to one of claims 2 to 4, in which the method further comprises a step of: - Acquisition (240) of first data determining a state of a vehicle and / or second data determining a state of a passenger of said vehicle and / or third data determining an environment close to said vehicle; Said prompt message is determined (250) from said prompt message type, from said segmented text, and from said first, second and / or third data.

6. Method according to the preceding claim, in which the second or third data comprise at least one image acquired by an image acquisition system.

7. Method according to one of the preceding claims, in which, following the acquisition of said audio signal, the method further comprises a step of recognizing a predetermined sound sequence without generating a text in order to determine a triggering of the step of generating (210) a text from said audio signal.

8. Device (101) comprising a memory (102) associated with at least one processor (103) configured to implement the method according to one of the preceding claims.

9.

10. Vehicle comprising the device according to the preceding claim. Computer program comprising instructions which, when the program is executed by the device (101) according to claim 8, cause the latter to implement the method according to one of claims 1 to 7.