Method, device, equipment and system for processing in-vehicle infotainment service request

By constructing a neural network model and visual language model processing supplemented by mixed features, the problem of inaccurate recognition of large-screen controls in the car computer is solved, and more accurate control recognition and intelligent service response are achieved, improving user experience.

CN120259683APending Publication Date: 2025-07-04GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510333959.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When the prior art uses visual language models to understand the content of the large screen of the vehicle computer, the detection granularity is too coarse and the target control cannot be accurately identified, resulting in bias in service response and affecting the user experience.

Method used

A neural network model with mixed feature supplement function is built, and feature element recognition of screen images is recognized through optical character recognition and functional description model, text content and location information are extracted, and visual language models are used to process it to generate more accurate service responses.

Benefits of technology

It improves the recognition accuracy of the large-screen controls of the car machine, enhances the fluency and satisfaction of the user experience, and achieves a more intelligent and personalized service response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259683A_ABST
    Figure CN120259683A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a method, a device, machine equipment and a system for processing a vehicle machine service request. The method comprises the following steps: in response to a service request of a user, obtaining a screen image displayed on a large screen of the vehicle-mounted terminal; performing feature element recognition on the screen image to obtain at least one feature element, and detecting the at least one feature element to obtain at least one text content and / or position information; constructing a neural network model with a mixed feature supplementing function, and inputting the at least one piece of text content and / or position information into the neural network model; and performing mixed feature supplementation on the at least one text content and / or position information by using a neural network model, processing the at least one feature element after supplementation through a visual language model to obtain a service response, and feeding back the service response to a user. According to the method, the target control is recognized by utilizing the supplemented feature elements, so that a recognition task with finer granularity is realized, and the accuracy of control recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, device, equipment, and system for processing vehicle-mounted service requests. Background Art

[0002] In the process of the in-vehicle computer controller providing intelligent driving services to users, the in-vehicle large screen needs to understand any scenario during driving. A process of recognizing or understanding a scenario includes: completing operations such as searching, clicking, or swiping on the target control on the in-vehicle large screen according to the instructions issued by the user. Currently, in the process of using a Vision Language Model (VLM) to understand the in-vehicle large screen and provide service responses, there are often problems such as a relatively coarse granularity in detecting icons or text on the large screen content, inability to accurately identify the position of the target control, or inability to correctly identify the target control, resulting in deviations or errors in the feedback service responses and seriously affecting the user experience. Summary of the Invention

[0003] In view of this, this application provides a method, device, equipment, and system for processing vehicle-mounted service requests to solve or improve the problem that the system's misinterpretation of the content on the in-vehicle large screen affects the user's service experience.

[0004] In a first aspect, this application provides a method for processing vehicle-mounted service requests, the method including:

[0005] In response to a user's service request, obtain a screen image displayed on the in-vehicle large screen; perform feature element recognition on the screen image to obtain at least one feature element, and perform detection on the at least one feature element to obtain the corresponding at least one text content and / or position information; construct a neural network model with a hybrid feature supplementation function, and input the at least one text content and / or position information into the neural network model; use the neural network model to perform hybrid feature supplementation on the at least one text content and / or position information, and output the supplemented at least one feature element; process the supplemented at least one feature element through a vision language model to obtain a service response, and feedback the service response to the user.

[0006] The processing method provided in this aspect uses a constructed neural network model with a hybrid feature supplementation function and a UI component parsing model to perform hybrid feature supplementation on the at least one text content and / or position information to obtain the supplemented feature element, and then uses the supplemented feature element to be processed through a vision language model to obtain a service response. This method uses the supplemented feature element to identify the target control, thereby realizing a finer-grained recognition task, improving the accuracy of control recognition, and further improving the user's service experience.

[0007] In combination with the first aspect, in a possible implementation, the feature element includes text and an icon;

[0008] Detect at least one feature element to obtain the corresponding at least one text content and / or location information, including: using optical character recognition technology to detect at least one text and icon in the screen image, identifying the content and the border position of each text and icon; according to the border position, if there is an overlapping situation in the border positions of the text and / or icon, then using the principle of taking the inner frame, perform deduplication and merging processing on the overlapping border positions of the text and / or icon to obtain at least one location information and the content information corresponding to the content of at least one icon.

[0009] In combination with the first aspect, in another possible implementation, construct a neural network model with a hybrid feature supplementation function, including: obtaining an initial neural network model, and obtaining an optical character recognition model and a function description model; adding the optical character recognition model and the function description model to the initial neural network model to construct and generate a neural network model with a hybrid feature supplementation function. Among them, the optical character recognition model is a model for identifying text content, and the function description model is a model for generating a text description of an icon or an interface element.

[0010] In combination with the first aspect, in yet another possible implementation, before performing feature element recognition on the screen image, it further includes: dividing the screen image into multiple image blocks.

[0011] Use the neural network model to perform hybrid feature supplementation on at least one text content and / or location information, and output at least one supplemented feature element, including: inputting multiple image blocks into the neural network model, and using the optical character recognition model and the function description model to perform hybrid feature supplementation on the text content and / or location information in each image block, and output at least one feature element.

[0012] In combination with the first aspect, in yet another possible implementation, before processing at least one supplemented feature element through a vision-language model, it further includes: inputting a predetermined image and a partial area delimited in the predetermined image into the vision-language model; according to the predetermined image, training the vision-language model to learn the delimited partial area to obtain a trained vision-language model.

[0013] In combination with the first aspect, in yet another possible implementation, process at least one supplemented feature element through a vision-language model to obtain a service response and feedback the service response to the user, including: according to the content of the service request, using the trained vision-language model to reason about at least one feature element to generate a service response, and the service response includes: text description, image or action operation instruction; display the service response on the in-vehicle large screen or broadcast it to the user through voice.

[0014] In a second aspect, the present application further provides a processing device for a vehicle-mounted device service request. The device includes:

[0015] An acquisition module, configured to acquire a screen image displayed on a large screen of a vehicle-mounted device in response to a service request from a user;

[0016] A processing module, configured to perform feature element recognition on the screen image to obtain at least one feature element, and perform detection on the at least one feature element to obtain corresponding at least one text content and / or location information;

[0017] A construction module, configured to construct a neural network model with a hybrid feature supplementation function, and input the at least one text content and / or location information into the neural network model;

[0018] The processing module is further configured to use the neural network model to perform hybrid feature supplementation on the at least one text content and / or location information, output the at least one supplemented feature element, and process the at least one supplemented feature element through a vision-language model to obtain a service response, and feedback the service response to the user.

[0019] In a third aspect, the present application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method for processing a vehicle-mounted device service request according to the first aspect or any corresponding embodiment thereof.

[0020] In a fourth aspect, the present application provides a computer-readable storage medium, on which computer instructions are stored. The computer instructions are used to cause a computer to execute the method for processing a vehicle-mounted device service request according to the first aspect or any corresponding embodiment thereof.

[0021] In addition, the present application provides a computer program product, including computer instructions, which are used to cause a computer to execute the method for processing a vehicle-mounted device service request according to the first aspect or any corresponding embodiment thereof.

[0022] The method, device, equipment and system for processing a vehicle-mounted device service request provided by the present application have the following beneficial effects:

[0023] Enhance user experience: By acquiring the screen image on the large screen of the vehicle-mounted device in real time and performing feature element recognition on it, the system can quickly understand the current interaction environment and possible intentions of the user, so as to provide a service response that better meets the user's needs. This instant feedback mechanism significantly enhances the fluency and satisfaction of the user experience.

[0024] Enhanced information processing capabilities: The hybrid detection identifies the characteristic elements and extracts their corresponding text content and / or location information, enabling the system to comprehensively understand the screen content. This meticulous information extraction ability provides a solid foundation for subsequent intelligent processing and helps to achieve more accurate service recommendations or instruction executions.

[0025] Optimize the performance of the neural network model: Construct a neural network model with the function of hybrid feature supplementation and input the extracted text content and / or location information. It can make full use of the deep learning ability of the neural network to perform more complex and meticulous supplementary processing on the characteristic elements. This not only improves the accuracy of feature recognition but also enhances the model's adaptability to complex scenarios.

[0026] Achieve intelligent service response: The supplementary characteristic elements processed by the vision-language model can generate more intelligent and personalized service responses. This intelligent service method not only improves service efficiency but also brings a more convenient and efficient interaction experience for users.

[0027] Promote technological innovation and upgrading: The entire technical process involves multiple technical fields such as image recognition, neural network model construction, and vision-language model processing. Its successful application will drive the continuous innovation and upgrading of related technologies. At the same time, this cross-field technology integration also provides new ideas and directions for the intelligent development of future intelligent driving and in-vehicle infotainment systems. Description of the Drawings

[0028] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0029] Figure 1 It is a schematic flowchart of a method for processing a vehicle-mounted service request according to an embodiment of the present application;

[0030] Figure 2 It is a schematic flowchart of another method for processing a vehicle-mounted service request according to an embodiment of the present application;

[0031] Figure 3 It is a schematic flowchart of a hybrid detection method according to an embodiment of the present application;

[0032] Figure 4 It is a schematic flowchart of yet another method for processing a vehicle-mounted service request according to an embodiment of the present application;

[0033] Figure 5It is a schematic flowchart of a hybrid feature extraction according to an embodiment of the present application;

[0034] Figure 6 It is a structural block diagram of a processing device for a vehicle-mounted service request according to an embodiment of the present application;

[0035] Figure 7 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application;

[0036] Figure 8 It is a schematic diagram of the structure of a processing system for a vehicle-mounted service request according to an embodiment of the present application. Detailed implementation manners

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0038] It should be noted that in the description of the present application, the terms "including", "comprising", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements that are not explicitly listed, or elements that are inherent to such process, method, article, or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0039] To enable those skilled in the art of the present technology to better understand the solution of the present application, the application scenarios and related technical terms involved in the technical solution of the present application will be described first.

[0040] VLM (Vision Language Model), that is, a vision language model, is an artificial intelligence model that combines visual information and language information. Through deep learning technology, especially the combination of convolutional neural networks and Transformer architectures, the model can understand and generate content involving images and texts.

[0041] VLM is a multimodal model capable of learning from images and text. By integrating visual information with language information, it achieves the understanding and description of complex visual scenes. Its features include: (1) Multimodal processing: VLM can simultaneously process data in two modalities, images and text, achieving cross-modal information fusion. (2) Advanced reasoning ability: VLM can not only understand text input but also provide advanced reasoning and generate text responses, while also having the ability to process image input. (3) Powerful performance: VLM demonstrates powerful performance in various visual and language tasks, such as image caption generation, visual question answering, image retrieval, etc.

[0042] In this application, it is mainly applied to handle Visual Question Answering (VQA) problems. Specifically, VLM can understand natural language questions and provide answers based on image content. This technology can be used in educational software, virtual assistants, and interactive customer service systems.

[0043] This solution can be applied to the technical scenario of computers. In the in-vehicle infotainment (IVI) large screen, multiple application APPs and prompt messages can be displayed. Among them, the IVI large screen can be used to display and provide relevant service responses to users, such as providing navigation services, playing audio / video, and controlling various functional modules in the vehicle.

[0044] The Agent module is used to parse or respond to service requests proposed by users. Among them, the Agent module can further include one or more sub-modules, such as a planning module, a reflection module, a parsing module, an execution module, etc. Optionally, the Agent module can be arranged on the server.

[0045] In the Agent module or the screen understanding scenario, there is a task requirement: to complete operations such as clicking, searching, or swiping on the target control on the IVI large screen according to text instructions. However, for visual multimodal models, it is quite difficult. For example, VLM cannot directly complete the task of returning the region coordinates [x1, y1, x2, y2] according to text instructions. The main reasons are that the detection granularity is too coarse and it cannot correctly identify the target control or content.

[0046] To solve or improve the above problems, the embodiments of this application provide a method for processing in-vehicle service requests. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0047] In this embodiment, a method for processing in-vehicle service requests is provided, which can be used in a server or a server cluster. In addition, it can also be applied to other network devices, such as data centers, etc. Figure 1It is a schematic flowchart of a method for processing a vehicle-mounted device service request according to an embodiment of the present application. As Figure 1 shown, the method includes:

[0048] Step S101: In response to a service request from a user, obtain a screen image displayed on the large screen of the vehicle-mounted device.

[0049] Among them, the service request can be generated by the vehicle-mounted device controller and sent to the server. When the user issues a voice command in the vehicle, a service request is generated according to the content of the voice command. Based on this service request, a screenshot of the large screen of the vehicle-mounted device is taken to obtain a screen image, and the screen image is sent to the server.

[0050] Optionally, the sent screen image can be one or more. The screen image includes multiple applications (APPs), and at least one target APP among these APPs can be used to provide a service response for the user.

[0051] Step S102: Identify at least one feature element from the screen image and detect the at least one feature element to obtain the corresponding at least one text content and / or location information.

[0052] Among them, one feature element can be detected with one or more text contents, or the location information of the feature element can be detected. The location information can be represented by coordinates.

[0053] Specifically, this step includes: first preprocessing the screen image, loading the screen image: using an image processing library (such as OpenCV) to load the screen capture or the real-time captured screen image, and then performing grayscale conversion, denoising, etc. on the screen image to make it easier to identify feature elements.

[0054] Then, identify the feature elements from the screen image. For example, edge detection and other algorithms can be used to identify the edge information in the image. And use contour detection algorithms (such as findContours) to identify the shapes in the image (such as rectangles, circles, polygons, etc.). These shapes may represent feature elements such as buttons, icons, and text boxes. Optionally, for possible text areas, use optical character recognition (OCR) technology (such as Tesseract) to extract the text content. Finally, for known feature elements (such as specific icons or buttons), template matching algorithms can be used to detect their positions in the image.

[0055] The above process of obtaining text content and / or location information includes: extracting text content: for the identified text regions, directly extract the text content from the OCR results. Recording location information: for each identified feature element, record its location information in the image (such as the coordinates of the upper left corner, width, and height). Data integration: integrate the extracted text content and location information into a data structure (such as a dictionary or object) for subsequent use.

[0056] Step S103: Construct a neural network model with a hybrid feature supplementation function, and input at least one text content and / or location information into the neural network model.

[0057] Specifically, constructing a neural network model with a hybrid feature supplementation function includes: obtaining an initial neural network model, and obtaining an optical character recognition model (such as an OCR model) and a function description model; adding the optical character recognition model and the function description model to the initial neural network model to construct and generate a neural network model with a hybrid feature supplementation function.

[0058] Among them, the optical character recognition model is a model for recognizing text content (text), and the function description model is a model for generating a text description of an icon or an interface element.

[0059] Step S104: Use the neural network model to perform hybrid feature supplementation on at least one text content and / or location information, and output at least one supplemented feature element.

[0060] Among them, at least one supplemented feature element includes: hybrid detection (such as UI element detection, OCR text location detection), hybrid feature supplementation (text region OCR, icon region description feature), etc.

[0061] Step S105: Process at least one supplemented feature element through a vision-language model to obtain a service response, and feedback the service response to the user.

[0062] Specifically: this step includes: according to the content of the service request, use the trained vision-language model to reason about at least one feature element to generate a service response, and the service response includes: text description, image, or action operation instruction; display the service response on the in-vehicle infotainment system screen, or broadcast it to the user through voice.

[0063] The method for processing in-vehicle service requests provided by this application includes the following beneficial effects:

[0064] Enhance user experience: By obtaining the screen image on the in-vehicle large screen in real time and identifying the characteristic elements therein, the system can quickly understand the current interaction environment and possible intentions of the user, so as to provide service responses that better meet the user's needs. This instant feedback mechanism significantly enhances the fluency and satisfaction of the user experience.

[0065] Enhance information processing capabilities: Mix the detected characteristic elements and extract their corresponding text content and / or location information, enabling the system to more comprehensively understand the screen content. This meticulous information extraction ability provides a solid foundation for subsequent intelligent processing, helping to achieve more accurate service recommendations or instruction executions.

[0066] Optimize the performance of the neural network model: Construct a neural network model with a hybrid feature supplementation function and input the extracted text content and / or location information, which can make full use of the deep learning ability of the neural network to perform more complex and meticulous supplementation processing on the characteristic elements. This not only improves the accuracy of feature recognition but also enhances the model's adaptability to complex scenarios.

[0067] Achieve intelligent service responses: The supplemented characteristic elements processed by the vision-language model can generate more intelligent and personalized service responses. This intelligent service method not only improves service efficiency but also brings a more convenient and efficient interaction experience to users.

[0068] Promote technological innovation and upgrading: The entire technical process involves multiple technical fields such as image recognition, neural network model construction, and vision-language model processing. Its successful application will promote the continuous innovation and upgrading of related technologies. At the same time, this cross-field technology integration also provides new ideas and directions for the intelligent development of future intelligent driving and in-vehicle infotainment systems.

[0069] Furthermore, in a possible implementation manner of this embodiment, the above-mentioned characteristic elements include at least one of text and icons. As Figure 2 shown, the above-mentioned step S102: Detect at least one characteristic element to obtain at least one text content and / or location information corresponding to the at least one characteristic element, specifically including:

[0070] Step S1021: Use optical character recognition technology to detect at least one text and icon in the screen image, and identify the content and border position of each text and icon.

[0071] Step S1022: According to the border positions, if there is an overlap in the border positions of the text and / or icon, use the principle of taking the inner frame to perform duplicate removal and merging processing on the overlapping border positions of the text and / or icon to obtain at least one location information and the content information corresponding to the content of at least one icon.

[0072] See Figure 3 shown in the figure, which is a schematic flowchart of a hybrid detection method according to an embodiment of the present application.

[0073] Combining the optical character recognition OCR detection frame and the UI element detection frame for filtering and deduplication, the deduplication method follows the principle of taking the innermost frame, and merging into the finest-grained detection result as Figure 3 shown in the figure, and then performing visual box filtering processing on the detection result to obtain the detection result, which includes the at least one position information and the content information corresponding to at least one icon or text content.

[0074] In one example, assume the coordinates of the rectangle are: rectangle Box 1: {(x 11 , y 11 , x 12 , y 12 )}; rectangle Box2: {(x 21 , y 21 , x 22 , y 22 )}. It is necessary to calculate the intersection of the two rectangles to deduplicate the information of the intersection set, so as to improve the recognition accuracy.

[0075] Optionally, in another implementation manner of this embodiment, as Figure 4 shown in the figure, the above step S102: Before performing feature element recognition on the screen image, further includes: dividing the screen image into multiple image blocks. The size of each image block can be the same or different.

[0076] Step S104: Using a neural network model to perform hybrid feature supplementation on at least one text content and / or position information, and outputting at least one supplemented feature element, including:

[0077] Step S104': Inputting multiple image blocks into the neural network model, and using an optical character recognition model and a function description model to perform hybrid feature supplementation on the text content and / or position information in each image block, and outputting at least one feature element.

[0078] Specifically, as Figure 5 shown in the figure, inputting multiple image blocks into the neural network model, using the OCR rec model and the function description model to jointly construct the core features of the text text and the icon icon. The main purpose of the hybrid features is: to improve the efficiency of the hybrid feature module, and then after being processed by the OCR rec model and the function description model, the recognition result is obtained.

[0079] In this embodiment, a recognition method using hybrid features is adopted, and the OCR rec model is cited. The OCR rec model generally refers to a recognition model for optical character recognition. OCR technology is a technology that can convert printed or handwritten text into machine-editable and searchable digital text. The OCR rec model is the key component to achieve this conversion.

[0080] The OCR rec model recognizes the text in the input image or scanned document through deep learning or traditional image processing techniques. It can automatically locate, segment, and recognize the characters in the image and convert them into a text format that can be understood by a computer. The technical principles of the OCR rec model generally include: Preprocessing: Preprocess the input image, including operations such as denoising, binarization, and image enhancement, to improve the accuracy of character recognition. Character localization: Use image processing techniques, such as connected component analysis, maximally stable extremal regions (MSER), etc., to locate the character regions in the image. Character segmentation: Segment the located character regions into individual characters for subsequent recognition processing. Character recognition: Use deep learning models (such as convolutional neural network CNN, recurrent neural network RNN and its variants LSTM, GRU, etc.) or traditional classifiers (such as logistic regression, support vector machine SVM, etc.) to recognize the segmented characters. Postprocessing: Combine language models (such as hidden Markov model HMM) and rules to postprocess the recognition results to improve the accuracy and effectiveness of recognition, etc.

[0081] In addition, in this embodiment, before the above step S105: processing the at least one supplemented feature element through the vision language model, the method further includes: inputting a predetermined image and a partial area delimited in the predetermined image into the vision language model; training the vision language model to learn the delimited partial area according to the predetermined image to obtain a trained vision language model.

[0082] This step is the process of generating a trained VLM for iocn icon recognition. Specifically, one implementation is: first obtain the training task, based on the given detected picture or image, and then give the serial number area that needs to be described, so that the VLM learns the core function description of this area, and a new VLM model is trained.

[0083] For example, first obtain the data, combine xml parsing to filter out the icon (icon) category, split the entire screen image into 1, 2, 3, 4, 5, 6....n image blocks, and then input each image block into the generative pre-trained Transformer model (such as GPT 4.0), use this GPT 4.0 to pre-annotate the core function of each difference, and perform manual verification. Finally, the following data set is formed:

[0084] Python

[0085] Input:

[0086] Please describe the core function of the icon in area 76 within 10 characters;

[0087] Output:

[0088] Add product "Orange C Americano".

[0089] The method provided in this embodiment analyzes the problems of past solutions and proposes a module combining hybrid features and hybrid detection to solve or improve the problems of rough detection strength and inaccurate VLM recognition target. This method combines the strategies of detection mixing and feature mixing, proposes an OCR det model and a UI component parsing model to complete the fine-grained supplementary task of boxes, and proposes an OCR rec and a function description model to complete the rapid supplementary tasks of text and icon features.

[0090] In addition, this method also proposes an alignment method for all features of full-automatic local data construction to enhance the ability to provide local feature alignment for VLM.

[0091] In this embodiment, a processing device for in-vehicle service requests is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0092] This embodiment provides a processing device for in-vehicle service requests, as Figure 6 shown. This device includes: an acquisition module 610, a processing module 620, and a construction module 630. In addition, this device may also include other more or fewer modules, such as a storage module, a sending module, etc.

[0093] Among them, the acquisition module 610 is used to acquire the screen image displayed on the in-vehicle large screen in response to the user's service request.

[0094] The processing module 620 is used to identify at least one feature element from the screen image and detect the at least one feature element to obtain the corresponding at least one text content and / or location information.

[0095] The construction module 630 is used to construct a neural network model with a hybrid feature supplementation function and input the at least one text content and / or location information into the neural network model.

[0096] The processing module 620 is further configured to use a neural network model to supplement hybrid features for at least one text content and / or location information, output at least one supplemented feature element, and process the at least one supplemented feature element through a vision-language model to obtain a service response, and feedback the service response to the user.

[0097] In some alternative embodiments, the feature elements include text and icons.

[0098] The processing module 620 is specifically configured to use optical character recognition technology to detect at least one text and icon in the screen image, identify the content and border positions of each text and icon; according to the border positions, if there is an overlap in the border positions of the text and / or icon, the inner frame principle is used to perform de-duplication and merging processing on the overlapping border positions of the text and / or icon to obtain at least one location information and content information corresponding to the content of at least one icon.

[0099] In some other alternative embodiments, the building module 630 is specifically configured to obtain an initial neural network model, and obtain an optical character recognition model and a function description model; add the optical character recognition model and the function description model to the initial neural network model to build and generate a neural network model with a hybrid feature supplement function. Among them, the optical character recognition model is a model for identifying text content, and the function description model is a model for generating a text description of an icon or an interface element.

[0100] In some other alternative embodiments, the processing module 620 is further configured to divide the screen image into multiple image blocks before performing feature element recognition on the screen image; input the multiple image blocks into the neural network model, and use the optical character recognition model and the function description model to supplement hybrid features for the text content and / or location information in each image block, and output at least one feature element.

[0101] In some other alternative embodiments, the building module 630 is further configured to input a predetermined image and a partial area delimited in the predetermined image into the vision-language model before processing the at least one supplemented feature element through the vision-language model; train the vision-language model to learn the delimited partial area according to the predetermined image to obtain a trained vision-language model.

[0102] In some other alternative embodiments, the processing module 620 is further configured to reason about at least one feature element according to the content of the service request by using the trained vision-language model to generate a service response, where the service response includes: a text description, an image, or an action operation instruction; display the service response on the in-vehicle large screen or broadcast it to the user through voice.

[0103] The processing device provided in this aspect uses a neural network model and a UI component parsing model with a hybrid feature supplementation function to perform hybrid feature supplementation on at least one text content and / or location information to obtain supplemented feature elements, and then uses the supplemented feature elements to be processed by a vision-language model to obtain a service response. This method uses the supplemented feature elements to identify target controls, thereby achieving a finer-grained identification task and improving the accuracy of control identification.

[0104] The further functional descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.

[0105] The processing device for in-vehicle service requests in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0106] This application embodiment also provides an electronic device having the above-mentioned Figure 6 processing device for in-vehicle service requests shown.

[0107] Please refer to Figure 7 , which is a schematic structural diagram of an electronic device provided by an optional embodiment of this application. The electronic device includes: one or more processors 10, a memory 20, and an interface for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface).

[0108] In some optional implementation manners, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 7 In

[0109] Processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0110] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes a processing method for a vehicle-mounted service request shown in the above embodiments.

[0111] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0112] The memory 20 may include a volatile memory, for example, a random access memory; the memory may also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.

[0113] The electronic device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 7 Taking connection through a bus as an example. Among them, the input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function control of the electronic device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (for example, an LED), and a tactile feedback device (for example, a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0114] The electronic device further includes a communication interface for the electronic device to communicate with other devices or communication networks.

[0115] Embodiments of the present application also provide a computer-readable storage medium. The methods according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware.

[0116] Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the processing method of the in-vehicle infotainment service request shown in the above embodiments is implemented.

[0117] Embodiments of the present application can also provide a computer program product, including computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the above methods. Among them, the computer program product can be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, executed as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0118] Embodiments of the present application also provide a processing system for in-vehicle infotainment service requests, as Figure 8 shown. The processing system includes: a vehicle 100 and a server 200; among them, the vehicle 100 includes an in-vehicle infotainment controller 1001 and an in-vehicle infotainment large screen 1002, and among them, the in-vehicle infotainment controller 1001 is connected to the server 200. Specifically, the connection method can be through a wireless network, such as a WLAN connection.

[0119] In addition, the in-vehicle infotainment controller 1001 and the in-vehicle infotainment large screen 1002 can be connected by wire or wirelessly.

[0120] It should be understood that Figure 8In the shown system architecture, the number of the server 200 and the vehicle 100 is both 1. In fact, there can also be multiple servers and vehicles. This embodiment does not limit the number to be one or more, and it can be customized or designed according to the actual situation of the system.

[0121] Among them, the in-vehicle controller 1001 is used to generate a service request of the user and send the screen image displayed on the in-vehicle large screen 1002 to the server 200.

[0122] The server 200 is used to execute the Figures 1 to 5 processing method of the in-vehicle service request for vehicle driving shown above and send the generated service response to the in-vehicle controller 1001.

[0123] The in-vehicle controller 1001 is further used to receive the service response fed back by the server 200 and display the service response through the in-vehicle large screen 1002.

[0124] Specifically, for the respective functions and method steps of the in-vehicle controller 1001 and the server 200, refer to the description in the foregoing method embodiment, and details are not described herein again in this embodiment.

[0125] The above embodiments are only used to illustrate the technical solutions of the embodiments of the present application, rather than to limit them; although the embodiments of the present application have been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for processing a vehicle machine service request, characterized in that, The method includes: In response to a user's service request, obtaining a screen image displayed on the in-vehicle large screen; Performing feature element recognition on the screen image to obtain at least one feature element, and detecting the at least one feature element to obtain corresponding at least one text content and / or position information; Constructing a neural network model with a hybrid feature supplementation function, and inputting the at least one text content and / or position information into the neural network model; Using the neural network model to perform hybrid feature supplementation on the at least one text content and / or position information, and outputting at least one supplemented feature element; Processing the at least one supplemented feature element through a vision-language model to obtain a service response, and feeding back the service response to the user.

2. The method according to claim 1, wherein The feature elements include text and icons; Detecting the at least one feature element to obtain corresponding at least one text content and / or position information includes: Using optical character recognition technology to detect at least one text and icon in the screen image, and identifying the content and border position of each text and icon; According to the border positions, if there is an overlap in the border positions of the text and / or the icon, then using the principle of taking the inner frame, performing duplicate removal and merging processing on the border positions of the overlapping text and / or icon to obtain the at least one position information and content information corresponding to the content of at least one of the icons.

3. The method according to claim 1 or 2, characterized in that, Constructing the neural network model with a hybrid feature supplementation function includes: Obtaining an initial neural network model, and obtaining an optical character recognition model and a function description model; Adding the optical character recognition model and the function description model to the initial neural network model to construct and generate the neural network model with a hybrid feature supplementation function; Wherein, the optical character recognition model is a model for identifying text content, and the function description model is a model for generating a text description of an icon or an interface element.

4. The method according to claim 3, characterized in that, Before performing feature element recognition on the screen image, it further includes: Dividing the screen image into multiple image blocks; Using the neural network model to perform hybrid feature supplementation on the at least one text content and / or position information and outputting at least one supplemented feature element includes: Inputting the multiple image blocks into the neural network model, and using the optical character recognition model and the function description model to perform hybrid feature supplementation on the text content and / or position information in each image block, and outputting the at least one feature element.

5. The method according to claim 4, characterized in that Before processing the at least one supplemented feature element through the vision-language model, it further includes: Inputting a predetermined image and a partial area delimited in the predetermined image into the vision-language model; Training the vision-language model to learn the delimited partial area according to the predetermined image to obtain a trained vision-language model.

6. The method according to claim 5, characterized in that, Processing the at least one supplemented feature element through the vision-language model to obtain a service response, and feeding back the service response to the user includes: Based on the content of the service request, use the trained vision-language model to reason about the at least one feature element to generate the service response, where the service response includes: a text description, an image, or an action operation instruction; Display the service response on the in-vehicle large screen or announce it to the user by voice.

7. A processing device for in-vehicle infotainment service requests, characterized in that, The device includes: An acquisition module, configured to acquire a screen image displayed on the in-vehicle large screen in response to a user's service request; A processing module, configured to perform feature element recognition on the screen image to obtain at least one feature element, and detect the at least one feature element to obtain corresponding at least one text content and / or position information; A construction module, configured to construct a neural network model with a hybrid feature supplementation function, and input the at least one text content and / or position information into the neural network model; The processing module is further configured to use the neural network model to perform hybrid feature supplementation on the at least one text content and / or position information, output the supplemented at least one feature element, and process the supplemented at least one feature element through a vision-language model to obtain a service response, and feedback the service response to the user.

8. An electronic device, characterized in that, It includes a memory and a processor, and the memory is connected to the processor; The memory stores computer instructions, and the processor executes the computer instructions to execute the method for processing an in-vehicle service request according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium; The computer instructions are used to cause a computer to execute the method for processing an in-vehicle service request according to any one of claims 1 to 6.

10. A processing system for in-vehicle infotainment service requests, characterized in that, The system includes: a vehicle and a server; wherein, the vehicle includes an in-vehicle controller and an in-vehicle large screen, and the in-vehicle controller is connected to the server; The in-vehicle controller is configured to generate a user's service request and send the screen image displayed on the in-vehicle large screen to the server; The server is configured to execute the method for processing an in-vehicle service request for vehicle driving according to any one of claims 1 to 6; The in-vehicle controller is further configured to receive the service response fed back by the server and display the service response on the in-vehicle large screen.