Hierarchical label generation method, device, electronic device and storage medium
By generating hierarchical labels in the in-vehicle system, the problems of shallow semantic understanding and architectural isolation in existing technologies are solved, more accurate personalized recommendations and in-vehicle safety applications are achieved, and the intelligent driving and safety performance of the in-vehicle system are improved.
Patent Information
- Application Number
- CN202510748648.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Existing in-vehicle audio tag generation technology has problems such as shallow semantic understanding, architectural isolation, and lack of structured association in the tag system. As a result, the generated tags cannot form a hierarchical knowledge graph, which limits the effectiveness of downstream applications.
By acquiring the audio information of the in-vehicle environment, performing sound recognition and extracting text information, using the first prompt instruction to generate basic labels, constructing label classification guidance information based on the second prompt instruction, and using the large model to establish a hierarchical structure between labels, a hierarchical label system is formed.
It improves the accuracy of label application in the fields of personalized recommendation and vehicle safety, can better handle hierarchical relationships in the context, and enhances the intelligent driving safety and personalized service capabilities of the vehicle system.
Smart Images

Figure CN120319243B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle-mounted technology, and in particular to a hierarchical label generation method, device, electronic device, and storage medium. Background Art
[0002] Current in-vehicle audio tagging technology primarily relies on Automatic Speech Recognition (ASR) systems and traditional Natural Language Processing (NLP) methods. The core process typically involves converting in-vehicle audio into text using ASR, then generating tags using keyword extraction algorithms (such as TF-IDF and TextRank) or matching against predefined rule bases.
[0003] Although the above methods can achieve basic tag generation, they have significant limitations, mainly reflected in the shallowness of semantic understanding and the isolation of the architecture. First, the keyword matching-based method can only recognize explicit words and cannot parse the implicit semantics of the context. Secondly, the long text processing capability is insufficient, and the keyword extraction method is prone to semantic loss due to information dispersion. In addition, the coverage and update mechanism of the predefined tag library restrict the adaptability of the technology. Therefore, the tag system lacks structured associations, the generated tags are discretely distributed, and a hierarchical knowledge graph cannot be formed, which limits the effectiveness of downstream applications. Summary of the Invention
[0004] In order to solve the technical problem that the labels generated by the existing methods lack semantic association and have limitations in application, the present invention provides a hierarchical label generation method, device, electronic device and storage medium. By converting audio information into text information and generating multiple basic labels based on a first prompt instruction, it is possible to capture deeper semantic associations, guide label classification based on a second prompt instruction, discover implicit relationships between labels, form a hierarchical structure between isolated labels, and establish a label system with hierarchical relationships. It can better handle hierarchical relationships in context and improve the accuracy of label application in personalized recommendations and vehicle safety.
[0005] In a first aspect, an embodiment of the present application provides a hierarchical label generation method, the method comprising:
[0006] The hierarchical label generation method is characterized by comprising:
[0007] Acquire audio information in the vehicle environment;
[0008] Performing sound recognition on the audio information to extract text information;
[0009] In response to the received first prompt instruction, constructing a tag to generate guidance information;
[0010] Generate guidance information based on the label, and generate multiple basic labels for the text information using a first preset large model;
[0011] In response to the received second prompt instruction, constructing label classification guidance information;
[0012] Based on the tag classification guidance information, hierarchical information of the multiple basic tags is generated using a second preset large model.
[0013] In an optional embodiment, the tag generation guidance information includes first task specification information, first constraint information and first example information;
[0014] The step of generating the guidance information based on the label and generating a plurality of basic labels for the text information using a first preset large model includes:
[0015] Based on the first task specification information, the first constraint information, and the first example information, a plurality of original labels are generated using the first preset large model; each of the original labels has a corresponding confidence score;
[0016] If the confidence score corresponding to the original label is greater than or equal to the first confidence threshold, the original label is determined to be the base label.
[0017] In an optional embodiment, after generating a plurality of original labels using the first preset large model based on the first task specification information, the first constraint condition information, and the first example information, the method further includes:
[0018] If the confidence score corresponding to the original label is less than the first confidence threshold but greater than or equal to the second confidence threshold, determine that the original label is an intermediate label, and in response to the received adjustment information, modify the intermediate label to the base label; or
[0019] If the confidence score corresponding to the original label is less than the second confidence threshold, the original label is determined to be an incorrect label.
[0020] In an optional embodiment, the audio information includes real-time audio information and offline audio information;
[0021] The step of generating the hierarchical information of the plurality of basic tags by using a second preset large model based on the tag classification guidance information includes:
[0022] If the audio information corresponding to the multiple basic tags is offline audio information, clustering the multiple basic tags based on density clustering to obtain multiple tag sets, and generating hierarchical information of the multiple tag sets using the second preset large model based on the tag classification guidance information; or
[0023] If the audio information corresponding to the multiple basic tags is real-time audio information, based on the tag classification guidance information, the second preset large model is used to generate hierarchical information of the multiple basic tags.
[0024] In an optional embodiment, the tag classification guidance information includes second task specification information, second constraint condition information, and second example information;
[0025] The step of generating the hierarchical information of the plurality of basic tags by using the second preset large model based on the tag classification guidance information includes:
[0026] If there is a target set matching the basic tag in the multiple tag sets, determining the hierarchical information of the basic tag based on the hierarchical information of the target set; or;
[0027] If there is no target set matching the basic tag in the multiple tag sets, hierarchical information of the basic tag is generated using the second preset large model based on the second task specification information, the second constraint information and the second example information.
[0028] In an optional embodiment, performing sound recognition on the audio information to extract text information includes:
[0029] Preprocessing the audio information to obtain preprocessed audio information; the preprocessing includes at least one of segmentation processing, noise reduction processing, and;
[0030] Performing sound recognition on the pre-processed audio information based on a sound recognition model to obtain initial text information;
[0031] The initial text information is cleaned to obtain text information; the cleaning process includes at least one of removing stop words, correcting spelling errors, and normalizing numbers.
[0032] In an optional embodiment, the method further includes:
[0033] Receive ambient audio information;
[0034] determining the basic tag of the ambient audio information and the hierarchical information corresponding to the basic tag;
[0035] If the hierarchical information is a safety fault, obtaining vehicle safety information; the vehicle safety information includes at least one of acceleration information, angular velocity information, rotation speed information, air pressure information, and inertial measurement information;
[0036] In a case where the vehicle safety information indicates that a safety hazard currently exists, a safety strategy is determined based on the basic label.
[0037] In a second aspect, an embodiment of the present application provides a hierarchical label generation device, the device comprising:
[0038] An acquisition module, used to acquire audio information in a vehicle environment;
[0039] A voice recognition module, used to perform voice recognition on the audio information and extract text information;
[0040] A first building module is configured to build a tag to generate guidance information in response to the received first prompt instruction;
[0041] a hierarchical label generation module, configured to generate guidance information based on the label, and generate a plurality of basic labels for the text information using a first preset large model;
[0042] A second building module is configured to build label classification guidance information in response to the received second prompt instruction;
[0043] A hierarchical generation module is used to generate hierarchical information of the multiple basic tags based on the tag classification guidance information using a second preset large model.
[0044] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the hierarchical label generation method of the first aspect.
[0045] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction or at least one program is stored, and the at least one instruction or at least one program is loaded and executed by a processor to implement the hierarchical label generation method of the first aspect.
[0046] In a fifth aspect, embodiments of the present application provide a computer program product or computer program, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the hierarchical tag generation method of the first aspect.
[0047] The hierarchical label generation method, device, electronic device, and storage medium provided in the embodiments of the present application have the following technical effects:
[0048] Acquire audio information in the vehicle environment; perform sound recognition on the audio information and extract text information; in response to the received first prompt instruction, construct tag generation guidance information; based on the tag generation guidance information, use the first preset large model to generate multiple basic tags for the text information; in response to the received second prompt instruction, construct tag classification guidance information; based on the tag classification guidance information, use the second preset large model to generate hierarchical information of the multiple basic tags. In the embodiment of the present application, by converting audio information into text information and generating multiple basic tags based on the first prompt instruction, it is possible to capture deeper semantic associations, guide tag classification based on the second prompt instruction, discover implicit relationships between tags, form a hierarchical structure between isolated tags, and establish a tag system with hierarchical relationships. This can better handle hierarchical relationships in context and improve the accuracy of tag application in personalized recommendations and vehicle safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0050] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;
[0051] Figure 2 This is a schematic diagram of a hierarchical label generation method provided in an embodiment of the present application. Figure 1 ;
[0052] Figure 3 This is a schematic diagram of a hierarchical label generation method provided in an embodiment of the present application. Figure 2 ;
[0053] Figure 4 This is a flow chart of a method for generating an original label provided in an embodiment of the present application;
[0054] Figure 5 This is a schematic diagram of a hierarchical label generation method provided in an embodiment of the present application. Figure 3 ;
[0055] Figure 6 This is a schematic diagram of the structure of a hierarchical label generation device provided in an embodiment of the present application;
[0056] Figure 7 This is a hardware structure block diagram of a server for a hierarchical label generation method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0058] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0059] See also Figure 1 , Figure 1 1 is a schematic diagram of an application environment provided in an embodiment of the present application, including an audio acquisition system 101 and an in-vehicle server 102.
[0060] In a possible embodiment, the audio acquisition system 101 is used to collect audio information in a vehicle environment. Specifically, it may include a vehicle-mounted storage module that stores offline audio information, such as songs, and may also include a communication interface connected to a cloud server to obtain audio information in the cloud server.
[0061] Furthermore, the audio collection system 101 may also include a sensor module for collecting real-time audio information in the vehicle environment, such as a microphone array provided in the cabin.
[0062] In a possible embodiment, the vehicle-mounted server 102 receives the audio information collected by the audio collection system 101, and is used to perform sound recognition on the audio information and extract text information; in response to the received first prompt instruction, constructs label generation guidance information; based on the label generation guidance information, uses the first preset large model to generate multiple basic labels for the text information; in response to the received second prompt instruction, constructs label classification guidance information; based on the label classification guidance information, uses the second preset large model to generate hierarchical information of the multiple basic labels. In the embodiment of the present application, by converting audio information into text information and generating multiple basic labels based on the first prompt instruction, it is possible to capture deeper semantic associations, guide label classification based on the second prompt instruction, discover implicit relationships between labels, form a hierarchical structure between isolated labels, and establish a label system with hierarchical relationships. It can better handle hierarchical relationships in context and improve the accuracy of label application in personalized recommendations and vehicle safety.
[0063] The following describes a specific embodiment of a hierarchical label generation method of the present application. Figure 2 This is a schematic diagram of a hierarchical label generation method provided in an embodiment of the present application. Figure 1 , this specification provides method operation steps such as embodiments or flow charts, but may include more or fewer operation steps based on routine or non-creative work. The order of steps listed in the embodiments is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in the order shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically, Figure 2 As shown, the method is applied to the vehicle-mounted server and may include:
[0064] S201: Acquire audio information in a vehicle environment.
[0065] S202: Perform sound recognition on the audio information and extract text information.
[0066] S203: In response to the received first prompt instruction, construct a tag to generate guidance information.
[0067] S204: Generate guidance information based on the label, and use a first preset large model to generate multiple basic labels for the text information.
[0068] S205: In response to the received first prompt instruction, construct tag classification guidance information.
[0069] S206: Based on the tag classification guidance information, generate hierarchical information of the multiple basic tags using a second preset large model.
[0070] Figure 3 This is a schematic diagram of a hierarchical label generation method provided in an embodiment of the present application. Figure 2 , the method may include:
[0071] S301: Acquire audio information in the vehicle environment.
[0072] In a possible embodiment, in terms of the source of the audio information, the audio information includes audio information obtained from a cloud server, audio information stored in a vehicle storage module, and audio information collected by sensors in real time.
[0073] In a possible embodiment, in terms of the content of the audio information, the audio information in the vehicle environment mainly includes various sounds in the interior and exterior environments of the vehicle.
[0074] Optional: internal sounds such as driver instructions, passenger conversations and voice broadcasts, other voice interaction information; media audio information such as played music, podcasts or audiobooks; mechanical sounds during vehicle operation (engine running sound, braking sound, gear shift prompt sound); fault alarms (abnormal tire pressure prompt, low battery warning, ABS activation prompt) and other vehicle status audio information; Bluetooth or Wi-Fi connection prompt sound, system assisted driving voice sound and other in-vehicle auxiliary equipment audio information.
[0075] Optionally, external sounds such as other vehicle sounds (horn sounds, engine roars, sudden brake sounds), traffic facility sounds, road condition sounds, rescue signals (fire truck / ambulance sirens, road construction warning sounds) and other traffic environment audio information, weather sounds, pedestrian and road user sounds and other natural environment audio information.
[0076] By classifying and analyzing the above-mentioned various sound information, various environmental sounds during vehicle driving can be labeled and hierarchically managed. The established hierarchical labeling system can provide users with more accurate personalized services based on labels. It can also enhance the safety of intelligent driving by identifying sounds with safety hazards, providing fault location prompts and adjusting driving strategies, and providing a comprehensive technical foundation for in-vehicle voice recognition, environmental perception and intelligent decision-making.
[0077] S302: Perform sound recognition on the audio information and extract text information.
[0078] In a possible embodiment, performing sound recognition on the audio information to extract text information includes:
[0079] S3021: Preprocess the audio information to obtain preprocessed audio information.
[0080] In a possible embodiment, the preprocessing includes at least one of segmenting processing, noise reduction processing, and so on. By preprocessing the audio information, the audio quality can be improved, noise interference can be reduced, and it can be adapted to the subsequent processing model.
[0081] Among them, the segmenting processing can specifically be to split a continuous audio stream into segments of a fixed duration (such as 10 seconds) or event-driven segments (such as speech pause intervals). This avoids the model processing timeout caused by long audio and facilitates parallel computing. The noise reduction processing can be to analyze the noise spectrum through FFT and cancel the background noise in reverse, or it can be based on beamforming technology to directionally enhance the target sound source (such as the driver's speech) through a microphone array and suppress the noise in other directions. It can also be based on deep learning noise reduction, for example, using models such as SEGAN to separate human voices and noise.
[0082] S3022: Perform voice recognition on the preprocessed audio information based on a voice recognition model to obtain initial text information.
[0083] In the embodiment of the present application, automatic speech recognition is performed on the preprocessed audio information based on the locally built ASR model, and the initial text information is cached.
[0084] S3023: Perform cleaning processing on the initial text information to obtain text information.
[0085] In a possible embodiment, the cleaning processing includes at least one of stop word removal, spelling correction, and digital standardization. By cleaning the initial text information, the text quality can be optimized, which is convenient for subsequent semantic analysis processing.
[0086] Among them, stop word removal is based on a stop word list in the transportation field, such as "um", "ah", "then", etc., to remove redundant words. It can effectively reduce the text length and retain key information. Spelling correction can match common spelling mistakes and correct semantic contradictions, and a spelling correction rule library can also be established, especially a correction table for transportation-specific nouns. Digital standardization includes unifying the time format (such as converting "3 pm" to "15:00") and converting digital units (such as converting "three kilometers" to "3 km").
[0087] S303: In response to the received first prompt instruction, construct label generation guiding information.
[0088] In a possible embodiment, the label generation guiding information includes first task specification information, first constraint condition information, and first example information.
[0089] In an embodiment of the present application, the first task specification information, the first constraint information and the first example information constitute a three-level prompt word structure. Through the first task specification information, the model is explicitly required to perform "multi-dimensional label generation". The number of labels generated (for example, 3), the output format (for example, Json structure) and the output results (label text and label confidence score) are specified through the first constraint information specification. Finally, the first example information is used as a few-shot layer to inject vehicle-borne domain knowledge and provide some sample inputs and outputs.
[0090] For example, the first task specification requires the model to generate multi-dimensional labels based on text information, covering the three dimensions of scene, entity, and emotion, and reflecting semantic relevance. The first constraint requirement requires the generated labels to be three, with the output format being JSON. The output results include the label text and the label confidence score. The first example information might be: the audio information is a vehicle voice reminder message, "The battery is currently low and you need to find a charging station." The output labels are "battery" (0.98), "urgent" (0.75), and "navigation" (0.62). The audio information is a playback of the media message "Fold a thousand paper cranes and tie a red belt." The output labels are "festive" (0.88), "Spring Festival" (0.90), and "thousand paper cranes" (0.96).
[0091] S304: Generate guidance information based on the label, and use a first preset large model to generate multiple basic labels for the text information.
[0092] Figure 4 is a flow chart of a method for generating original tags provided by an embodiment of the present application. In one possible embodiment, the method of generating guiding information based on the tags and using a first preset large model to generate multiple basic tags for the text information includes:
[0093] S3041: Based on the first task specification information, the first constraint information and the first example information, generate a plurality of original labels using the first preset large model.
[0094] In the embodiment of the present application, each of the original labels has a corresponding confidence score.
[0095] S3042: Determine whether the confidence score is greater than or equal to the first confidence threshold. If so, execute S3043; if not, execute S3044.
[0096] S3043: Determine the original tag as the base tag.
[0097] S3044: Determine whether the confidence score is greater than or equal to the second confidence threshold. If so, execute S3045; if not, execute S3047.
[0098] S3045: Determine that the original label is an intermediate label.
[0099] S3046: In response to the received adjustment information, modify the intermediate label to the basic label.
[0100] S3047: Determine that the original label is an erroneous label.
[0101] In an embodiment of the present application, if the confidence score corresponding to the original label is greater than or equal to a first confidence threshold, the original label is determined to be a base label.
[0102] Optionally, if the confidence score corresponding to the original label is less than the first confidence threshold but greater than or equal to the second confidence threshold, the original label is determined to be an intermediate label, and in response to the received adjustment information, the intermediate label is corrected to the basic label.
[0103] Optionally, if the confidence score corresponding to the original label is less than the second confidence threshold, the original label is determined to be an erroneous label.
[0104] Through the first confidence threshold, high-confidence basic labels are screened out and directly used as original labels. The second confidence threshold is used to distinguish intermediate labels from erroneous labels. Intermediate labels can be adjusted and corrected to correct labels through manual or other feedback adjustments, and also used as original labels to filter out erroneous labels. Through the dual-threshold grading mechanism and correction process, accurate label stratification can be achieved. Basic labels directly support business, and dynamic repair of intermediate labels improves availability, significantly improving the efficiency and reliability of label generation in scenarios such as in-vehicle voice interaction and intelligent driving decision-making.
[0105] S305: In response to the received second prompt instruction, construct tag classification guidance information.
[0106] In a possible embodiment, the tag classification guidance information includes second task specification information, second constraint condition information, and second example information.
[0107] Similar to the label generation guidance information, the second task specification information, the second constraint information and the second example information also constitute a three-level prompt word structure.
[0108] The second task specification information explicitly requires the model to "build a hierarchical system for basic labels." The second constraint information specifies the number of layers (for example, three), the output format (for example, a JSON structure), and the output results (layer text and category names). Finally, the second example information is used as a few-shot layer to specify some desired category names (e.g., easy-to-understand label names such as military and politics), and provide some sample inputs and outputs.
[0109] S306: Based on the tag classification guidance information, generate hierarchical information of the multiple basic tags using a second preset large model.
[0110] In the embodiment of the present application, the audio information includes real-time audio information and offline audio information. For offline audio information and real-time audio information, the hierarchical process is different.
[0111] In a possible embodiment, generating the hierarchical information of the plurality of basic tags using a second preset large model based on the tag classification guidance information includes:
[0112] S3061: If the audio information corresponding to the multiple basic tags is offline audio information, cluster the multiple basic tags based on density clustering to obtain multiple tag sets.
[0113] In a possible embodiment, multiple basic tags of offline audio information are clustered using a density-based clustering (DBSCAN) method, and semantic clusters are automatically discovered in a data-driven manner, avoiding the need to preset a number of categories, and obtaining multiple tag sets.
[0114] S3062: Based on the tag classification guidance information, use the second preset large model to generate hierarchical information of the multiple tag sets.
[0115] Then, the label classification guidance information is used to generate the hierarchical information of the plurality of label sets using the second preset large model.
[0116] In the embodiment of the present application, the first preset large model and the second preset large model may be the same or different.
[0117] In a possible embodiment, the first preset large model responsible for basic label generation may be Conformer-Tiny, which combines a self-attention mechanism (Transformer) with a convolutional neural network (CNN) to generate multi-dimensional basic labels based on label generation guidance information.
[0118] In a possible embodiment, the second preset large model responsible for hierarchical label generation can be GPT-3.5-turbo, which is based on the Transformer decoder architecture and generates a tree-like label system according to label classification guidance information.
[0119] S3063: If the audio information corresponding to the multiple basic tags is real-time audio information, based on the tag classification guidance information, the second preset large model is used to generate hierarchical information of the multiple basic tags.
[0120] In a possible embodiment, after the offline audio information is divided into categories, multiple basic tags of the real-time audio information can be directly hierarchically set.
[0121] In a possible embodiment, generating the hierarchical information of the plurality of basic tags using the second preset large model based on the tag classification guidance information includes:
[0122] Optionally, if there is a target set matching the basic tag in the multiple tag sets, the hierarchical information of the basic tag is determined based on the hierarchical information of the target set.
[0123] Optionally, if there is no target set matching the basic tag in the multiple tag sets, the hierarchical information of the basic tag is generated using the second preset large model based on the second task specification information, the second constraint information and the second example information.
[0124] When matching hierarchical information exists, the target set's hierarchical information is directly used to avoid repeated calculations and speed up processing. In the absence of matching hierarchical information, a new hierarchical structure can be autonomously generated to adapt to newly emerging tags or scenarios.
[0125] The basic tags and hierarchical information generated through the above steps are stored in the database, the tag system is updated in real time, and applied in the search and recommendation fields to improve the accuracy of tag application in personalized recommendation and search fields.
[0126] Figure 5 This is a schematic diagram of a hierarchical label generation method provided in an embodiment of the present application. Figure 3 , the method may include:
[0127] S401: Receive ambient audio information.
[0128] S402: Determine the basic tag of the ambient audio information and the hierarchical information corresponding to the basic tag.
[0129] S403: If the hierarchical information is a safety fault, obtain vehicle safety information.
[0130] In an optional embodiment, the vehicle safety information includes at least one of acceleration information, angular velocity information, rotation speed information, air pressure information and inertial measurement information.
[0131] S404: When the vehicle safety information indicates that a safety hazard currently exists, determine a safety strategy based on the basic tag.
[0132] In one possible embodiment, a vehicle microphone captures ambient audio information, including the voice command "emergency brake" and the sound of metal impact. The basic tags and hierarchical information corresponding to the voice command and the sound of metal impact are then determined. For example, the basic tags and hierarchical information corresponding to "emergency brake" are "accident → collision → brake".
[0133] Safety failure is a hierarchical information category that includes various categories related to vehicle safety failures, such as accidents, extreme weather, and traffic incidents. If the hierarchical information of "Emergency Braking" corresponds to the safety failure category, obtain the current vehicle safety information.
[0134] In the embodiment of the present application, the Z-axis acceleration peak is obtained by the accelerometer, the lateral angular velocity is obtained by the gyroscope, the left front wheel speed is obtained by the wheel speed sensor, and the left front tire pressure is collected by the barometer. By integrating various vehicle safety information, it is possible to comprehensively determine whether there are safety hazards at present.
[0135] Finally, if vehicle safety information indicates a potential safety hazard, a safety strategy is developed based on the basic tags. For example, if the accelerometer's Z-axis acceleration exceeds the threshold, a severe collision is detected; and if the wheel speed sensor's left front wheel speed is abnormal, with a specific differential of 35%, a suspected tire blowout is suspected, thus determining a potential safety hazard.
[0136] Generate an automatic braking safety strategy for the basic label "brake" of the above "emergency brake".
[0137] Through the basic tags of the vehicle's ambient audio information and the hierarchical information corresponding to the basic tags, vehicle safety information is actively obtained, and active safety strategies are formulated to improve driving safety.
[0138] The present application also provides a hierarchical label generation device. Figure 6 This is a schematic diagram of the structure of a hierarchical label generation device provided in an embodiment of the present application. Figure 6 As shown, the apparatus 500 includes:
[0139] An acquisition module 501 is used to acquire audio information in a vehicle environment;
[0140] A voice recognition module 502 is used to perform voice recognition on the audio information and extract text information;
[0141] A first constructing module 503 is configured to construct a tag generation guide information in response to the received first prompt instruction;
[0142] A label generation module 504 is configured to generate guidance information based on the label, and generate a plurality of basic labels for the text information using a first preset large model;
[0143] The second constructing module 505 is configured to construct tag classification guidance information in response to the received first prompt instruction;
[0144] The hierarchy generation module 506 is configured to generate hierarchy information of the plurality of basic tags using a second preset large model based on the tag classification guidance information.
[0145] In an optional implementation, the tag generation guidance information includes first task specification information, first constraint information, and first example information; and further includes:
[0146] a first label generation module, configured to generate a plurality of original labels using the first preset large model based on the first task specification information, the first constraint information, and the first example information; each of the original labels having a corresponding confidence score;
[0147] A first determining module is configured to determine the original label as a base label if the confidence score corresponding to the original label is greater than or equal to a first confidence threshold.
[0148] In an optional embodiment, the method further includes:
[0149] a second determining module, configured to determine that the original label is an intermediate label if the confidence score corresponding to the original label is less than the first confidence threshold but greater than or equal to a second confidence threshold, and, in response to the received adjustment information, modify the intermediate label to the base label; or
[0150] The third determination module is configured to determine that the original label is an incorrect label if the confidence score corresponding to the original label is less than the second confidence threshold.
[0151] In an optional implementation manner, the audio information includes real-time audio information and offline audio information; and further includes:
[0152] A first hierarchical generation module is configured to, if the audio information corresponding to the multiple basic tags is offline audio information, cluster the multiple basic tags based on density clustering to obtain multiple tag sets, and generate hierarchical information of the multiple tag sets using the second preset large model based on the tag classification guidance information; or
[0153] The second hierarchical generation module is used to generate hierarchical information of the multiple basic tags using the second preset large model based on the tag classification guidance information if the audio information corresponding to the multiple basic tags is real-time audio information.
[0154] In an optional implementation, the tag classification guidance information includes second task specification information, second constraint condition information, and second example information; and further includes:
[0155] a fourth determining module, configured to determine the hierarchical information of the basic tag based on the hierarchical information of the target set if there is a target set matching the basic tag in the multiple tag sets; or;
[0156] The third hierarchical generation module is used to generate hierarchical information of the basic label using the second preset large model based on the second task specification information, the second constraint information and the second example information if there is no target set matching the basic label in the multiple label sets.
[0157] In an optional implementation, performing sound recognition on the audio information and extracting text information includes:
[0158] A preprocessing module, configured to preprocess the audio information to obtain preprocessed audio information; the preprocessing includes at least one of segmentation processing, noise reduction processing, and;
[0159] A sound recognition module, configured to perform sound recognition on the pre-processed audio information based on a sound recognition model to obtain initial text information;
[0160] A cleaning module is used to perform a cleaning process on the initial text information to obtain text information; the cleaning process includes at least one of removing stop words, correcting spelling errors, and normalizing numbers.
[0161] In an optional embodiment, the method further includes:
[0162] A receiving module, configured to receive ambient audio information;
[0163] a fifth determining module, configured to determine the basic tag of the ambient audio information and the hierarchical information corresponding to the basic tag;
[0164] A first acquisition module is configured to acquire vehicle safety information if the hierarchical information indicates a safety fault; the vehicle safety information includes at least one of acceleration information, angular velocity information, rotation speed information, air pressure information, and inertial measurement information;
[0165] The sixth determination module is configured to determine a safety strategy based on the basic label when the vehicle safety information indicates that a safety hazard currently exists.
[0166] The device and method embodiments in the embodiments of this application are based on the same application concept.
[0167] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 7This is a hardware structure diagram of a server for a hierarchical tag generation method provided by an embodiment of the present application. Figure 7 As shown, the server 600 may vary significantly depending on its configuration or performance. It may include one or more central processing units (CPUs) 610 (processor 610 may include, but is not limited to, a microprocessor (MCU) or a processing device such as a programmable logic device (FPGA), a memory 630 for storing data, and one or more storage media 620 (e.g., one or more mass storage devices) for storing application programs 623 or data 622. The memory 630 and storage media 620 may be either transient or persistent storage. The program stored in the storage medium 620 may include one or more modules, each of which may include a series of instruction operations on the server. Furthermore, the CPU 610 may be configured to communicate with the storage medium 620 to execute the series of instruction operations in the storage medium 620 on the server 600. The server 600 may also include one or more power supplies 660, one or more wired or wireless network interfaces 650, one or more input and output interfaces 640, and / or one or more operating systems 621, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0168] The input / output interface 640 can be used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the server 600. In one embodiment, the input / output interface 640 may include a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the input / output interface 640 may be a radio frequency (RF) module for wireless communication with the Internet.
[0169] It can be understood by those skilled in the art that Figure 7 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 7 More or fewer components than shown, or with Figure 7 Different configurations shown.
[0170] An embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the above-mentioned data processing method.
[0171] An embodiment of the present application also provides a computer-readable storage medium, which can be set in a server to store at least one instruction, at least one program, code set or instruction set related to a hierarchical label generation method in an embodiment of the method. The at least one instruction, the at least one program, the code set or instruction set is loaded and executed by the processor to implement the above-mentioned hierarchical label generation method.
[0172] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers in the computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard drive, a magnetic disk, or an optical disk, among other media capable of storing program code.
[0173] It can be seen from the embodiments of the hierarchical label generation method, device, electronic device or storage medium provided by the present application that the present application obtains audio information in a vehicle environment; performs sound recognition on the audio information to extract text information; constructs label generation guidance information in response to the received first prompt instruction; based on the label generation guidance information, generates multiple basic labels for the text information using a first preset large model; in response to the received second prompt instruction, constructs label classification guidance information; based on the label classification guidance information, generates hierarchical information of the multiple basic labels using a second preset large model. In the embodiment of the present application, by converting audio information into text information and generating multiple basic labels based on the first prompt instruction, it is possible to capture deeper semantic associations, guide label classification based on the second prompt instruction, discover implicit relationships between labels, form a hierarchical structure between isolated labels, and establish a label system with hierarchical relationships. It can better handle hierarchical relationships in context and improve the accuracy of label application in personalized recommendations and vehicle safety.
[0174] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0175] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0176] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0177] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A hierarchical label generation method, characterized in that: include: Acquire audio information in the vehicle environment; Performing sound recognition on the audio information to extract text information; In response to the received first prompt instruction, constructing a tag to generate guidance information; The tag generation guidance information includes first task specification information, first constraint condition information and first example information; Generate guidance information based on the label, and generate multiple basic labels for the text information using a first preset large model; In response to the received second prompt instruction, constructing label classification guidance information; The tag classification guidance information includes second task specification information, second constraint condition information and second example information; Based on the tag classification guidance information, generating hierarchical information of the plurality of basic tags using a second preset large model; Receive ambient audio information; determining the basic tag of the ambient audio information and the hierarchical information corresponding to the basic tag; If the level information is a safety fault category, obtaining vehicle safety information; The vehicle safety information includes at least one of acceleration information, angular velocity information, rotation speed information, air pressure information and inertial measurement information; When the vehicle safety information indicates that a safety hazard currently exists, determining a safety strategy based on the basic tag; The audio information includes real-time audio information and offline audio information; The step of generating the hierarchical information of the plurality of basic tags by using a second preset large model based on the tag classification guidance information includes: If the audio information corresponding to the multiple basic tags is offline audio information, clustering the multiple basic tags based on density clustering to obtain multiple tag sets, and generating hierarchical information of the multiple tag sets using the second preset large model based on the tag classification guidance information; If the audio information corresponding to the multiple basic tags is real-time audio information, and if there is a target set matching the basic tag in the multiple tag sets, the hierarchical information of the basic tag is determined based on the hierarchical information of the target set; or; if the audio information corresponding to the multiple basic tags is real-time audio information, and if there is no target set matching the basic tag in the multiple tag sets, the hierarchical information of the basic tag is generated using the second preset large model based on the second task specification information, the second constraint condition information and the second example information.
2. A hierarchical label generation method according to claim 1, characterized in that: The tag generation guidance information includes first task specification information, first constraint condition information and first example information; The step of generating the guidance information based on the label and generating a plurality of basic labels for the text information using a first preset large model includes: Based on the first task specification information, the first constraint information, and the first example information, a plurality of original labels are generated using the first preset large model; each of the original labels has a corresponding confidence score; If the confidence score corresponding to the original label is greater than or equal to the first confidence threshold, the original label is determined to be the base label.
3. A hierarchical label generation method according to claim 2, characterized in that: After generating a plurality of original labels using the first preset large model based on the first task specification information, the first constraint condition information, and the first example information, the method further includes: If the confidence score corresponding to the original label is less than the first confidence threshold but greater than or equal to the second confidence threshold, determine that the original label is an intermediate label, and in response to the received adjustment information, modify the intermediate label to the base label; or If the confidence score corresponding to the original label is less than the second confidence threshold, the original label is determined to be an incorrect label.
4. The hierarchical label generation method according to claim 1, wherein: The performing sound recognition on the audio information and extracting text information includes: Preprocessing the audio information to obtain preprocessed audio information; the preprocessing includes at least one of segmentation processing, noise reduction processing, and; Performing sound recognition on the pre-processed audio information based on a sound recognition model to obtain initial text information; The initial text information is cleaned to obtain text information; the cleaning process includes at least one of removing stop words, correcting spelling errors, and normalizing numbers.
5. A hierarchical label generation device, characterized in that: include: An acquisition module, used to acquire audio information in a vehicle environment; A voice recognition module, used to perform voice recognition on the audio information and extract text information; A first building module is configured to build a tag to generate guidance information in response to the received first prompt instruction; The tag generation guidance information includes first task specification information, first constraint condition information and first example information; a hierarchical label generation module, configured to generate guidance information based on the label, and generate a plurality of basic labels for the text information using a first preset large model; A second building module is configured to build label classification guidance information in response to the received second prompt instruction; The tag classification guidance information includes second task specification information, second constraint condition information and second example information; A hierarchy generation module, configured to generate hierarchy information of the plurality of basic tags using a second preset large model based on the tag classification guidance information; Receive ambient audio information; determining the basic tag of the ambient audio information and the hierarchical information corresponding to the basic tag; If the level information is a safety fault category, obtaining vehicle safety information; The vehicle safety information includes at least one of acceleration information, angular velocity information, rotation speed information, air pressure information and inertial measurement information; When the vehicle safety information indicates that a safety hazard currently exists, determining a safety strategy based on the basic tag; The audio information includes real-time audio information and offline audio information; The step of generating the hierarchical information of the plurality of basic tags by using a second preset large model based on the tag classification guidance information includes: If the audio information corresponding to the multiple basic tags is offline audio information, clustering the multiple basic tags based on density clustering to obtain multiple tag sets, and generating hierarchical information of the multiple tag sets using the second preset large model based on the tag classification guidance information; If the audio information corresponding to the multiple basic tags is real-time audio information, and if there is a target set matching the basic tag in the multiple tag sets, the hierarchical information of the basic tag is determined based on the hierarchical information of the target set; or; if the audio information corresponding to the multiple basic tags is real-time audio information, and if there is no target set matching the basic tag in the multiple tag sets, the hierarchical information of the basic tag is generated using the second preset large model based on the second task specification information, the second constraint condition information and the second example information.
6. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the hierarchical label generation method as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the hierarchical label generation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Intelligent information acquisition method
CN116798235A
Text classification model training method and device, electronic equipment and storage medium
CN117493557A
Text classification method, device and system based on large language model, storage medium and product
CN119647408A