Cold start method and device based on large model, electronic equipment and storage medium
By extracting the user's driving scenario, emotions and type characteristics in the on-board system, and using the large model to match the audio data at the label level, the problems of zero-sample cold start and insufficient semantic understanding are solved, and accurate personalized audio push is achieved.
Patent Information
- Application Number
- CN202510717980.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In vehicle-mounted scenarios, the prior art is difficult to achieve zero-sample cold start, and the semantic understanding is insufficient, resulting in the recommendation results being biased towards surface statistical features rather than deep semantic matching, and it is impossible to accurately push personalized audio content to users.
By obtaining the user's voice command information and vehicle driving information, the user's driving scenario, emotional characteristics, type characteristics and content tendency characteristics are extracted, and the large model matches the audio data carrying the scene layer, user layer and content layer tags to achieve zero-sample cold start.
In the absence of user historical data and privacy data, the precise push of personalized audio content improves the accuracy and semantic understanding of recommendations.
Smart Images

Figure CN120256668A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of intelligent cockpit systems and recommendation algorithms, and particularly to a cold start method, device, electronic device, and storage medium based on a large model. Background Art
[0002] With the popularization of intelligent vehicle systems, personalized recommendation has become one of the core technologies to enhance the user experience. However, in the in-vehicle scenario, the interaction data between users and audio content often shows a high degree of sparsity. Especially when users log in for the first time or switch scenarios (such as when new users use for the first time), the lack of sufficient interaction records causes traditional recommendation models to fail.
[0003] Existing methods mainly rely on Collaborative Filtering (CF) and Graph Neural Networks (GNNs) for cold start optimization. On the one hand, the hybrid collaborative filtering model alleviates the sparsity problem of the interaction matrix by fusing user portraits (such as age, gender) and audio content features (such as metadata, label word frequency), but its essence still relies on historical interaction data and it is difficult to achieve "zero-shot" cold start - that is, when there is no historical behavior of the user or content, the model cannot generate effective recommendations.
[0004] In addition, there are significant deficiencies in the understanding of semantics in existing technologies. Traditional content features are difficult to capture complex semantic associations, resulting in recommendation results being biased towards surface statistical features rather than deep semantic matching. In the in-vehicle scenario, the dynamic driving environment and fragmented user behavior further exacerbate the data sparsity. Therefore, it is difficult to accurately push content to users during the zero-shot cold start process. Summary of the Invention
[0005] To solve the technical problems that the existing data push methods cannot solve zero-shot cold start and have insufficient semantic understanding ability, the present invention provides a cold start method, device, electronic device, and storage medium based on a large model. By extracting the user's current driving scenario, user emotion characteristics, user type characteristics, and content tendency characteristics from the user's voice command information and vehicle driving information, and matching them with the third audio data carrying corresponding scene layer labels, user layer labels, and content layer labels, the fourth audio data that the user may be interested in is recalled, realizing zero-shot cold start without the need for the user's historical data and user privacy data, and accurately pushing personalized audio for the user. At the same time, through multiple hierarchical label levels with hierarchical relationships, support is provided for semantic understanding, making the pushed audio more accurate.
[0006] In a first aspect, an embodiment of the present application provides a cold start method based on a large model, and the method includes: Obtain the voice command information of the user and the vehicle driving information; the vehicle driving information includes the current location information and the current time information; Analyze the voice command information to obtain the user emotion feature and the user type feature; Analyze the vehicle driving information to determine the corresponding driving scenario; Obtain the first audio data and the second audio data; the first audio data is the popular audio associated with the current time information; the second audio data is the popular audio associated with the current location information; Analyze the first audio data and the second audio data to obtain the content tendency feature; Obtain the third audio data; the third audio data carries the corresponding scenario layer label, user layer label and content layer label; there is a hierarchical relationship between the scenario layer label, the user layer label and the content layer label; Based on the driving scenario, the user emotion feature, the user type feature and the content tendency feature, recall the fourth audio data from the third audio data.
[0007] In an alternative embodiment, before recalling the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature and the content tendency feature, it further includes: Perform content recognition on the third audio data and convert it into third text data; In response to the received first structured prompt word, extract multiple basic labels of the third text data based on a preset large model; In response to the received second structured prompt word, construct a hierarchical relationship based on the preset large model for the multiple basic labels to obtain the scenario layer label, the user layer label and the content layer label of the third audio data.
[0008] In an alternative embodiment, recalling the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature and the content tendency feature includes: If the driving scenario matches the scenario layer label, determine the first similarity based on the user emotion feature, the user type feature and the user layer label; When the first similarity is greater than the preset similarity threshold, determine the second similarity based on the content tendency feature and the content layer label; Recall the fourth audio data from the third audio data based on the second similarity.
[0009] In an alternative embodiment, the scenario layer label has a corresponding information density; If the driving scenario matches the scenario layer label, determining a first similarity based on the user emotion feature, the user type feature, and the user layer label includes: Based on the driving scenario and the information density, determining a safety factor based on the preset large model; If the safety factor is greater than a preset safety factor threshold, determining that the driving scenario matches the scenario layer label; Determining the first similarity based on the user emotion feature, the user type feature, and the user layer label.
[0010] In an alternative embodiment, analyzing the voice command information to obtain user emotion features and user type features includes: Performing intonation analysis on the voice command information to extract the user emotion feature; Performing timbre analysis on the voice command information to extract the user type feature; the user type feature includes age feature and gender feature.
[0011] In an alternative embodiment, the vehicle driving information further includes driving state information and driving environment information; Analyzing the vehicle driving information to determine a corresponding driving scenario includes: If the driving state information and the driving environment information indicate that the current driving state is safe, determining the information demand scenario as the corresponding driving scenario; or; If the driving state information and the driving environment information indicate that current emotion guidance is needed, determining the emotion demand scenario as the corresponding driving scenario.
[0012] In an alternative embodiment, it further includes: Obtaining updated fifth audio data; Determining the scenario layer label, the user layer label, and the content layer label corresponding to the fifth audio data; Obtaining a plurality of zero-shot users; the plurality of zero-shot users respectively carry the corresponding driving scenario, user emotion feature, user type feature, and content tendency feature; Recalling a plurality of users to be shared from the plurality of zero-shot users based on the scenario layer label, the user layer label, and the content layer label.
[0013] In a second aspect, an embodiment of the present application provides a cold start device based on a large model, and the device includes: A first acquisition module, configured to acquire voice command information of a user and vehicle driving information; the vehicle driving information includes current location information and current time information; A first analysis module, configured to analyze the voice command information to obtain user emotion characteristics and user type characteristics; A second analysis module, configured to analyze the vehicle driving information to determine a corresponding driving scenario; A second acquisition module, configured to acquire first audio data and second audio data; the first audio data is popular audio associated with the current time information; the second audio data is popular audio associated with the current location information; A third analysis module, configured to analyze the first audio data and the second audio data to obtain content tendency characteristics; A third acquisition module, configured to acquire third audio data; the third audio data carries corresponding scenario layer tags, user layer tags, and content layer tags; there is a hierarchical relationship among the scenario layer tags, the user layer tags, and the content layer tags; A first recall module, configured to recall fourth audio data from the third audio data based on the driving scenario, the user emotion characteristics, the user type characteristics, and the content tendency characteristics.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the large model-based cold start method in the first aspect.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction or at least one program is stored, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the large model-based cold start method in the first aspect.
[0016] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the large model-based cold start method in the first aspect.
[0017] The large model-based cold start method, device, electronic device, and storage medium provided by the embodiments of the present application have the following technical effects: Obtain the voice command information and vehicle driving information of the user; the vehicle driving information includes the current location information and the current time information; analyze the voice command information to obtain the user emotion feature and the user type feature; analyze the vehicle driving information to determine the corresponding driving scenario; obtain the first audio data and the second audio data; the first audio data is the popular audio associated with the current time information; the second audio data is the popular audio associated with the current location information; analyze the first audio data and the second audio data to obtain the content tendency feature; obtain the third audio data; the third audio data carries the corresponding scenario layer label, user layer label and content layer label; there is a hierarchical relationship between the scenario layer label, the user layer label and the content layer label; recall the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature and the content tendency feature. In the embodiment of the present application, the current driving scenario, user emotion feature, user type feature and content tendency feature of the user are extracted through the user's voice command information and vehicle driving information, and are matched with the third audio data carrying the corresponding scenario layer label, user layer label and content layer label, and the fourth audio data that the user may be interested in is recalled. Without the need for the user's historical data and user privacy data, zero-sample cold start is achieved, personalized audio is accurately pushed for the user. At the same time, through the multi-layer label hierarchy with a hierarchical relationship, support is provided for semantic understanding, making the pushed audio more accurate. Brief Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the drawings required to be used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application; Figure 2 is a flowchart of a cold start method based on a large model provided by an embodiment of the present application Figure 1 ; Figure 3 is a flowchart of a cold start method based on a large model provided by an embodiment of the present application Figure 2 ; Figure 4 is a flowchart of a cold start method based on a large model provided by an embodiment of the present application Figure 3 ; Figure 5It is a schematic structural diagram of a cold start device based on a large model provided by an embodiment of the present application; Figure 6 It is a hardware structure block diagram of a server for a cold start method based on a large model provided by an embodiment of the present application. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0022] Please refer to Figure 1 , Figure 1 It is a schematic diagram of an application environment provided by an embodiment of the present application, including an information collection device 101 and a server 102.
[0023] In a possible embodiment, the information collection device 101 is used to collect the voice command information and vehicle driving information of the user, and specifically includes a voice collection device and a vehicle information collection device.
[0024] In the embodiment of the present application, the voice collection device may include a microphone array disposed inside the vehicle. By arranging 4-8 microphones, sound source localization, noise suppression, and echo cancellation can be achieved, improving the clarity of the voice signal.
[0025] In the embodiment of the present application, the vehicle information collection device may include those disposed in the vehicle Positioning and navigation sensors, such as the Global Positioning System (GPS) that can obtain the vehicle's position, speed, and heading in real time, and the Inertial Measurement Unit (IMU) that detects the vehicle's acceleration, angular velocity, and attitude changes through accelerometers and gyroscopes; it can also include vehicle state sensors, such as wheel speed sensors that monitor the wheel speed, calculate the driving mileage and vehicle speed, pressure sensors that detect the tire pressure in real time and monitor the status of the hydraulic system, and temperature sensors that monitor the engine coolant temperature and the temperature inside the vehicle air conditioner; it can also include driving behavior sensors, such as a steering wheel angle sensor that detects the steering angle and rate, or a pedal travel sensor that collects the position of the accelerator / brake pedal: it can also include cameras installed inside and outside the vehicle to detect the images inside and around the vehicle.
[0026] In an embodiment of the present application, the information acquisition device 101 may further include an interface connected to a local storage device or a cloud storage device for obtaining audio data from the local storage device or the cloud storage device.
[0027] In a possible embodiment, obtain the user's voice command information and vehicle driving information; the vehicle driving information includes the current position information and the current time information; analyze the voice command information to obtain the user's emotion characteristics and user type characteristics; analyze the vehicle driving information to determine the corresponding driving scenario; obtain the first audio data and the second audio data; the first audio data is the popular audio associated with the current time information; the second audio data is the popular audio associated with the current position information; analyze the first audio data and the second audio data to obtain the content tendency characteristics; obtain the third audio data; the third audio data carries the corresponding scenario layer label, user layer label, and content layer label; there is a hierarchical relationship between the scenario layer label, the user layer label, and the content layer label; based on the driving scenario, the user's emotion characteristics, the user type characteristics, and the content tendency characteristics, recall the fourth audio data from the third audio data. In an embodiment of the present application, the user's current driving scenario, user emotion characteristics, user type characteristics, and content tendency characteristics are extracted through the user's voice command information and vehicle driving information, and are matched with the third audio data carrying the corresponding scenario layer label, user layer label, and content layer label, and the fourth audio data that the user may be interested in is recalled, realizing zero-sample cold start without the need for the user's historical data and user privacy data, and accurately pushing personalized audio for the user.
[0028] The following introduces a specific embodiment of a cold start method based on a large model in the present application. Figure 2 It is a flowchart of a cold start method based on a large model provided by an embodiment of the present application. Figure 1, this specification provides method operation steps such as in the embodiments or flowcharts, but may include more or fewer operation steps based on routine or non-creative labor. The order of steps listed in the embodiments is just one way among many execution orders of steps and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, when this method is applied to a server, it may include: S201: Obtain the voice instruction information of the user and the vehicle driving information; the vehicle driving information includes the current location information and the current time information.
[0029] S202: Analyze the voice instruction information to obtain the user emotion feature and the user type feature.
[0030] S203: Analyze the vehicle driving information to determine the corresponding driving scenario.
[0031] S204: Obtain the first audio data and the second audio data; the first audio data is the popular audio associated with the current time information; the second audio data is the popular audio associated with the current location information.
[0032] S205: Analyze the first audio data and the second audio data to obtain the content tendency feature.
[0033] S206: Obtain the third audio data; the third audio data carries the corresponding scenario layer label, user layer label, and content layer label; there is a hierarchical relationship among the scenario layer label, the user layer label, and the content layer label.
[0034] S207: Recall the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature, and the content tendency feature.
[0035] Figure 3 is the flowchart illustration of a cold start method based on a large model provided by an embodiment of the present application Figure 2 , this method may include: S301: Obtain the voice instruction information of the user and the vehicle driving information.
[0036] In a possible embodiment, the vehicle driving information includes the current location information and the current time information. Specifically, the voice instruction information is obtained through a microphone, the current location information is obtained through the global positioning system set in the vehicle, and the current time information is synchronized through a time synchronization module.
[0037] S302: Analyze the voice command information to obtain the user's emotional characteristics and user type characteristics.
[0038] In a possible embodiment, the analyzing the voice command information to obtain the user's emotional characteristics and user type characteristics includes: S3021: Conduct intonation analysis on the voice command information to extract the user's emotional characteristics.
[0039] In a possible embodiment, based on a large model, conduct intonation analysis on the received voice command information.
[0040] First, perform voice preprocessing on the obtained voice command information. Specifically, it can be to use the RNN-Turbo model to separate the human voice from background noise (such as the in-vehicle air conditioner sound, navigation announcements).
[0041] Then segment the continuous speech into short frames of 20 - 40 ms, and extract time-domain / frequency-domain features. Further extract acoustic features such as speech rate (SR), Mel-frequency cepstral coefficients (MFCC), volume (RMS Energy), etc., and extract prosodic features such as fundamental frequency (F0) curve, formants, etc.
[0042] Since people will show manifestations such as higher pitch, faster speech rate, and higher volume in an anxious emotional state, acoustic features can accurately capture the intonation contour, reflect the number of syllables per unit time and emotional intensity, and prosodic features can reflect the intonation rise and fall pattern and analyze the vocal tract shape, which also helps to distinguish emotional states.
[0043] In the embodiment of the present application, based on the acoustic features and prosodic features, use a large model to determine the user's current emotional characteristics. Specifically, it can include multi-level emotion labels. For example, at the upper-level secondary classification labels: positive (confidence level 0.8), neutral (0.5), negative (0.3), and at the lower-level fine-grained labels: angry (0.9), anxious (0.7), pleasant (0.6).
[0044] For example, if the voice command information input by the user is "Navigate to the nearest gas station, hurry up!", it can output: {"Emotion": "Anxious", "Confidence level": 0.85}.
[0045] S3022: Conduct timbre analysis on the voice command information to extract the user type characteristics.
[0046] In a possible embodiment, based on a large model, conduct timbre analysis on the voice command information to extract the user type characteristics. The user type characteristics include age characteristics and gender characteristics.
[0047] First, extract the voiceprint features and timbre features of the voice command information, including performing fundamental frequency (F0) statistics. The average fundamental frequency of males is in the range of 85 - 180 Hz, and that of females is in the range of 165 - 255 Hz. It also includes performing harmonic structure analysis. Female vocal cords are shorter and have richer high-frequency harmonics. Input the above voiceprint features and timbre features into the classification model to infer the user's gender characteristics. For example, when inputting a piece of voice command information, the output can be: {"Gender": "Male", "Confidence": 0.92}.
[0048] Secondly, extract the formants, speed - intonation combined features, and voiceprint aging features of the voice command information, and input them into the age segmentation model to infer the user's age characteristics. For example, when inputting a piece of voice command information, the output can be: {"Age range": "30 - 45 years old", "Confidence": 0.88}.
[0049] Through the above steps, the age characteristics and gender characteristics of the user can be extracted, which is convenient for subsequently recommending more likely interesting content to users of different genders and different age groups, achieving stronger pertinence.
[0050] S303: Analyze the vehicle driving information to determine the corresponding driving scenario.
[0051] In a possible embodiment, the vehicle driving information further includes driving state information and driving environment information.
[0052] In a possible embodiment, the driving state information may include vehicle speed, acceleration (longitudinal / transverse), braking frequency, steering operation, steering angle, lane keeping deviation rate, etc.
[0053] In a possible embodiment, the driving environment information may involve road conditions (congestion index 0 - 1), weather conditions (rain / snow / fog), time (day or night), density of surrounding vehicles (vehicles per 100 meters), etc.
[0054] By combining the above driving state information and driving environment information, it can be comprehensively determined whether the current is a normal safe driving state or a state that requires emotional guidance.
[0055] In a possible embodiment, the analyzing the vehicle driving information to determine the corresponding driving scenario includes: In a possible embodiment, if the driving state information and the driving environment information indicate that the current driving state is safe, determine the information demand scenario as the corresponding driving scenario.
[0056] In a possible embodiment, if the driving state information and the driving environment information indicate that the current requires emotional guidance, determine the emotional demand scenario as the corresponding driving scenario.
[0057] In the embodiments of the present application, the meaning of the current driving state being safe is normal driving under stable road conditions, and the safety factor can support the listening of audio with a certain amount of information. Or rather, audio with a certain amount of information will not affect the user's normal driving.
[0058] In the embodiments of the present application, the meaning of the current need for emotional guidance is that in complex or high-pressure driving environments (such as congestion, emergencies), the safety factor cannot support the listening of audio with a certain amount of information. Or rather, audio with a certain amount of information will affect the user's normal driving, and the user needs audio with low or no information such as white noise and light music for emotional guidance.
[0059] In a possible embodiment, when all of the following conditions are met, it is determined that the current driving state is safe, and the information demand scenario is determined as the corresponding driving scenario: (1) Vehicle speed stability: The current vehicle speed change rate < 5%; (2) Driving smoothness: The standard deviation of lateral acceleration < 0.3 m / s²; (3) Low environmental risk: The road condition index ≤ 0.3; (4) No bad weather (not rain / snow / fog); (5) The time period is during the day. Otherwise, it is determined that the current need for emotional guidance, and the emotional demand scenario is determined as the corresponding driving scenario.
[0060] In another possible embodiment, when any of the following conditions is met, it is determined that the current need for emotional guidance due to congestion pressure, and the emotional demand scenario is determined as the corresponding driving scenario: (1) The road condition index > 0.7 and the duration ≥ 10 minutes; (2) Sudden interference: The number of rapid acceleration and deceleration times > 3 times / minute and the lane departure rate > 15%; (3) Environmental pressure: The weather is rain or snow or fog; (4) The time is at night (18:00 - 6:00) and the vehicle lights are not turned on. Otherwise, it is determined that the current driving state is safe, and the information demand scenario is determined as the corresponding driving scenario.
[0061] S304: Obtain the first audio data and the second audio data.
[0062] In a possible embodiment, the first audio data is a popular audio associated with the current time information; the second audio data is a popular audio associated with the current location information.
[0063] S305: Analyze the first audio data and the second audio data to obtain the content tendency characteristics.
[0064] That is to say, the first audio data represents the popular audio of the current month, day or time, the second audio data represents the popular audio of the current province, city or region, and then the large model is used to analyze the popular audio in two dimensions of the first audio data and the second audio data to obtain the content tendency characteristics in the current time and space state.
[0065] Specifically, the first audio data and the second audio data are first merged and de-duplicated, and then content analysis is performed to extract text features (keywords such as "technology" and "finance"), sentiment tendencies (positive, neutral, or negative), rhythm features (music styles such as electronic music), and timbre features from the title or description.
[0066] Then, the above features are input into a large model, prompt words are designed to generate content tendency features. Specifically, a three-level structured prompt word is constructed, including task definition, assigning a specific identity, activating relevant domain knowledge, then defining a constraint condition layer (JSON output format, including relevant data albums, singles, and radio names), and finally giving relevant cases as few-shot samples to inject in-vehicle domain knowledge.
[0067] For example, if the first audio data is current popular content: {XXX,XXX}, and the second audio data is regional popular content: {XXX play count ↑35%, XXX collection rate ↑22%}, some audio features are finally output as potential listening audio resources for users.
[0068] S306: Obtain the third audio data.
[0069] In a possible embodiment, the third audio data is a to-be-recommended audio resource locally or in the cloud, including various types such as podcasts, audiobooks, songs, light music, and white noise.
[0070] S307: Perform content recognition on the third audio data and convert it into third text data.
[0071] In a possible embodiment, for each third audio data, preprocessing such as segmenting and noise reduction can be performed on the third audio data first to improve the audio quality, reduce noise interference, and adapt to subsequent processing models.
[0072] Then, voice recognition is performed on the preprocessed third audio data based on a voice recognition model to obtain initial text information. Specifically, automatic speech recognition is performed on the preprocessed audio information based on a locally built ASR model, and the initial text information is cached.
[0073] In a possible embodiment, cleaning processing such as stop word removal, spelling correction, and digital standardization can also be performed on the initial text information to obtain the third text data, optimize the text quality, and facilitate subsequent semantic analysis processing.
[0074] S308: In response to the received first structured prompt word, extract multiple basic tags of the third text data based on a preset large model.
[0075] In a possible embodiment, the first structured prompt is also a three-level structured prompt, including first task specification information, first constraint condition information, and first example information.
[0076] In the embodiment of the present application, the first task specification information, the first constraint condition information, and the first example information constitute a three-level prompt structure. Through the first task specification information, it is clearly required that the model execute "basic label generation". Through the first constraint condition information, the number of generated labels (such as 3), the output format (such as Json structure), and the output result (label text and label confidence score) are specified. Finally, through the first example information as the Few-shot layer, vehicle domain knowledge is injected, and some sample inputs and outputs are given.
[0077] For example, the first task specification information requires the model to generate basic labels based on text information. The first constraint condition information requires that the number of generated labels is 3, the output format is Json format, and the output result includes label text and label confidence score. The first example information may be: the audio information is the playing media information "Fold a thousand paper cranes and tie a red belt", and the output labels are "festive" (0.88), "Spring Festival" (0.64), and "thousand paper cranes" (0.96).
[0078] S309: In response to the received second structured prompt, based on a preset large model, construct a hierarchical relationship for the multiple basic labels to obtain the scenario layer label, the user layer label, and the content layer label of the third audio data.
[0079] In a possible embodiment, the second structured prompt is also a three-level structured prompt, including second task specification information, second constraint condition information, and second example information.
[0080] Through the second task specification information, it is clearly required that the model execute "build a hierarchical system for the basic labels, including a scenario layer, a user layer, and a content layer, and there needs to be an association between the layers". Through the second constraint condition information, the number of layers (such as 3), the output format (such as Json structure), and the output result (layer text and category name) are specified. Finally, through the second example information as the Few-shot layer, some desired category names are specified, and some sample inputs and outputs are given.
[0081] For example, the scenario layer label: medium information volume scenario (0.40), the user layer label: "festive" (0.88), and the content layer labels: "Spring Festival" (0.64), "thousand paper cranes" (0.96).
[0082] Through the above process, the third audio data carries corresponding scene layer tags, user layer tags, and content layer tags, and there is a hierarchical relationship among the scene layer tags, the user layer tags, and the content layer tags.
[0083] S310: Recall the fourth audio data from the third audio data based on the driving scene, the user emotion feature, the user type feature, and the content tendency feature.
[0084] In a possible embodiment, the recalling the fourth audio data from the third audio data based on the driving scene, the user emotion feature, the user type feature, and the content tendency feature includes: For each third audio data, perform: S3101: Determine whether the driving scene matches the scene layer tag. If it does, execute S3102; if not, execute S3106.
[0085] S3102: Determine the first similarity based on the user emotion feature, the user type feature, and the user layer tag.
[0086] In a possible embodiment, vectorize the above features and calculate the cosine similarity between the vectors as the first similarity.
[0087] S3103: Determine whether the first similarity is greater than a preset similarity threshold. If it is, execute S3104; if not, execute S3106.
[0088] S3104: Determine the second similarity based on the content tendency feature and the content layer tag.
[0089] In a possible embodiment, vectorize the above features and calculate the cosine similarity between the vectors as the second similarity.
[0090] S3105: Recall the fourth audio data from the third audio data based on the second similarity.
[0091] In a possible embodiment, sort the third audio data based on the second similarity, and take the top N required audio data as the fourth audio data.
[0092] S3106: Determine that this third audio data does not meet the requirements.
[0093] Among them, the scene layer tags have a one-to-one corresponding information density. That is to say, for an audio data, it has an information density. For example, podcasts and audiobooks have the highest information density, and their scene layer tags are high-information scenes (above 0.5). Ordinary songs have a medium information density, and their scene layer tags are medium-information scenes (0.2 - 0.5). Light music or white noise has the lowest information density, and their scene layer tags are low-information scenes (below 0.2).
[0094] In a possible embodiment, if the driving scene matches the scene layer tag, determining a first similarity based on the user emotion feature, the user type feature, and the user layer tag includes: Based on the driving scene and the information density, determine a safety factor based on the preset large model. Specifically, simulate the safety factor of receiving audio data with this information density in this driving scene through the preset large model to obtain an accurate value, which is convenient for subsequent calculation and comparison.
[0095] Optionally, if the safety factor is greater than the preset safety factor threshold, determine that the driving scene matches the scene layer tag; otherwise, if the safety factor is less than or equal to the preset safety factor threshold, determine that the driving scene does not match the scene layer tag.
[0096] Determine the first similarity based on the user emotion feature, the user type feature, and the user layer tag.
[0097] Figure 4 It is a schematic flow of a cold start method based on a large model provided by an embodiment of the present application Figure 3 The method may include: S401: Obtain updated fifth audio data.
[0098] In a possible embodiment, the fifth audio data is updated audio data, such as a new album, single, or broadcast, and needs to be recommended to the user.
[0099] S402: Determine the scene layer tag, the user layer tag, and the content layer tag corresponding to the fifth audio data.
[0100] S403: Obtain multiple zero-shot users.
[0101] In the embodiment of the present application, each of the multiple zero-shot users carries the corresponding driving scene, user emotion feature, user type feature, and content tendency feature. The specific feature extraction process is as described above and will not be elaborated here.
[0102] S404: Recall multiple users to be shared among the multiple zero-shot users based on the scenario layer tags, the user layer tags, and the content layer tags.
[0103] Similar to recalling multiple fourth audio data from multiple third audio data described above, it is also to first perform scenario matching, then calculate the first similarity and the second similarity, and finally sort based on the second similarity, and take the top N zero-shot users as the users to be shared.
[0104] By setting multi-level tags and multi-level user-side features in this application, horizontal audio matching alignment can be achieved, effectively improving the accuracy and comprehensiveness of the user portraits of zero-shot users, and also effectively enhancing the scene-based and hierarchical nature of audio resource classification, which plays an important role both in the process of recommending audio data to users and in the process of finding users for audio data.
[0105] The embodiment of this application also provides a cold start device based on a large model. Figure 5 It is a schematic structural diagram of a cold start device based on a large model provided by the embodiment of this application, as Figure 5 shown. The device 500 includes: A first acquisition module 510, configured to acquire the voice command information and vehicle driving information of the user; the vehicle driving information includes the current location information and the current time information; A first analysis module 520, configured to analyze the voice command information to obtain the user emotion feature and the user type feature; A second analysis module 530, configured to analyze the vehicle driving information to determine the corresponding driving scenario; A second acquisition module 540, configured to acquire the first audio data and the second audio data; the first audio data is the popular audio associated with the current time information; the second audio data is the popular audio associated with the current location information; A third analysis module 550, configured to analyze the first audio data and the second audio data to obtain the content tendency feature; A third acquisition module 560, configured to acquire the third audio data; the third audio data carries the corresponding scenario layer tags, user layer tags, and content layer tags; there is a hierarchical relationship among the scenario layer tags, the user layer tags, and the content layer tags; A first recall module 570, configured to recall the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature, and the content tendency feature.
[0106] In an optional embodiment, it further includes: A content conversion module for performing content recognition on the third audio data and converting it into third text data; A first tagging module for, in response to receiving a first structured prompt word, extracting multiple basic tags of the third text data based on a preset large model; A second tagging module for, in response to receiving a second structured prompt word, constructing a hierarchical relationship based on the preset large model for the multiple basic tags to obtain a scenario layer tag, a user layer tag, and a content layer tag of the third audio data.
[0107] In an alternative embodiment, it includes: A first determination module for, if the driving scenario matches the scenario layer tag, determining a first similarity based on the user emotion feature, the user type feature, and the user layer tag; A second determination module for, when the first similarity is greater than a preset similarity threshold, determining a second similarity based on the content tendency feature and the content layer tag; A second recall module for recalling fourth audio data from the third audio data based on the second similarity.
[0108] In an alternative embodiment, the scenario layer tag has a corresponding one-to-one information density; it includes: A third determination module for determining a safety factor based on the driving scenario and the information density based on the preset large model; A fourth determination module for, if the safety factor is greater than a preset safety factor threshold, determining that the driving scenario matches the scenario layer tag; A fifth determination module for determining the first similarity based on the user emotion feature, the user type feature, and the user layer tag.
[0109] In an alternative embodiment, it includes: A first feature extraction module for performing intonation analysis on the voice command information to extract the user emotion feature; A second feature extraction module for performing timbre analysis on the voice command information to extract the user type feature; the user type feature includes an age feature and a gender feature; In an alternative embodiment, the vehicle driving information further includes driving state information and driving environment information; it includes: A sixth determination module for, if the driving state information and the driving environment information indicate that the current driving state is safe, determining that the information demand scenario is the corresponding driving scenario; or; A seventh determination module, configured to determine that the emotion demand scenario is the corresponding driving scenario if the driving state information and the driving environment information indicate that emotion guidance is currently required.
[0110] In an optional embodiment, it further includes: A third acquisition module, configured to acquire updated fifth audio data; An eighth determination module, configured to determine the scenario layer label, the user layer label, and the content layer label corresponding to the fifth audio data; A fourth acquisition module, configured to acquire a plurality of zero-sample users; each of the plurality of zero-sample users carries the corresponding driving scenario, user emotion feature, user type feature, and content tendency feature; A third recall module, configured to recall a plurality of users to be shared from the plurality of zero-sample users based on the scenario layer label, the user layer label, and the content layer label.
[0111] The device and method embodiments in this application are based on the same application concept.
[0112] The method embodiments provided in the embodiments of this application can be executed on a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 6 It is a hardware structure block diagram of a server for a cold start method based on a large model provided in the embodiments of this application. As Figure 6 shown, the server 600 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 610 (the central processing unit 610 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 630 for storing data, and one or more storage media 620 for storing application programs 623 or data 622 (for example, one or more mass storage devices). Among them, the memory 630 and the storage media 620 may be transient storage or persistent storage. The program stored in the storage media 620 may include one or more modules, and each module may include a series of instruction operations on the server. Further, the central processing unit 610 may be configured to communicate with the storage media 620 and execute a series of instruction operations in the storage media 620 on the server 600. The server 600 may further include one or more power supplies 660, one or more wired and wireless network interfaces 650, one or more input / output interfaces 640, and / or, one or more operating systems 621, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0113] The input / output interface 640 can be used to receive or transmit data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the server 600. In one example, the input / output interface 640 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the input / output interface 640 can be a RadioFrequency (RF) module, which is used to communicate with the Internet wirelessly.
[0114] Those of ordinary skill in the art can understand that Figure 6 The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the server 600 may further include more or fewer components than those shown Figure 6 in the figure, or have a different configuration from that shown Figure 6 in the figure.
[0115] An embodiment of the present application provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned data processing method.
[0116] An embodiment of the present application further provides a computer-readable storage medium. The storage medium can be disposed in the server to store at least one instruction, at least one program, a code set or an instruction set related to a cold start method based on a large model in the method embodiment. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned cold start method based on a large model.
[0117] Optionally, in this embodiment, the above-mentioned storage medium may be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above-mentioned storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs.
[0118] As can be seen from the embodiments of the cold start method, device, electronic device, or storage medium based on the large model provided by the present application, the present application obtains the voice command information and vehicle driving information of the user; the vehicle driving information includes the current location information and the current time information; analyzes the voice command information to obtain the user emotion feature and the user type feature; analyzes the vehicle driving information to determine the corresponding driving scenario; obtains the first audio data and the second audio data; the first audio data is the popular audio associated with the current time information; the second audio data is the popular audio associated with the current location information; analyzes the first audio data and the second audio data to obtain the content tendency feature; obtains the third audio data; the third audio data carries the corresponding scenario layer label, user layer label, and content layer label; there is a hierarchical relationship between the scenario layer label, the user layer label, and the content layer label; based on the driving scenario, the user emotion feature, the user type feature, and the content tendency feature, recalls the fourth audio data from the third audio data. In the embodiments of the present application, the current driving scenario, user emotion feature, user type feature, and content tendency feature of the user are extracted through the user's voice command information and vehicle driving information, and are matched with the third audio data carrying the corresponding scenario layer label, user layer label, and content layer label, and the fourth audio data that the user may be interested in is recalled. Without the need for the user's historical data and user privacy data, zero-shot cold start is achieved, and personalized audio is accurately pushed to the user.
[0119] It should be noted that: the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the advantages or disadvantages of the embodiments. And the above description of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0120] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0121] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, an optical disc, etc.
[0122] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A cold start method based on a large model, characterized in that including; Obtain the voice command information of the user and the vehicle driving information; the vehicle driving information includes the current location information and the current time information; Analyze the voice command information to obtain the user emotion feature and the user type feature; Analyze the vehicle driving information to determine the corresponding driving scenario; Obtain the first audio data and the second audio data; the first audio data is the popular audio associated with the current time information; The second audio data is the popular audio associated with the current location information; Analyze the first audio data and the second audio data to obtain the content tendency feature; Obtain the third audio data; the third audio data carries the corresponding scenario layer label, user layer label and content layer label; There is a hierarchical relationship among the scenario layer label, the user layer label and the content layer label; Recall the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature and the content tendency feature.
2. The cold start method based on a large model according to claim 1, wherein Before recalling the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature and the content tendency feature, it further includes: Perform content recognition on the third audio data and convert it into the third text data; In response to the received first structured prompt word, extract multiple basic labels of the third text data based on the preset large model; In response to the received second structured prompt word, construct a hierarchical relationship based on the preset large model for the multiple basic labels to obtain the scenario layer label, the user layer label and the content layer label of the third audio data.
3. The cold start method based on a large model according to claim 2, wherein The recalling the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature and the content tendency feature includes: If the driving scenario matches the scenario layer label, determine the first similarity based on the user emotion feature, the user type feature and the user layer label; In the case where the first similarity is greater than the preset similarity threshold, determine the second similarity based on the content tendency feature and the content layer label; Recall the fourth audio data from the third audio data based on the second similarity.
4. The cold start method based on a large model according to claim 3, characterized in that The scenario layer label has a corresponding information density; The if the driving scenario matches the scenario layer label, determine the first similarity based on the user emotion feature, the user type feature and the user layer label, includes: Based on the driving scenario and the information density, determine the safety factor based on the preset large model; If the safety factor is greater than the preset safety factor threshold, determine that the driving scenario matches the scenario layer label; Determine the first similarity based on the user emotion feature, the user type feature and the user layer label.
5. A cold start method based on a large model according to claim 1, characterized in that, The analyzing the voice command information to obtain the user emotion feature and the user type feature includes: Perform intonation analysis on the voice command information to extract the user emotion feature; Perform timbre analysis on the voice command information to extract the user type features; the user type features include age features and gender features.
6. The cold start method based on a large model according to claim 1, characterized in that The vehicle driving information further includes driving state information and driving environment information; The analyzing the vehicle driving information to determine the corresponding driving scenario includes: If the driving state information and the driving environment information indicate that the current driving state is safe, determining the information demand scenario as the corresponding driving scenario; or; If the driving state information and the driving environment information indicate that current emotional guidance is needed, determining the emotional demand scenario as the corresponding driving scenario.
7. A cold start method based on a large model according to claim 1, characterized in that, Further includes: Obtain updated fifth audio data; Determine the scenario layer label, the user layer label, and the content layer label corresponding to the fifth audio data; Obtain multiple zero-shot users; the multiple zero-shot users respectively carry the corresponding driving scenario, the user emotion feature, the user type feature, and the content tendency feature; Recall multiple users to be shared from the multiple zero-shot users based on the scenario layer label, the user layer label, and the content layer label.
8. A cold start device based on a large model, characterized in that, Includes: A first acquisition module for acquiring the voice command information and the vehicle driving information of the user; The vehicle driving information includes the current location information and the current time information; A first analysis module for analyzing the voice command information to obtain the user emotion feature and the user type feature; A second analysis module for analyzing the vehicle driving information to determine the corresponding driving scenario; A second acquisition module for acquiring the first audio data and the second audio data; the first audio data is a popular audio associated with the current time information; the second audio data is a popular audio associated with the current location information; A third analysis module for analyzing the first audio data and the second audio data to obtain the content tendency feature; A third acquisition module for acquiring the third audio data; the third audio data carries the corresponding scenario layer label, user layer label, and content layer label; There is a hierarchical relationship among the scenario layer label, the user layer label, and the content layer label; A first recall module for recalling the fourth audio data from the third audio data based on the driving scenario, the user emotion feature, the user type feature, and the content tendency feature.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the large model-based cold start method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, At least one instruction or at least one program is stored in the computer-readable storage medium, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the large model-based cold start method according to any one of claims 1-7.
Citation Information
Patent Citations
Music pushing method and device, electronic equipment and readable storage medium
CN115114473A
Multimedia data determination method and device, equipment, storage medium and communication system
CN118695006A
Vehicle-mounted music interaction system based on context awareness
CN119597958A
Vehicle-mounted multimedia intelligent recommendation method and device based on scene and preference and vehicle
CN119821299A
Method and apparatus for pushing multimedia content
US20190147058A1