A vehicle-mounted video ringtone generation method, device, equipment and program product
By recognizing landmarks in the vehicle's surroundings to generate real-time virtual driving scenes and stylizing them, the problem of in-vehicle systems being unable to display video ringback tones has been solved, improving the visual experience during driving.
Patent Information
- Application Number
- CN202410568107.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2044-05-09
AI Technical Summary
Existing in-vehicle systems cannot display video ringback tones, resulting in ringback tones when receiving phone calls while driving typically being a single audio ringback tone, lacking a rich visual experience.
By identifying landmarks in the current vehicle's surrounding environment, a real-time virtual driving scene is generated, and the landmarks are stylized according to the characteristics of the ringback tone audio to generate an in-vehicle video ringback tone.
It enables the display of rich video ringback tones on the in-vehicle system, improving the user's call experience while driving and avoiding the limitations of mobile terminal systems.
Smart Images

Figure CN118803141B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video ringback tone, and in particular to a vehicle-mounted video ringback tone generation method and device, a terminal device and a computer program product. BACKGROUND
[0002] At present, more and more users use video ringback tone services on mobile terminals, so that users can see a video content on the mobile terminal when waiting for a telephone call to be connected or being called, and compared with the traditional audio ringback tone service, the video ringback tone service can provide better call experience for users.
[0003] When driving, users usually connect the mobile terminal to the vehicle-mounted system through a connection mode such as vehicle-mounted Bluetooth, so as to answer the phone during driving. However, since the existing vehicle-mounted system usually cannot display the video ringback tone, the ringback tone mode when receiving a telephone call during driving is usually only the audio ringback tone, and the ringback tone mode is single. SUMMARY
[0004] The present application provides a vehicle-mounted video ringback tone generation method, device, equipment and program product, which can generate a real-time virtual driving scene based on the current vehicle surrounding environment when detecting an incoming call signal, and perform stylized processing on the identifiers in the real-time virtual driving scene according to the audio features of the current ringback tone audio, so as to present a vehicle-mounted video ringback tone with rich effects and improve the call experience of users during driving.
[0005] To solve the above technical problems, the first aspect of the present application provides a vehicle-mounted video ringback tone generation method, comprising:
[0006] When detecting an incoming call signal, identifying a plurality of identifiers in the current vehicle surrounding environment, and generating a real-time virtual driving scene according to the identification information of the plurality of identifiers;
[0007] According to the identification information, calculating the feature value of each identifier in the real-time virtual driving scene;
[0008] According to the preset audio features of the current ringback tone audio and the feature value, screening the plurality of identifiers in the real-time virtual driving scene to obtain a plurality of target identifiers;
[0009] According to the audio features, performing stylized processing on the plurality of target identifiers in the real-time virtual driving scene, generating and displaying a vehicle-mounted video ringback tone.
[0010] As a preferred scheme, the calculation of the feature value of each identifier in the real-time virtual driving scene according to the identification information specifically comprises:
[0011] acquire distance information and pixel size value of each of the markers from the identification information;
[0012] determine a feature value of each of the markers in the real-time virtual driving scene according to a product of the distance information and a first weight value and a product of the pixel size value and a second weight value; wherein a sum of the first weight value and the second weight value is 1; the first weight value and the second weight value are pre-set or change with the change of the vehicle driving scene.
[0013] As a preferred solution, when the first weight value and the second weight value change with the change of the vehicle driving scene, the first weight value and the second weight value are adjusted in the following way:
[0014] when the vehicle driving scene is identified as a non-driving scene or a red light waiting scene, the first weight value is adjusted to a first pre-set weight threshold, and the second weight value is adjusted to a second pre-set weight threshold; wherein the first pre-set weight threshold is less than the second pre-set weight threshold;
[0015] when the vehicle driving scene is identified as a driving scene or a green light passing scene, the first weight value is adjusted to a third pre-set weight threshold, and the second weight value is adjusted to a fourth pre-set weight threshold; wherein the third pre-set weight threshold is greater than the fourth pre-set weight threshold.
[0016] As a preferred solution, the several target markers are obtained by screening the several markers in the real-time virtual driving scene according to the pre-set audio features of the current ringtone audio and the feature values, specifically including:
[0017] acquire the total number of scales from the audio features;
[0018] sort the feature values of the several markers according to a pre-set sorting rule;
[0019] screen the several markers in the real-time virtual driving scene according to the total number of scales and the sorting order of the feature values of the several markers, and obtain the several target markers; wherein the number of the target markers is equal to or less than the total number of scales.
[0020] As a preferred solution, the several target markers in the real-time virtual driving scene are stylized according to the audio features, specifically including:
[0021] acquire the occurrence frequency and the playing time point information of each note from the audio features;
[0022] determine the target object corresponding to each of the notes in the real-time virtual driving scene according to the frequency of occurrence and the order of the characteristic values of the target objects;
[0023] When the current ringtone audio is detected, perform style processing on the target object corresponding to each of the notes in the real-time virtual driving scene according to the playing time point information of each of the notes.
[0024] As a preferred solution, the style processing on the target object corresponding to each of the notes in the real-time virtual driving scene according to the playing time point information of each of the notes specifically includes:
[0025] perform style processing on the target object corresponding to the note at the current playing time point in the real-time virtual driving scene according to a preset image style based on the playing time point information of each of the notes;
[0026] adjust the style attribute of the target object corresponding to the note at the current playing time point in the real-time virtual driving scene according to a preset strength level of each of the notes, wherein the strength level of each of the notes is obtained from the audio characteristic.
[0027] As a preferred solution, the audio characteristic includes a ringtone number of the ringtone audio corresponding thereto; and the method specifically obtains the audio characteristic of the current ringtone audio by the following steps:
[0028] receive a to-be-processed ringtone number sent by a ringtone server;
[0029] determine whether the ringtone number in any one of the audio characteristics matches the to-be-processed ringtone number based on a plurality of pre-stored audio characteristics;
[0030] when the ringtone number in any one of the audio characteristics matches the to-be-processed ringtone number, take the any one of the audio characteristics as the audio characteristic of the current ringtone audio;
[0031] The audio characteristic is specifically string information received from a ringtone server, and the string information is constructed by the ringtone server based on a set ringtone audio of a current vehicle user, and the ringtone server analyzes the total number of scales of the set ringtone audio, the frequency of occurrence of each note, the playing time point information, and the strength level.
[0032] The second aspect of the embodiment of the application provides a vehicle-mounted video ringtone generation device, which comprises:
[0033] The real-time virtual driving scene generation module is configured to, when detecting the incoming call signal, identify a plurality of markers in the current vehicle surrounding environment, and generate a real-time virtual driving scene according to the identification information of the plurality of markers.
[0034] The feature value calculation module is configured to calculate a feature value of each of the markers in the real-time virtual driving scene according to the identification information.
[0035] The marker screening module is configured to screen the plurality of markers in the real-time virtual driving scene according to the preset audio feature of the ringtone audio and the feature value, and obtain a plurality of target markers.
[0036] The vehicle-mounted video ringtone generation module is configured to perform stylization processing on the plurality of target markers in the real-time virtual driving scene according to the audio feature, and generate and display a vehicle-mounted video ringtone.
[0037] The third aspect of the embodiment of the present application provides a terminal device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the vehicle-mounted video ringtone generation method according to any one of the first aspect when executing the computer program.
[0038] The fourth aspect of the embodiment of the present application provides a computer program product, which comprises computer programs / instructions, and the computer programs / instructions implement the steps of the vehicle-mounted video ringtone generation method according to any one of the first aspect when executed by a processor.
[0039] Compared with the prior art, the embodiment of the present application has the beneficial effects that, when detecting the incoming call signal, a real-time virtual driving scene can be generated by identifying a plurality of markers in the current vehicle surrounding environment, then the feature values of the markers can be calculated according to the identification information of the markers, and the markers in the real-time virtual driving scene can be stylized in combination with the audio feature of the current ringtone audio, so that a vehicle-mounted video ringtone matching the current vehicle surrounding environment and the current ringtone audio can be generated, the vehicle-mounted system can be used to present the vehicle-mounted video ringtone with rich effects, and the user's call experience during driving is improved. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1 is a flowchart of the vehicle-mounted video ringtone generation method in the embodiment of the present application;
[0041] Figure 2 is a flowchart of the marker feature value calculation in the embodiment of the present application;
[0042] Figure 3 is a flowchart of the marker screening in the embodiment of the present application;
[0043] Figure 4 is a flowchart of the style processing of the marker in the embodiment of the present application;
[0044] Figure 5 is a structural diagram of the vehicle-mounted video ringtone generation device in the embodiment of the present application;
[0045] Figure 6 is a structural diagram of the terminal device in the embodiment of the present application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0047] Referring to Figure 1 , Figure 1 is a flowchart of the vehicle-mounted video ringtone generation method provided by the embodiment of the present application. The first aspect of the embodiment of the present application provides a vehicle-mounted video ringtone generation method, including the following steps S1 to S4:
[0048] Step S1, when a call signal is detected, a plurality of markers in the current vehicle surrounding environment are identified, and a real-time virtual driving scene is generated according to the identification information of the plurality of markers;
[0049] Step S2, according to the identification information, a characteristic value of each of the markers in the real-time virtual driving scene is calculated;
[0050] Step S3, according to a preset audio characteristic of a current ringtone audio and the characteristic value, a plurality of the markers in the real-time virtual driving scene are screened to obtain a plurality of target markers;
[0051] Step S4, according to the audio characteristic, a plurality of the target markers in the real-time virtual driving scene are stylized to generate a vehicle-mounted video ringtone and display the vehicle-mounted video ringtone.
[0052] It is worth noting that in the prior art, the video ringtone is not only difficult to display on the vehicle-mounted system, but also is limited by the mobile terminal system. For example, the video ringtone usually does not support display in the iOS system, but can only be displayed in the Android system. The embodiment can realize display of the video ringtone to the user on the vehicle-mounted system, and since the display of the video ringtone is completely taken over by the vehicle-mounted system, the operation system of the mobile terminal is not limited, thereby ensuring the application range of the vehicle-mounted video ringtone.
[0053] In step S1, the embodiment establishes a connection with the user's mobile terminal by means of, for example, vehicle-mounted Bluetooth, WiFi, etc., so that when the user's mobile terminal receives an incoming call signal, the embodiment can detect the incoming call signal and identify a plurality of markers in the current vehicle surroundings, obtaining identification information of each marker, which at least includes marker number, marker name, image contour information, distance information, and pixel size value. Optionally, the embodiment activates a plurality of cameras equipped in the vehicle to collect images of the vehicle surroundings, and then uses a trained AI graphic recognition engine to identify the markers in each collected image, and determines the real-time distance information and pixel size value of each marker based on deep learning technology and stereo vision technology. Specifically, the above-mentioned markers include but are not limited to vehicles of different models and colors, road signs, lane lines, buildings, and people; during the training process of the AI graphic recognition engine, a large number of marker image data can be used as a training set, which can include marker information such as number, name, color, distance, etc., so that the AI graphic recognition engine can learn the characteristics of each marker in the image based on deep learning technology. In addition, stereo vision technology is a computer vision technology, which aims to infer the depth information of each pixel point in the image from two or more images. This depth information is the distance information of the identified object. In the identification process, the contour information of each marker in the image is first extracted, the pixel size value of the marker is determined by calculating the number of pixel points, and the distance information of the marker is determined by calculating the depth information of the pixel points. Further, since the identification information of the markers around the vehicle has been obtained at this time, a real-time virtual driving scene can be modeled and generated. It should be noted that since the current vehicle can be in a driving state or a stationary state, and the identified markers can be in a moving state, the identification information of the markers around the vehicle is constantly updated, so the real-time virtual driving scene is also constantly updated.
[0054] As one of the optional embodiments, based on the plurality of markers identified in step S1, a marker sequence is constructed, which records the identification information of each marker around the current vehicle, including marker number, marker name, image contour information, distance information, and pixel size value, etc. As mentioned above, since the identification information of the markers around the vehicle is constantly updated, the identification information of each marker in the marker sequence is also constantly updated.
[0055] Further, the embodiment considers that the stylization requirements of the markers in the real-time virtual driving scene are different for different driving scenarios or user preferences, and therefore calculates a feature value based on the identification information of each marker as one of the screening factors for the markers that need to be stylized. In addition, the embodiment considers that the finally generated vehicle-mounted video ringtone not only needs to fit the vehicle surrounding environment, but also needs to fit the current user's ringtone audio, and therefore the audio features of the current ringtone audio are also used as one of the screening factors for the markers that need to be stylized, and the target markers are screened in combination with the audio features of the current ringtone audio and the feature values of the markers.
[0056] Further, the embodiment considers that different ringtone audios have different audio features, and accordingly, the display effect of the vehicle-mounted video ringtone should also fit different audio features to present effect diversity, for example, for audio with strong rhythm, the picture of the vehicle-mounted video ringtone changes color or animation rhythm accordingly, and for audio with slow rhythm, the picture of the vehicle-mounted video ringtone presents a soft style accordingly. Therefore, the embodiment performs stylization processing on a plurality of target markers in the real-time virtual driving scene based on the audio features, such as contour highlighting processing, two-dimensional stylization processing, national animation stylization processing, sketch stylization processing, etc., which are not specifically limited in the embodiment, so as to generate and display the vehicle-mounted video ringtone. Optionally, the AIGC technology is used to realize the stylization processing of the target markers and generate the vehicle-mounted video ringtone.
[0057] For example, the embodiment can be displayed through vehicle-mounted screens such as HUD head-up display screens, center screens, etc. Preferably, when the vehicle is equipped with a HUD head-up display screen, the generated vehicle-mounted video ringtone is directly projected onto the front windshield to provide better visual experience for the user.
[0058] In the embodiment, when the incoming call signal is detected, a real-time virtual driving scene can be generated by identifying a plurality of markers in the current vehicle surrounding environment, and then the feature values of the markers can be calculated according to the identification information of the markers, and the markers in the real-time virtual driving scene are stylized in combination with the audio features of the current ringtone audio, so as to generate a vehicle-mounted video ringtone that matches the current vehicle surrounding environment and the current ringtone audio, realize the presentation of the vehicle-mounted video ringtone with rich effects by using the vehicle-mounted system, and improve the call experience of the user during driving.
[0059] Referring to Figure 2 , Figure 2 is a flowchart of the marker feature value calculation provided by the embodiment. As a preferred solution, the calculation of the feature value of each marker in the real-time virtual driving scene based on the identification information specifically includes the following steps S21 to S22:
[0060] Step S21, obtaining distance information and pixel size value of each of the markers from the identification information;
[0061] Step S22, determining a feature value of each of the markers in the real-time virtual driving scene according to a product of the distance information and a first weight value and a product of the pixel size value and a second weight value; wherein a sum of the first weight value and the second weight value is 1; the first weight value and the second weight value are pre-set or change with the change of the vehicle driving scene.
[0062] Specifically, the embodiment first obtains distance information and pixel size value of each marker, and then obtains a feature value of each marker in the real-time virtual driving scene according to the formula: feature value = (distance information * first weight value + pixel size value * second weight value) / 2.
[0063] It is worth noting that the embodiment determines the marker with a relatively close or far distance as the marker to be stylized by the first weight value and the second weight value. It can be understood that the marker with a larger distance value in the distance information indicates a marker with a relatively far distance, such as a building, a mountain, a sunset, etc. in the distance, and the marker with a larger pixel size value indicates a marker with a relatively close distance, such as a vehicle, a pedestrian, a lane line, a road sign, a tree, etc. in the distance. When a marker has a larger distance value, its pixel size value in the image is generally smaller, and when a marker has a larger pixel size value in the image, its distance value is generally smaller. Therefore, when the first weight value is larger and the second weight value is smaller, it indicates that the marker selection process is biased towards selecting markers with a relatively far distance, and when the first weight value is smaller and the second weight value is larger, it indicates that the marker selection process is biased towards selecting markers with a relatively close distance.
[0064] It is worth noting that the first weight value and the second weight value in the embodiment can be pre-set. For example, if the user prefers to stylize the markers in the distance, the first weight value can be set to be larger and the second weight value can be set to be smaller, and if the user prefers to stylize the markers in the distance, the first weight value can be set to be smaller and the second weight value can be set to be larger. Alternatively, the first weight value and the second weight value can change with the change of the vehicle driving scene, so as to realize flexible transformation of the presentation picture of the vehicle-mounted video ringtone.
[0065] As a preferred solution, when the first weight value and the second weight value change with the change of the vehicle driving scene, the first weight value and the second weight value are adjusted in the following manner:
[0066] When it is identified that the vehicle driving scene is a non-driving scene or a red light waiting scene, the first weight value is adjusted to a first preset weight threshold, and the second weight value is adjusted to a second preset weight threshold; wherein the first preset weight threshold is less than the second preset weight threshold.
[0067] When it is identified that the vehicle driving scene is a driving scene or a green light passing scene, the first weight value is adjusted to a third preset weight threshold, and the second weight value is adjusted to a fourth preset weight threshold; wherein the third preset weight threshold is greater than the fourth preset weight threshold.
[0068] Specifically, when it is detected that the vehicle is in a static state, or the image collected by the camera determines that the signal light of the current intersection is a red light, it is determined that the current vehicle driving scene is a non-driving scene or a red light waiting scene, at this time the first weight value is adjusted to a first preset weight threshold, and the second weight value is adjusted to a second preset weight threshold, so that the first weight value is less than the second weight value, the purpose is to style as many as possible the identification objects with a relatively short distance, thereby improving the display effect of the vehicle-mounted video color bell. Preferably, the first preset weight threshold is less than 0.4, such as 0.35, 0.3, 0.2, 0.1, etc., which is not specifically limited in the embodiment, and the second preset weight threshold is correspondingly greater than 0.6, such as 0.65, 0.7, 0.8, 0.9, etc., which is not specifically limited in the embodiment, thereby better ensuring that the identification objects for style processing are identification objects with a relatively short distance.
[0069] When it is detected that the vehicle is in a non-static state, or the image collected by the camera determines that the signal light of the current intersection is a green light, it is determined that the current vehicle driving scene is a driving scene or a green light passing scene, at this time the first weight value is adjusted to a third preset weight threshold, and the second weight value is adjusted to a fourth preset weight threshold, so that the first weight value is greater than the second weight value, the purpose is to style as many as possible the identification objects with a relatively long distance, so as not to affect the user driving, which helps to avoid traffic accidents and safety problems. Preferably, the third preset weight threshold is greater than 0.7, such as 0.75, 0.8, 0.9, 0.95, etc., which is not specifically limited in the embodiment, and the fourth preset weight threshold is correspondingly less than 0.3, such as 0.25, 0.2, 0.1, 0.05, etc., which is not specifically limited in the embodiment, thereby better ensuring that the identification objects for style processing are identification objects with a relatively long distance. Further, the display of the vehicle-mounted video color bell in this vehicle driving scene needs to cooperate with the auxiliary driving system of the vehicle, for example, when the reverse auxiliary system is triggered and needs to display a reverse image on the center control screen, the vehicle-mounted video color bell is not displayed on the center control screen at this time, thereby further avoiding traffic accidents and safety problems.
[0070] Referring toFigure 3 , Figure 3 is a flowchart of the identification object screening provided by the embodiment of the present application. As a preferred solution, the several identification objects in the real-time virtual driving scene are screened according to the preset audio features of the current color ring audio and the feature values, and several target identification objects are obtained, which specifically include the following steps S31 to S33:
[0071] Step S31, the total number of scales is obtained from the audio features.
[0072] Step S32, the feature values of the several identification objects are sorted according to a preset sorting rule.
[0073] Step S33, the several identification objects in the real-time virtual driving scene are screened according to the total number of scales and the sorting order of the feature values of the several identification objects, and several target identification objects are obtained; wherein the number of the target identification objects is equal to or less than the total number of scales.
[0074] It is worth noting that the total number of scales in the embodiment is the number of scales in the interval from the lowest note to the highest note in the current color ring audio. For example, assuming that there are 20 scales from the lowest note to the highest note in a color ring audio, the total number of scales is 20. Further, the feature values of each identification object are sorted according to a preset sorting rule. Optionally, the above sorting rule can be a descending sorting rule, i.e., the feature values of each identification object are sorted in descending order; or a descending sorting rule for distant identification objects, i.e., the feature values of each identification object with a distance value greater than a preset distance value (such as 10 km) are sorted in descending order; or a descending sorting rule for close identification objects, i.e., the feature values of each identification object with a distance value less than a preset distance value (such as 500 m) are sorted in descending order. Of course, the sorting rule can also be selected according to the actual display effect requirement, and the embodiment does not specifically limit the sorting rule.
[0075] Further, when the number of identification objects after sorting of the feature values is less than or equal to the total number of scales, each identification object after sorting of the feature values is regarded as a target identification object to be stylized; when the number of identification objects after sorting of the feature values is greater than the total number of scales, the several identification objects are screened from front to back according to the total number of scales and the sorting order of the feature values, so that the number of the target identification objects obtained finally is equal to the total number of scales. For example, assuming that the total number of scales of the current color ring audio is 20, the identification objects ranked in the top 20 according to the sorting order of the feature values are selected as the target identification objects.
[0076] In the embodiment of the present application, the target identifiers are screened based on the total number of scales of the current ringtone audio and the sequence of the characteristic values of the identifiers, so that the generated vehicle-mounted video ringtone can be consistent with the scale characteristics of the current ringtone audio.
[0077] Referring to Figure 4 , Figure 4 is a flowchart of the identifier stylization provided by the embodiment of the present application. As a preferred solution, the stylization of the target identifiers in the real-time virtual driving scene according to the audio characteristics specifically includes the following steps S41 to S43:
[0078] Step S41, the occurrence frequency and the playing time point information of each note are obtained from the audio characteristics;
[0079] Step S42, the target identifier corresponding to each note in the real-time virtual driving scene is determined according to the occurrence frequency and the sequence of the characteristic values of the target identifiers;
[0080] Step S43, when the current ringtone audio is detected to be played, the target identifier corresponding to each note in the real-time virtual driving scene is sequentially stylized based on the playing time point information of each note.
[0081] Specifically, in order to ensure that the generated vehicle-mounted video ringtone can be better consistent with the audio characteristics of the current ringtone audio, the embodiment first obtains the occurrence frequency and the playing time point information of each note from the audio characteristics, so that the generated vehicle-mounted video ringtone can produce corresponding picture changes with the rhythm of the current ringtone audio. Further, since the sequence of the characteristic values of the identifiers has been determined before the target identifiers are determined, the sequence of the characteristic values of the target identifiers is known, and in combination with the sequence of the occurrence frequency of the notes, the notes and the target identifiers are one-to-one corresponding, so that when a note is played in the playing process of the current ringtone audio, the corresponding target identifier will be stylized.
[0082] For example, it is assumed that in the playing process of the current ringtone audio, the current playing time point is the note "re" in the C scale, and the occurrence frequency of the note "re" is the highest in the current ringtone audio, so that the target identifier with the first sequence of the characteristic values is stylized at the current playing time point.
[0083] In the embodiment of the present application, the target identifiers in the real-time virtual driving scene are stylized based on the occurrence frequency and the playing time point information of each note in the current ringtone audio, so that the generated vehicle-mounted video ringtone can produce corresponding picture changes with the rhythm of the current ringtone audio.
[0084] As a preferred solution, the playing time point information of each note is used to sequentially perform stylized processing on the target object corresponding to each note in the real-time virtual driving scene, specifically including:
[0085] Based on the playing time point information of each note, the target object corresponding to the note at the current playing time point is stylized in the real-time virtual driving scene according to a preset image style.
[0086] According to the preset strength level of each note, the stylized attribute of the target object corresponding to the note at the current playing time point in the real-time virtual driving scene is adjusted, wherein the strength level of each note is obtained from the audio features.
[0087] Specifically, the embodiment first performs stylized processing on the target object corresponding to the note at the current playing time point based on the playing time point information of each note according to a preset image style, such as a high-profile highlight cool style, a two-dimensional style, a national cartoon style, a sketch style, etc. The embodiment does not make specific limitations here.
[0088] Further, since notes have different strengths, for example, note "f" represents strong, note "mf" represents medium strong, note "ff" represents very strong, note "fff" represents extra strong, note "p" represents weak, note "mp" represents medium weak, note "pp" represents very weak, and note "ppp" represents extra weak, so that different audio has different strengths in rhythm. In order to make the finally generated car video ringtone be able to produce corresponding picture changes with the rhythm of the current ringtone audio, the embodiment adjusts the stylized attribute of the target object corresponding to the note at the current playing time point according to the strength level of each note obtained from the audio features. It can be understood that the stylized attribute includes but is not limited to color, brightness, sharpening degree, contrast, etc. For each strength level of the note, the embodiment pre-sets a corresponding stylized attribute. Taking the high-profile highlight style as an example, the color of the outline is yellow. When the strength level of the note is extra weak, and the stylized attribute corresponding to the extra weak level is that the color is set to dark yellow, the color of the outline is adjusted to dark yellow when playing to the note, so as to correspond to the weaker audio rhythm. When the strength level of the note is extra strong, and the stylized attribute corresponding to the extra strong level is that the color is set to bright yellow, the color of the outline is adjusted to bright yellow when playing to the note, so as to correspond to the stronger audio rhythm.
[0089] In the embodiment of the present application, the style attribute of the several target markers in the real-time virtual driving scene is adjusted based on the strength level of each note in the current ringtone audio, so that the generated vehicle-mounted video ringtone can produce corresponding picture changes according to the rhythm and strength of the current ringtone audio.
[0090] As a preferred solution, the audio feature includes the ringtone number of the ringtone audio corresponding thereto; then, the method specifically acquires the audio feature of the current ringtone audio by the following steps:
[0091] Receiving the to-be-processed ringtone number sent by the ringtone server;
[0092] Based on the pre-stored several audio features, it is judged whether the ringtone number in any one of the audio features matches the to-be-processed ringtone number;
[0093] When the ringtone number in any one of the audio features matches the to-be-processed ringtone number, the any one of the audio features is taken as the audio feature of the current ringtone audio;
[0094] Wherein, the audio feature is specifically the string information received from the ringtone server, and the string information is constructed by the ringtone server based on the set ringtone audio of the current vehicle user, after analyzing the total number of scales, the occurrence frequency of each note, the playing time point information and the strength level of the set ringtone audio.
[0095] It is worth noting that for mobile terminals, whether it is an audio ringtone or a video ringtone, the user-set audio ringtone or video ringtone usually needs to be issued to the mobile terminal for display by the ringtone server, so when the incoming call signal is detected, the embodiment can receive the to-be-processed ringtone number of the current ringtone audio from the ringtone server, and it can be understood that the current ringtone audio can be the music of the user-set audio ringtone or the background music of the user-set video ringtone. Further, since the audio ringtone or video ringtone set by different users is usually different, accordingly, the music of different audio ringtone or the background music of different video ringtone also has different audio features, in order to accurately generate the vehicle-mounted video ringtone that fits the current ringtone audio, the embodiment selects the audio feature with the ringtone number matching the current to-be-processed ringtone number from the pre-stored several audio features as the audio feature of the current ringtone audio.
[0096] It is worth mentioning that each audio feature pre-stored in the embodiment is a string information received from the ringback tone server. It can be understood that the embodiment analyzes the set ringback tone audio of the current vehicle user through the ringback tone server, and the set ringback tone audio is the music of the audio ringback tone or the background music of the video ringback tone set by the user, so as to determine the total number of scales, the occurrence frequency of each note, the playing time point information and the strength level of the set ringback tone audio, and combine the ringback tone number pre-assigned to the current set ringback tone audio to construct a string information. Preferably, the embodiment classifies the scales of the set ringback tone audio and sorts the occurrence frequency of each note in descending order in the analysis process of the set ringback tone audio through the ringback tone server, so as to facilitate the subsequent screening of the target markers. For example, it is assumed that the set ringback tone audio has 18 scales from the lowest note to the highest note, and the scales of the set ringback tone audio are divided into 18 levels, which can be optionally represented by numbers from the 0th level to the 17th level.
[0097] As one of the optional embodiments, the vehicle-mounted video ringback tone generation method provided by the embodiment of the application can also be applied to the vehicle-mounted music listening scene. Specifically, the embodiment pre-acquires the audio features of each to-be-played music, including the total number of scales, the occurrence frequency of each note, the playing time point information and the strength level. When any music is detected to be played, a plurality of markers in the current vehicle surrounding environment are identified, and a real-time virtual driving scene is generated according to the identification information of the plurality of markers. The feature values of each marker in the real-time virtual driving scene are calculated. Then, according to the total number of scales of the currently played music and the sorting order of the feature values of the plurality of markers, the plurality of markers in the real-time virtual driving scene are screened to obtain a plurality of target markers. Further, according to the occurrence frequency of each note in the currently played music and the sorting order of the feature values of the plurality of target markers, the target marker corresponding to each note in the real-time virtual driving scene is determined. Based on the playing time point information of each note, the target marker corresponding to the note at the current playing time point in the real-time virtual driving scene is stylized. According to the preset strength level of each note, the stylized attribute of the target marker corresponding to the note at the current playing time point in the real-time virtual driving scene is adjusted, so as to generate and display the vehicle-mounted video ringback tone of the currently played music.
[0098] Referring to Figure 5 , Figure 5 is a structural schematic diagram of the vehicle-mounted video ringback tone generation device 100 provided by the embodiment of the application. The second aspect of the embodiment of the application provides a vehicle-mounted video ringback tone generation device 100, which comprises:
[0099] The real-time virtual driving scene generation module 11 is configured to identify a plurality of markers in the current driving environment of the vehicle and generate a real-time virtual driving scene according to the identification information of the plurality of markers when a call signal is detected.
[0100] The feature value calculation module 12 is configured to calculate a feature value of each of the markers in the real-time virtual driving scene according to the identification information.
[0101] The marker screening module 13 is configured to screen the plurality of markers in the real-time virtual driving scene according to the audio features of the preset ringtone audio and the feature values, and obtain a plurality of target markers.
[0102] The vehicle-mounted video ringtone generation module 14 is configured to perform stylization processing on the plurality of target markers in the real-time virtual driving scene according to the audio features, and generate and display a vehicle-mounted video ringtone.
[0103] As a preferred solution, the feature value calculation module 12 is configured to calculate a feature value of each of the markers in the real-time virtual driving scene according to the identification information, and specifically includes:
[0104] The distance information and the pixel size value of each of the markers are obtained from the identification information.
[0105] The feature value of each of the markers in the real-time virtual driving scene is determined according to the product of the distance information and a first weight value and the product of the pixel size value and a second weight value, wherein the sum of the first weight value and the second weight value is 1, and the first weight value and the second weight value are pre-set or change with the change of the vehicle driving scene.
[0106] As a preferred solution, when the first weight value and the second weight value change with the change of the vehicle driving scene, the feature value calculation module 12 adjusts the first weight value and the second weight value in the following manner:
[0107] When the vehicle driving scene is identified as a non-driving scene or a red light waiting scene, the first weight value is adjusted to a first preset weight threshold, and the second weight value is adjusted to a second preset weight threshold, wherein the first preset weight threshold is less than the second preset weight threshold.
[0108] When the vehicle driving scene is identified as a driving scene or a green light passing scene, the first weight value is adjusted to a third preset weight threshold, and the second weight value is adjusted to a fourth preset weight threshold, wherein the third preset weight threshold is greater than the fourth preset weight threshold.
[0109] As a preferred solution, the identifier screening module 13 is configured to screen the identifiers in the real-time virtual driving scene according to the preset audio features of the current ringtone audio and the feature values, and obtain target identifiers, specifically including:
[0110] obtain the total number of scales from the audio features;
[0111] sort the feature values of the identifiers according to a preset sorting rule;
[0112] screen the identifiers in the real-time virtual driving scene according to the total number of scales and the sorting order of the feature values of the identifiers, and obtain the target identifiers; wherein the number of the target identifiers is equal to or less than the total number of scales.
[0113] As a preferred solution, the vehicle-mounted video ringtone generation module 14 is configured to style the target identifiers in the real-time virtual driving scene according to the audio features, specifically including:
[0114] obtain the occurrence frequency and the playing time point information of each note from the audio features;
[0115] determine the target identifier corresponding to each note in the real-time virtual driving scene according to the occurrence frequency and the sorting order of the feature values of the target identifiers;
[0116] when detecting the playing of the current ringtone audio, style the target identifier corresponding to each note in the real-time virtual driving scene according to the playing time point information of each note in sequence.
[0117] As a preferred solution, the vehicle-mounted video ringtone generation module 14 is configured to style the target identifier corresponding to each note in the real-time virtual driving scene according to the playing time point information of each note in sequence, specifically including:
[0118] style the target identifier corresponding to the note at the current playing time point in the real-time virtual driving scene according to a preset image style based on the playing time point information of each note;
[0119] adjust the styling attribute of the target identifier corresponding to the note at the current playing time point in the real-time virtual driving scene according to a preset strength level of each note; wherein the strength level of each note is obtained from the audio features.
[0120] As a preferred solution, the audio feature comprises a ringback tone number of the ringback tone audio corresponding to the audio feature; the device further comprises an audio feature acquisition module configured to:
[0121] receiving a to-be-processed ringback tone number sent by a ringback tone server;
[0122] based on a plurality of pre-stored audio features, determining whether a ringback tone number in any one of the audio features matches the to-be-processed ringback tone number;
[0123] when the ringback tone number in any one of the audio features matches the to-be-processed ringback tone number, taking the any one of the audio features as an audio feature of a current ringback tone audio;
[0124] wherein the audio feature is specifically string information received from the ringback tone server, and the string information is constructed by the ringback tone server based on a set ringback tone audio of a current vehicle user, after analyzing scale total number, occurrence frequency of each note, playing time point information and strength level of the set ringback tone audio.
[0125] The vehicle-mounted video ringback tone generation device 100 provided by the embodiment of the present application can generate a real-time virtual driving scene by recognizing a plurality of markers in the surrounding environment of the current vehicle when detecting an incoming call signal, can then calculate characteristic values of the markers according to the recognition information of the markers, and can further perform stylization processing on the markers in the real-time virtual driving scene in combination with the audio feature of the current ringback tone audio, so as to generate a vehicle-mounted video ringback tone matching the surrounding environment of the current vehicle and the current ringback tone audio, realize the presentation of a vehicle-mounted video ringback tone with rich effects by using a vehicle-mounted system, and improve the call experience of the user during driving.
[0126] Referring to Figure 6 , Figure 6 is a structural schematic diagram of the terminal device 200 provided by the embodiment of the present application. The third aspect of the embodiment of the present application provides a terminal device 200, which comprises a memory 22, a processor 21, and a computer program stored in the memory 22 and capable of running on the processor 21, and the processor 21 implements the vehicle-mounted video ringback tone generation method as described in any one of the embodiments of the first aspect when executing the computer program.
[0127] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 22 and executed by the processor 21 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the terminal device 200.
[0128] The terminal device 200 can include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that the schematic diagram is only an example of the terminal device 200, and does not constitute a limitation on the terminal device 200, and can include more or fewer components than the diagram, or combine certain components, or different components, for example, the terminal device 200 can also include an input / output device, a network access device, a bus, etc.
[0129] The processor 21 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor 21 is the control center of the terminal device 200, and connects all parts of the terminal device 200 through various interfaces and lines.
[0130] The memory 22 can be used to store computer programs and / or modules, and the processor 21 realizes various functions of the terminal device 200 by running or executing computer programs and / or modules stored in the memory 22, and calling data stored in the memory 22. The memory 22 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 22 can include a high-speed random access memory, and can also include a non-volatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0131] The modules / units of the terminal device 200 are stored in a computer readable storage medium if they are realized in the form of software function units and sold or used as independent products. Based on this understanding, all or part of the processes in the above-mentioned embodiments can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by the processor 21, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0132] The above is the preferred embodiment of the present application. It should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.
Claims
1. A method for generating in-vehicle video ringback tones, characterized in that, include: When an incoming call signal is detected, several landmarks in the current vehicle's surrounding environment are identified, and a real-time virtual driving scene is generated based on the identification information of the landmarks. Based on the identification information, the feature value of each of the markers in the real-time virtual driving scene is calculated; Based on the preset audio characteristics and feature values of the current ringback tone audio, several of the markers in the real-time virtual driving scene are filtered to obtain several target markers; Based on the audio features, several target markers in the real-time virtual driving scene are stylized to generate in-vehicle video ringback tones and display them.
2. The method for generating vehicle-mounted video ringback tones as described in claim 1, characterized in that, The step of calculating the feature value of each of the markers in the real-time virtual driving scene based on the identification information specifically includes: The distance information and pixel size value of each of the identifiers are obtained from the identification information; The feature value of each of the markers in the real-time virtual driving scene is determined based on the product of the distance information and the first weight value, and the product of the pixel size value and the second weight value; wherein the sum of the first weight value and the second weight value is 1; the first weight value and the second weight value are preset or change with the change of the vehicle driving scene.
3. The method for generating vehicle-mounted video ringback tones as described in claim 2, characterized in that, When the first weight value and the second weight value change with the vehicle driving scenario, the first weight value and the second weight value are adjusted in the following way: When the vehicle driving scenario is identified as a non-driving scenario or a red light waiting scenario, the first weight value is adjusted to a first preset weight threshold, and the second weight value is adjusted to a second preset weight threshold; wherein, the first preset weight threshold is less than the second preset weight threshold; When the vehicle driving scenario is identified as a driving scenario or a green light passage scenario, the first weight value is adjusted to the third preset weight threshold, and the second weight value is adjusted to the fourth preset weight threshold; wherein, the third preset weight threshold is greater than the fourth preset weight threshold.
4. The method for generating vehicle-mounted video ringback tones as described in claim 1, characterized in that, The step of filtering several markers in the real-time virtual driving scene based on the preset audio characteristics and feature values of the current ringback tone audio to obtain several target markers specifically includes: Obtain the total number of musical scales from the audio features; The feature values of several of the identifiers are sorted according to a preset sorting rule; Based on the total number of musical scales and the sorting order of the feature values of the aforementioned markers, the aforementioned markers in the real-time virtual driving scene are filtered to obtain a number of target markers; wherein the number of target markers is equal to or less than the total number of musical scales.
5. The method for generating vehicle-mounted video ringback tones as described in claim 4, characterized in that, The stylization processing of several target markers in the real-time virtual driving scene based on the audio features specifically includes: The frequency of occurrence and playback time information of each note are obtained from the audio features; Based on the frequency of occurrence and the sorting order of the feature values of several target identifiers, the target identifier corresponding to each musical note in the real-time virtual driving scene is determined; When the current ringback tone audio is detected to be playing, based on the playback time information of each note, the target identifier corresponding to each note in the real-time virtual driving scene is stylized sequentially.
6. The method for generating vehicle-mounted video ringback tones as described in claim 5, characterized in that, Based on the playback time information of each note, the target identifier corresponding to each note in the real-time virtual driving scene is stylized sequentially, specifically including: Based on the playback time information of each note, the target markers corresponding to the notes at the current playback time in the real-time virtual driving scene are stylized according to a preset image style. Based on the preset intensity level of each note, the stylistic attributes of the target marker corresponding to the note at the current playback time point in the real-time virtual driving scene are adjusted; wherein, the intensity level of each note is obtained from the audio features.
7. The method for generating vehicle-mounted video ringback tones as described in any one of claims 1 to 6, characterized in that, The audio features include the ringback tone number of the corresponding ringback tone audio; therefore, the method specifically obtains the audio features of the current ringback tone audio through the following steps: Receive the pending ringback tone number sent by the ringback tone server; Based on several pre-stored audio features, determine whether there exists a ringback tone number in any audio feature that matches the ringback tone number to be processed; When any one of the audio features contains a ringback tone number that matches the ringback tone number to be processed, that one audio feature is used as the audio feature of the current ringback tone audio. Specifically, the audio features are string information received from the ringback tone server. The string information is constructed by the ringback tone server based on the set ringback tone audio of the current vehicle user, after analyzing the total number of scales, the frequency of each note, the playback time information, and the intensity level of the set ringback tone audio.
8. A vehicle-mounted video ringback tone generator, characterized in that, include: The real-time virtual driving scene generation module is used to identify several landmarks in the current vehicle's surrounding environment when an incoming call signal is detected, and to generate a real-time virtual driving scene based on the identification information of the landmarks. The feature value calculation module is used to calculate the feature value of each of the markers in the real-time virtual driving scene based on the identification information; The identifier filtering module is used to filter several identifiers in the real-time virtual driving scene according to the audio characteristics and the feature values of the preset ringback tone audio, and obtain several target identifiers; The in-vehicle video ringback tone generation module is used to stylize several target markers in the real-time virtual driving scene according to the audio features, generate in-vehicle video ringback tones, and display them.
9. A terminal device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the in-vehicle video ringback tone generation method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, It includes a computer program / instructions that, when executed by a processor, implement the steps of the in-vehicle video ringback tone generation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Video polyphonic ringtone interaction method, server and storage medium
CN116055638A
Video polyphonic ringtone interactive playing method based on user environment
CN116095238A