Artificial intelligence-based memory saving and memory retrieval method and related device

By collecting and processing video and audio data from the user's surroundings, generating and storing memory summaries, the problem of user memory omissions and errors is solved, achieving efficient memory preservation and retrieval, and improving the user's work and life efficiency.

CN117033556BActive Publication Date: 2025-12-12SUGR ELECTRONICS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311051700.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-18
Publication Date
2025-12-12
Estimated Expiration
2043-08-18

AI Technical Summary

Technical Problem

In existing technologies, there are problems with users missing or misremembering information, and existing tools are not convenient to carry and use.

Method used

By collecting video and audio data from the user's scene, using scene change detection and rate prediction algorithms for dynamic frame-by-frame acquisition, and combining large language models to process image and audio information, a memory summary is generated and stored in a database. The memory is then retrieved by receiving user queries.

Benefits of technology

It reduces user memory omissions and errors, improves work efficiency and quality of life, forms a closed loop for memory storage and retrieval, and makes it convenient for users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117033556B_ABST
    Figure CN117033556B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, in particular to a memory saving and memory extracting method based on artificial intelligence and related equipment. The application collects video data and audio data of a scene where a user is located, respectively processes the video data and the audio data to obtain image information and audio information, and then performs classification and identification to obtain a memory original text; a large language model is called to process the memory original text, and a memory abstract obtained through processing and the memory original text are stored in a database. When a query question of a user about a past event is received, a memory abstract corresponding to the query question is queried and output in the database. The application can fill the deficiency of user memory, reduce the possibility of user memory omission and memory error, reduce the brain burden of the user, improve work efficiency and life quality; in addition, a query question input by the user is output to form a closed loop of memory saving and memory extracting, so that the user can use the application conveniently and work and life efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a memory saving and memory extraction method based on artificial intelligence and related equipment. BACKGROUND

[0002] The emergence and popularity of the Internet have brought a large amount of information to users, making humans need to process more and more information, leading to the phenomenon of missing memory information. With the continuous development of artificial intelligence, although many tools have appeared to help humans complete memory, such as paper notebooks, note software on computers, and conference summary software for video conferences, these all need humans to actively open and use, and there are problems of inconvenience in carrying and using. SUMMARY

[0003] Therefore, the present application provides a memory saving and memory extraction method based on artificial intelligence and related equipment to fill the lack of user memory, reduce the possibility of user memory omission and memory error, and solve the technical problems of inconvenience in carrying and using existing technologies.

[0004] The first aspect of the present application provides a memory saving and memory extraction method based on artificial intelligence, which comprises:

[0005] collecting video data and audio data of a scene where a user is located;

[0006] processing the video data to obtain image information and processing the audio data to obtain audio information;

[0007] classifying and identifying the image information and the audio information to obtain memory raw text;

[0008] calling a large language model to process the memory raw text, and storing a memory summary obtained by processing and the memory raw text in a database;

[0009] when receiving a query question of the user, querying and outputting a memory summary corresponding to the query question in the database.

[0010] In an optional embodiment, the processing of the video data to obtain image information comprises:

[0011] dynamic frame acquisition of the video data by combining a scene transformation detection algorithm and a rate prediction algorithm to obtain a plurality of image data;

[0012] content segmentation of each of the image data to obtain image data blocks;

[0013] image recognition of the image data blocks to obtain the image information.

[0014] In an optional implementation, the combination of the scene change detection algorithm and the rate prediction algorithm dynamically frame-captures the video data to obtain a plurality of image data, including:

[0015] The scene change detection algorithm is used to detect the scene of the video data to obtain a video scene type;

[0016] The rate prediction algorithm is used to predict the adaptive transform rate of the video data corresponding to each video scene type;

[0017] When the predicted transform rate is higher than a preset rate threshold, a first preset frame rate is used to capture the frame rate of the video data to obtain a plurality of image data corresponding to the video scene type;

[0018] When the predicted transform rate is lower than the preset rate threshold, a second preset frame rate is used to capture the frame rate of the video data to obtain a plurality of image data corresponding to the video scene type;

[0019] The first preset frame rate is greater than the second preset frame rate.

[0020] In an optional implementation, the processing of the audio data to obtain audio information includes:

[0021] The audio data is frame-captured to obtain a plurality of sub-audio data;

[0022] The scene change detection algorithm is used to detect whether the capture scene of the audio data changes;

[0023] When the capture scene of the audio data changes, the sub-audio data whose capture scene changes is classified to obtain an audio scene type;

[0024] Each of the sub-audio data is audio-layered to obtain layered audio;

[0025] The layered audio is audio-recognized to obtain the audio information.

[0026] In an optional implementation, the classification and recognition of the image information and the audio information to obtain a memory original text includes:

[0027] The image information is classified and recognized to obtain image text, the audio information is classified and recognized to obtain audio text, and the image text and the audio text are semantically associated to obtain the memory original text.

[0028] In an optional implementation, the semantic association of the image text and the audio text includes:

[0029] semantically associate the image text and the audio text based on a scene or a time or a location or a theme to structure and merge the image text and the audio text.

[0030] In an optional implementation, the method further includes:

[0031] According to the video scene type, the corresponding image information is classified and compressed and stored; and

[0032] The audio scene type and the corresponding audio information are stored.

[0033] In an optional implementation, when the query question is a voice query question input by the user in a voice form, the querying and outputting, in the database, of the memory summary corresponding to the query question includes:

[0034] voice recognizing the voice query question to obtain a text query question;

[0035] querying and outputting, in the database, of the memory summary corresponding to the text query question.

[0036] A second aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the memory saving and memory extracting method based on artificial intelligence when executing the computer program.

[0037] A third aspect of the present application provides a computer readable storage medium, having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps of the memory saving and memory extracting method based on artificial intelligence.

[0038] In summary, the memory saving and memory extracting method based on artificial intelligence and related devices provided by the embodiments of the present application collect video data and audio data of a scene where a user is located, and respectively process the video data and the audio data to obtain image information and audio information, classify and identify the image information and the audio information to obtain memory original text, call a large language model to process the memory original text, and store the memory summary obtained by processing and the memory original text in a database, which can fill the deficiency of user memory, reduce the possibility of user memory omission and memory error, reduce the mental burden of the user, improve work efficiency and life quality. In addition, the query question input by the user is output to form a closed loop of memory saving and memory extraction, which can facilitate the user to use and improve the work and life efficiency of the user. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 is a flowchart of a memory saving and memory extraction method based on artificial intelligence shown in an embodiment of the present application;

[0040] Figure 2 is a data flowchart of processing video data shown in an embodiment of the present application;

[0041] Figure 3 is a flowchart of processing video data shown in an embodiment of the present application;

[0042] Figure 4 is a data flowchart of processing audio data shown in an embodiment of the present application;

[0043] Figure 5 is a flowchart of processing audio data shown in an embodiment of the present application;

[0044] Figure 6 is a structural diagram of an electronic device shown in an embodiment of the present application;

[0045] Figure 7 is a structural diagram of another electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION

[0046] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or," as used herein, refers to any or all possible combinations of one or more of the associated listed items.

[0047] Hereinafter, the terms "first" and "second" are used only for the purpose of description and can not be understood as implying or suggesting relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise specified.

[0048] Figure 1 is a flowchart of a memory saving and memory extraction method based on artificial intelligence shown in an embodiment of the present application. The memory saving and memory extraction method based on artificial intelligence can be executed by an electronic device, and specifically includes the following steps.

[0049] S11, video data and audio data of a scene where a user is located are collected.

[0050] The data of the scene where the user is in is collected in real time as memory and stored, thereby playing a memory saving role. When the user forgets or is unclear, the stored data can be searched or retrieved, thereby playing a memory extraction role. The data of the scene where the user is in includes video data and audio data of the scene where the user is in.

[0051] The video data of the scene where the user is in can include visual information in front of the user and all image information browsed by the user on a display of a user terminal device (for example, a computer). The audio data of the scene where the user is in can include auditory information in front of the user and all sound information received by a microphone of the user terminal device (for example, a computer). Collecting the video data and the audio data of the scene where the user is in is equivalent to collecting all visual information that can be seen by the user and all auditory information that can be heard by the user.

[0052] In some embodiments, the video data of the scene where the user is in can be collected by using an image collection device. The image collection device can be built in an electronic device, for example, a camera of the electronic device, or can be independent of the electronic device. When the image collection device is independent of the electronic device, the image collection device can transmit the collected video data of the scene where the user is in to the electronic device in a wired or wireless manner.

[0053] In some embodiments, the audio data of the scene where the user is in can be collected by using an audio collection device. The audio collection device can be built in an electronic device, for example, a microphone of the electronic device, or can be independent of the electronic device. When the audio collection device is independent of the electronic device, the audio collection device can transmit the collected audio data of the scene where the user is in to the electronic device in a wired or wireless manner.

[0054] In some embodiments, the electronic device can be a bracelet type, glasses type, necklace type, helmet type or handheld type device, as long as the device can implement the memory saving and memory extraction method based on artificial intelligence provided in the embodiments of the present application, regardless of the structure form, and can be included in the present application.

[0055] S12, processing the video data to obtain image information, and processing the audio data to obtain audio information.

[0056] When the video data is processed, the video data is processed in combination with Figure 2 and Figure 3 As shown in the figures, the method for processing the video data to obtain image information specifically includes the following steps:

[0057] S21, combining a scene transformation detection algorithm and a rate prediction algorithm to dynamically frame the video data to obtain a plurality of image data.

[0058] Since it is considered that the video data is not used for watching or communication, the frame rate can be set to one frame per second or multiple frames per second, i.e. the video data is read frame by frame in a loop, one frame per second or multiple frames per second, to realize dynamic frame separation acquisition of the video data.

[0059] In an optional embodiment, the dynamic frame separation acquisition of the video data by combining the scene change detection algorithm and the rate prediction algorithm obtains a plurality of image data, including:

[0060] applying a scene change detection algorithm to the video data to obtain a video scene type;

[0061] applying a rate prediction algorithm to the video data corresponding to each video scene type to perform adaptive transform rate prediction;

[0062] when the predicted transform rate is higher than a preset rate threshold, a first preset frame rate is used to acquire the video data to obtain a plurality of image data corresponding to the video scene type;

[0063] when the predicted transform rate is lower than the preset rate threshold, a second preset frame rate is used to acquire the video data to obtain a plurality of image data corresponding to the video scene type;

[0064] wherein the first preset frame rate is greater than the second preset frame rate.

[0065] The embodiments of the present application can set an initial acquisition frame rate according to the required application scenarios and system requirements, for example, 15 frames per second. The initial acquisition frame rate is used to perform frame separation acquisition on the video data, and the time stamp and frame rate information of each acquired image data are recorded.

[0066] Meanwhile, the collected image data is analyzed to detect whether the collection scene of the video data has changed using scene change detection. Background difference method, inter-frame difference method, optical flow method, background modeling method, feature point matching, etc. can be used to detect whether the collection scene has changed, such as object movement, light change, appearance / disappearance of objects, etc. For example, the electronic device can obtain a weighted mean of color information of the image data, and calculate the inter-frame difference of the two image data according to the weighted mean of the color information. The color information can include, but is not limited to, brightness information, chroma information, saturation information, texture information, etc. The weighted mean of the color information of the image data can be obtained by converting the image data from RGB color space to HSV (hue, saturation, brightness) color space, extracting the value of the brightness channel of each pixel, assigning a weight to the brightness value of each pixel, multiplying the brightness value of each pixel by its corresponding weight to obtain a plurality of weighted values, and then adding the plurality of weighted values and dividing by the sum of the total weights. The inter-frame difference of the two image data can be obtained. When it is determined that the inter-frame difference is greater than a preset difference threshold, it indicates that there is a large difference between the two image data, and it is determined that the collection scene of the video data has changed. When it is determined that the inter-frame difference is less than the preset difference threshold, it indicates that the two image data are almost the same or have a small difference, and it is determined that the collection scene of the video data has not changed.

[0067] In the current collection scene, the logic of triggering the change rate prediction and the collection frame rate adjustment is triggered. When the collection scene changes, the logic of triggering the change rate prediction and the collection frame rate adjustment is triggered again. Machine learning, statistical analysis or other algorithms are used to predict the change rate of the scene. The frame rate, change frequency, change amplitude, etc. data in the past period of time are analyzed and modeled to estimate the future change rate. Regression analysis, time series analysis or model-based prediction methods of historical data can be used to predict the change rate.

[0068] According to the prediction result of the change rate, it is determined whether the collected frame rate should be increased or decreased. If the predicted change rate is fast, it means that the video data changes frequently, and a faster frame rate should be used for collection to provide smoother pictures. If the predicted change rate is slow, it means that the video data changes slowly, and the frame rate can be reduced for collection to save storage space and processing resources.

[0069] For example, in the current scenario, adaptive transformation rate prediction is performed, and fast frame rate (first preset frame rate) is used to collect video data for fast transformation rate (i.e., the predicted transformation rate is higher than the preset rate threshold), for example, when driving on the road, the frame rate of 120 frames per second can be used for video data collection for the front window of the car; slow frame rate (second preset frame rate) is used to collect video data for slow transformation rate (i.e., the predicted transformation rate is lower than the preset rate threshold), for example, when reading a book, turning a page every minute, a frame rate of 1 frame per minute can be used for video data collection. When detecting that the collection scene of the video data has changed, it is marked as a new scene type, and for the new scene type, adaptive transformation rate prediction is performed, and fast frame rate is used to collect for fast transformation rate, and slow frame rate is used to collect for slow transformation rate. Repeat this until the user turns off the image collection device, and obtain multiple video scene types, each of which corresponds to multiple image data.

[0070] In the above embodiments of the present application, the collection of video data is performed, and the scene transformation detection algorithm is used to detect whether the collection scene of the video data has changed. When it is detected that the collection scene has changed, the logic of transformation rate prediction and collection frame rate adjustment is triggered. The embodiments of the present application combine real-time video data collection, scene transformation detection, and transformation rate prediction to dynamically adjust the collection frame rate to adapt to changes in different scenes. For dynamic scenes with fast transformation, a faster frame rate is used for collection, which can accurately capture the details of the change of the collection scene and improve the storage efficiency and the quality of subsequent video data processing; for static or slow-changing scenes, a slower frame rate is used for collection, which can reduce unnecessary image data processing and computing resource usage, save storage space and transmission bandwidth, and thus improve the efficiency of processing video data.

[0071] It should be noted that scene classification is triggered only when there is scene transformation, and scene classification is not triggered when there is no scene transformation. Scene classification includes scene recognition. The video scene type obtained by performing scene classification on the image data can be flexibly designed and adjusted according to specific tasks or user needs.

[0072] In some embodiments, a scene classification model can be trained using a support vector machine (SVM), a random forest, a deep neural network, etc. The electronic device can use the trained scene classification model to perform scene classification on the image data in which the collection scene has changed to obtain the scene type. For the sake of distinction from the following, the scene type obtained by performing scene classification on the image data is referred to as the video scene type. The video scene type can include, but is not limited to, text, people, scenery, etc.

[0073] For example, when the image data contains text (text identification, billboards, etc.), the video scene type of the image data is a text type; when the image data contains a person (a person region in the image data can be identified using a face detection or object detection algorithm), the video scene type of the image data is a person type; when the image data does not contain obvious text and persons, for example, contains scenery, buildings, natural environment, etc., the video scene type of the image data is a scenery type.

[0074] In some embodiments, after obtaining the plurality of image data, the plurality of image data can be processed, for example, size normalization, image denoising, image enhancement, etc., to improve the accuracy and efficiency of detecting a change in the collection scene of the video data, and to improve the accuracy and efficiency of classifying the scene of the image data.

[0075] S22, performing content segmentation on each of the image data to obtain image data blocks.

[0076] The electronic device can use a pre-stored content segmentation algorithm to perform content segmentation on each of the image data, thereby extracting one or more object blocks in each image data. The content segmentation algorithm can include, but is not limited to, semantic segmentation, instance segmentation, etc.

[0077] For example, assuming that a certain image data corresponds to a library scene, after content segmentation of the image data by a content segmentation algorithm, three image data blocks can be obtained, one of which contains books corresponding to the foreground, one of which contains a table corresponding to the middle ground, and the other of which contains a library corresponding to the background.

[0078] S23, performing image recognition on the image data blocks to obtain the image information.

[0079] The electronic device can input each of the segmented image data blocks into a trained image recognition model for recognition to obtain corresponding image information. For example, the image information recognized from a person image data block is “person”, the image information recognized from a car image data block is “car”, the image information recognized from an animal image data block is “animal”, etc.

[0080] It should be understood that different recognition methods can be used for different types of image data blocks, for example, face recognition technology can be used for person image data blocks, vehicle model recognition technology can be used for car image data blocks, animal species classification technology can be used for animal image data blocks, etc.

[0081] In an optional embodiment, the electronic device stores the video scene type and the corresponding image information.

[0082] The electronic device can store the image information according to the video scene type. That is, the image information with the same video scene type is stored in the same location, and the image information with different video scene types is stored in different locations. For example, assuming that the first image data to the third image data correspond to the same video scene type, for example, a person type, the image information of the first image data to the third image data and the corresponding video scene type (person type) are stored. Assuming that the fourth image data and the fifth image data correspond to the same video scene type, for example, a conversation type, the image information of the fourth image data and the fifth image data and the corresponding video scene type (conversation type) are stored.

[0083] In an optional embodiment, the electronic device can also compress the corresponding image information according to the video scene type. For example, assuming that the image information is text, such as a book, a newspaper, a business card, or a license plate, a compression algorithm suitable for text is used for compression. Assuming that the image information is a scene, a general compression algorithm is used for compression.

[0084] Storing the video scene type and the corresponding image information can realize the structured storage of the image information.

[0085] The purpose of classifying and compressing the image information is to select the most suitable compression method for image information of different scene types, thereby reducing storage.

[0086] When processing the audio data, the audio data is processed in combination with the video data. Figure 4 and Figure 5 As shown in the figures, the method for processing the audio data to obtain audio information specifically includes the following steps:

[0087] S41, the audio data is collected by frame, and a plurality of sub-audio data is obtained.

[0088] The electronic device can preset an audio acquisition frame rate, and collect audio data according to the preset audio acquisition frame rate. Frame collection of audio data means that continuous audio signals are divided into a plurality of short period audio data, and each short period audio data is sub-audio data. Audio data usually exists in the form of continuous analog signals, and is sampled and discretized into digital signals in an audio acquisition device.

[0089] S42, whether the acquisition scene of the audio data changes is detected according to a scene change detection algorithm.

[0090] The scene change detection algorithm can include, but is not limited to, a statistical-based method, a machine learning algorithm (e.g., a support vector machine, a decision tree, a random forest, etc.), and a deep learning algorithm (e.g., a convolutional neural network, etc.).

[0091] In other embodiments, the electronic device can further extract features from the sub-audio data for detecting whether the collection scene of the audio data has changed. The features can include, but are not limited to, time-domain features (e.g., volume, energy, etc.), frequency-domain features (e.g., spectral centroid, spectral average energy, etc.), time-frequency features (e.g., short-time Fourier transform coefficients, etc.), and the like. Whether the collection scene of the audio data has changed is detected based on the extracted features according to a scene change detection algorithm.

[0092] S43, when the collection scene of the audio data changes, performing scene classification on the sub-audio data whose collection scene has changed to obtain an audio scene type.

[0093] The scene type obtained by performing scene classification on the sub-audio data is referred to as an audio scene type.

[0094] In some embodiments, the electronic device can utilize a machine learning algorithm (e.g., a support vector machine, a decision tree, a random forest, etc.) to extract a relevant feature vector from each of the sub-audio data for scene classification to obtain an audio scene type. The audio scene type can include, but is not limited to, a conversation class, a music class, a scene class, and the like.

[0095] For example, assuming that the sub-audio data contains people's conversation sound (e.g., a telephone call, a meeting discussion, etc.), the audio scene type is determined to be a conversation class. The conversation class audio data usually has obvious speech features, such as time-frequency features of speech, speech rate, tone, etc. Assuming that the sub-audio data contains audio representing music playing, playing, singing, etc., the audio scene type is determined to be a music class. The music class audio data usually has unique spectral features, rhythm, and instrument sound. Assuming that the sub-audio data contains audio representing a background environment (e.g., city street noise, natural environment sound, traffic sound, etc.), the audio scene type is determined to be a scene class. The scene class audio data usually contains environmental noise, sound texture, and the like.

[0096] It should be noted that scene classification is triggered only when there is a scene change, otherwise, scene classification is not triggered when there is no scene change, and scene classification includes scene recognition. The audio scene type obtained by performing scene classification on the sub-audio data can be flexibly designed and adjusted according to specific tasks or user requirements.

[0097] S44, performing audio layering on each of the sub-audio data to obtain layered audio.

[0098] In some embodiments, the sub-audio data can be audio layered based on the distribution characteristics of the audio signals in time or the scene context that people often encounter in daily life. The audio layering can be divided into front layer-conversation, middle layer-music, and back layer-background. Each of the sub-audio data can be input into a pre-trained audio layering recognition model, and the audio layering result of each of the sub-audio data is output by the audio layering recognition model, i.e., it is determined whether the sub-audio data belongs to the front layer-conversation, the middle layer-music, or the back layer-background.

[0099] For example, it is assumed that the sub-audio data is a scene in which a person has a conversation in a cafe. In this scene, the front layer-conversation can be the actual conversation sound (e.g., language, speaking sound, etc.) of the people; the middle layer-music can be the background music played in the cafe; and the back layer-background can be the background noise and environmental sound (e.g., the sound of people's footsteps, the sound of the coffee machine, the echo of the environment, etc.).

[0100] Through the above optional implementation, since the audio features of different levels of audio data correspond to different scene elements, by dividing the sub-audio data into three layers of front, middle, and back, or even more layers, and processing the audio information of the specified layer, the processing efficiency and quality are improved. For example, if it is intended to extract the conversation content of the people in the cafe, the front layer-conversation part can be focused on, and the influence of the middle layer-music and the back layer-background is reduced during processing.

[0101] S45, performing audio recognition on the layered audio to obtain audio information.

[0102] The electronic device can input each of the layered audio into a trained audio recognition model (e.g., a convolutional neural network (CNN), a recurrent neural network (RNN), and a Transformer based on deep learning, etc.) for recognition to obtain the corresponding audio information. The audio information can include the content and features of the audio data, etc.

[0103] For example, it is assumed that a certain layered audio is music data, and the electronic device can output "popular music", "rock music", "classical music", etc. as the recognition result through the trained audio recognition model.

[0104] In some implementations, the electronic device can also obtain the audio information of the layered audio, such as the title of the song and the name of the artist, etc., through the trained audio recognition model.

[0105] In an optional implementation, the electronic device stores the audio scene type and the corresponding audio information.

[0106] The electronic device can store the audio information according to the audio scene type. That is, audio information with the same audio scene type is stored in the same location, and audio information with different audio scene types is stored in different locations.

[0107] Storing the audio scene type and the corresponding audio information can achieve structured storage of audio information.

[0108] S13, classifying and identifying the image information and the audio information to obtain a memory original text.

[0109] The image information obtained by processing the video data and the audio scene type corresponding to each image information, and the audio information obtained by processing the audio data and the audio scene type corresponding to each audio information are stored in the memory arrangement module of the electronic device, so that the memory arrangement module of the electronic device classifies and identifies the image information and the audio information to obtain a memory original text.

[0110] The memory arrangement module of the electronic device classifies and identifies the image information and the audio information to obtain a memory original text can include classifying and identifying the image information to obtain image text, classifying and identifying the audio information to obtain audio text, and semantically associating the image text and the audio text to obtain the memory original text.

[0111] In some embodiments, the electronic device can extract image text from image information and audio text from audio information using natural language processing technology, speech recognition technology, etc., that is, convert image information and audio information into text description, text label, text keyword, etc. The memory original text can be a descriptive text related to image information or audio information, used to record key information, content summary or identification information. For example, classifying and identifying image information can obtain category labels such as "person", "scene", "object", etc.; classifying and identifying audio information can obtain sound type labels such as "talking", "music", etc.

[0112] In some embodiments, the image text and the audio text can be semantically associated based on the scene or the time or the place or the theme to structurally combine the image text and the audio text.

[0113] The semantic association of the image text and the audio text based on the scene or the time or the location or the theme refers to matching, corresponding or associating the image text and the audio text based on the scene or the time or the location or the theme, so as to realize semantic connection between the image text and the audio text. The semantic association of the image text and the audio text obtains the merged data after the structured combination of the image text and the audio text, as the memory original text. The merged data can be in a text format (for example, XML, JSON, etc.).

[0114] For example, it is assumed that the video data shows an indoor environment of a cafe. According to the image information, visual features can be extracted, which can include a cafe, coffee, an indoor environment, etc. Meanwhile, the name of the cafe is mentioned in the audio data, and audio features can be extracted from the audio information, which can include the name of the cafe, the price of coffee, etc. Since the visual features and the audio features both include “cafe”, the image information and the audio information can be semantically associated according to the image information and the audio information.

[0115] Through the above optional implementation, through semantic association, the image information and the audio information can be associated to a common semantic concept, and the image information and the audio information can be structured and combined into a unified data storage form, which helps to integrate and share multi-modal data, thereby improving the accuracy and relevance of search results, and providing users with richer and more accurate data description and analysis results.

[0116] In S14, a large language model is called to process the memory original text, and a memory summary obtained by processing is stored in the database together with the memory original text.

[0117] The large language model refers to a model trained through machine learning and artificial intelligence technology, which is used to understand and generate natural language text. The large language model can realize semantic understanding, text generation, question and answer, etc. through training on large-scale text data, and has strong language processing capability.

[0118] The electronic device can periodically (every minute or every 10 minutes, the period can be determined according to the amount of data generated by the change of the collection scene) call the application programming interface (Application Programming Interface, API) of the large language model, or directly call the locally stored large language model to process the memory original text obtained by the memory arrangement module, to obtain a memory summary.

[0119] The electronic device stores the memory summary and the memory original text in the database at the same time, and associates the memory summary and the memory original text for future retrieval.

[0120] S15, upon receiving the query question of the user, querying in the database and outputting the memory summary corresponding to the query question.

[0121] The user can input the query question in the electronic device through voice input or keyboard input. The query question is a query question about past events, which means that when the user forgets, the user can query according to the visual information and auditory information collected by the user in the past scene.

[0122] In an optional embodiment, when the user inputs the query question in the electronic device through the keyboard input, since the query question input through the keyboard is a text form query question (text query question), the electronic device can directly query in the database and output the memory summary corresponding to the text form query question. The electronic device can also output the memory original text corresponding to the memory summary.

[0123] In an optional embodiment, when the user inputs the query question in the electronic device through voice input, since the query question input through voice is a voice form query question (voice query question), the electronic device needs to first perform voice recognition on the voice form query question to obtain a text form query question, and then query in the database and output the memory summary corresponding to the text form query question.

[0124] In some embodiments, the voice query question of the user can be collected by the audio collection device or the voice recording device, and the voice query question can be converted into a text query question using voice recognition technology (for example, automatic speech recognition (ASR) technology). In other embodiments, before performing voice recognition on the voice query question, the voice query question can be preprocessed (for example, noise removal, audio level reduction, voice signal enhancement, etc.) to improve the accuracy of subsequent voice recognition on the voice query question. After voice recognition on the voice query question, the text query question obtained by recognition can be post-processed (for example, removing recognition errors, punctuation processing, etc.) to obtain a more accurate text query question.

[0125] The following are several application scenarios of the memory saving and memory extraction method based on artificial intelligence of the present application.

[0126] Application scenario one, user U1 chats with friend A today, talks about the apple company will hold a new product release conference at 1 o'clock in the morning on June 6, user U1 forgets after listening, user U1 inputs the query question as "today mentioned apple release conference", and outputs the question and answer result as "information in the conversation today: apple release conference, June 6, 1 o'clock in the morning" in the form of voice or text.

[0127] Application scenario two, user U2 had a video conference with customer B last week, the PPT of customer B in the conference shows the market capacity in the next year, user U2 forgets the market capacity, now user U2 inputs the query question as "the market capacity mentioned by customer last week", and outputs the question and answer result as market capacity: XXX in the form of voice or text.

[0128] Application scenario three, user U3 discussed with family C last week to go to XX hot pot restaurant for hot pot, user U3 forgets the location of the hot pot restaurant, now user U3 inputs the query question as "XX hot pot restaurant discussed by me last week is where", and outputs the question and answer result as "XX hot pot restaurant discussed by you last week is in XX street, XX building, XX number" in the form of voice or text. The electronic device can display the location of XX hot pot restaurant, the telephone of XX hot pot restaurant, and plan the related path, hot pot consumption package, etc. in the form of text, and also can display the pictures related to XX hot pot restaurant (such as facade photo) in the form of picture.

[0129] It should be noted that the memory saving and memory extraction method based on artificial intelligence described in the present application can be executed by an electronic device, or can be executed by a combination of cloud and electronic device, that is, the electronic device obtains video data and audio data of the scene where the user is in real time, processes the video data to obtain image information, processes the audio data to obtain audio information, the cloud runs a large language model to process, understand and extract memory summary based on the image information and the audio information, the electronic device saves the memory summary extracted by the cloud for query, so as to realize the reminding of memory.

[0130] In addition, it should be noted that the present application can only obtain audio data of the scene where the user is, and realize the saving and extraction of memory based on the audio data.

[0131] The application obtains video data and audio data of a scene where a user is located, and respectively processes the video data and the audio data to obtain image information and audio information, and then classifies and identifies the image information and the audio information to obtain a memory original text. A large language model is called to process the memory original text, and a memory summary obtained by processing is stored in a database together with the memory original text. When a query question about a past event is received from the user, a memory summary corresponding to the query question is queried from the database and output. The video data and / or audio data of the scene where the user is located are recorded and saved at any time and anywhere, which is equivalent to helping the user to add an external brain memory, filling the user's memory deficiency, reducing the possibility of memory omission and memory errors, reducing the user's mental burden, improving work efficiency and life quality; in addition, the query question input by the user is output to form a closed loop of memory saving and memory extraction, which can facilitate the user to use and improve the user's work and life efficiency.

[0132] Referring to Figure 6 As shown in FIG. 6, it is a structural schematic diagram of an electronic device according to an embodiment of the present application. In the preferred embodiment of the present application, the electronic device 6 can include a memory 61, at least one processor 62 and at least one communication bus 63.

[0133] In the embodiment, the electronic device 6 including the memory 61, the at least one processor 62 and the at least one communication bus 63 can obtain video data of a scene where a user is located from other devices, for example, from an image acquisition device, and obtain audio data of the scene where the user is located from an audio acquisition device.

[0134] In some embodiments, the memory 61 stores a computer program, and the computer program is executed by the at least one processor 62 to realize all or part of the steps of the artificial intelligence-based memory saving and memory extraction method.

[0135] In some embodiments, the at least one processor 62 is a control core of the electronic device 6, which connects various components of the entire electronic device 6 through various interfaces and lines, and executes or runs programs or modules stored in the memory 61 and calls data stored in the memory 61 to perform various functions and process data of the electronic device 6. For example, the at least one processor 62 executes the computer program stored in the memory 61 to realize all or part of the steps of the artificial intelligence-based memory saving and memory extraction method.

[0136] The at least one communication bus 63 is configured to realize the connection and communication between the memory 61, the at least one processor 62 and the like.

[0137] Referring to Figure 7 Fig. 7 is a structural schematic diagram of another electronic device according to an embodiment of the present application. The electronic device 7 can include a memory 71, at least one processor 72, a camera 73, a microphone 74, a speaker 75, a screen 76, at least one communication bus 77, and the like.

[0138] The electronic device 7 can include a user terminal device, which includes but is not limited to any electronic product that can be interacted with a user through a keyboard, a mouse, a remote controller, a touchpad, a voice control device, or the like, such as a personal computer, a tablet computer, a smart phone, a digital camera, and the like.

[0139] In some embodiments, the memory 71 stores a computer program, and the computer program, when executed by the at least one processor 72, implements all or part of the steps of the artificial intelligence-based memory saving and memory extracting method as described.

[0140] In some embodiments, the at least one processor 72 is a control unit of the electronic device 7, which connects various components of the entire electronic device 7 through various interfaces and lines, and performs various functions of the electronic device 7 and processes data by running or executing programs or modules stored in the memory 71 and calling data stored in the memory 71. For example, the at least one processor 72 executes the computer program stored in the memory 71 to implement all or part of the steps of the artificial intelligence-based memory saving and memory extracting method as described in the embodiments of the present application.

[0141] The camera 73 is used to collect video data of a scene where a user is located and transmit the video data to the at least one processor 72.

[0142] The microphone 74 is used to collect audio data of a scene where a user is located and transmit the audio data to the at least one processor 72.

[0143] The speaker 75 is used to play various voice information output by the electronic device 7.

[0144] The screen 76 can display various information output by the electronic device 7, and can also be used to receive touch operations of a user, and the like.

[0145] The at least one communication bus 77 is configured to realize connection and communication between the memory 71, the at least one processor 72, the camera 73, the microphone 74, the speaker 75, and the screen 76, and the like.

[0146] Those skilled in the art should understand that, Figure 6 and the like.Figure 7 The structure of the electronic device shown is not a limitation of the embodiments of the present application, and can be a bus type structure or a star type structure. The electronic device 6 and the electronic device 7 can also include more or less other hardware or software or different component arrangements than shown.

[0147] It should be noted that the electronic device 6 and the electronic device 7 are only examples, and other existing or future electronic products can also be applicable to the present application and should be included in the protection scope of the present application by reference.

[0148] Although not shown, the electronic device 6 and the electronic device 7 can also include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor, for example, the at least one processor 62 or the at least one processor 72, through a power management device, so as to realize functions such as management of charging, discharging, and power consumption management through the power management device. The power supply can also include one or more direct current or alternating current power supplies, recharging devices, power supply fault detection circuits, power supply converters or inverters, power supply status indicators, and any other components. The electronic device 6 and the electronic device 7 can also include various sensors, Bluetooth modules, Wi-Fi modules, and the like, which are not described here.

[0149] The integrated units in the form of software function modules described above can be stored in a computer readable storage medium. The software function modules described above are stored in a storage medium and include a plurality of instructions for causing an electronic device (which can be a personal computer, an electronic device, or a network device, etc.) or a processor to execute part of the method described in each embodiment of the present application.

[0150] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the described apparatus embodiments are merely schematic. For example, the division of the modules is merely a logical function division. There can be another division manner in actual implementation.

[0151] The modules illustrated as separate components can or can not be physically separated, and the components illustrated as modules can or can not be physical units. They can be located in one place, or distributed on a plurality of network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

Claims

1. An artificial intelligence-based memory saving and memory extraction method, characterized by, The method comprises: collecting video data and audio data of a scene where a user is located, processing the video data to obtain image information, including: collecting the video data by dynamic frame rate according to a scene change detection algorithm and a rate prediction algorithm, detecting the scene of the video data by the scene change detection algorithm to obtain a video scene type, predicting the adaptive transform rate of the video data corresponding to each video scene type by the rate prediction algorithm, collecting the video data by frame rate according to the predicted transform rate to obtain a plurality of image data corresponding to the video scene type, segmenting each image data to obtain an image data block, and identifying the image data block to obtain the image information; processing the audio data to obtain audio information, including: collecting the audio data by frame rate to obtain a plurality of sub-audio data, detecting whether the collection scene of the audio data changes according to the scene change detection algorithm, classifying the sub-audio data whose collection scene changes to obtain an audio scene type when the collection scene of the audio data changes, segmenting each sub-audio data to obtain layered audio, and identifying the layered audio to obtain the audio information; classifying and identifying the image information and the audio information to obtain a memory original text; calling a large language model to process the memory original text, and storing a memory summary obtained by processing and the memory original text in a database; when a query question of the user is received, querying and outputting a memory summary corresponding to the query question in the database.

2. The artificial intelligence-based memory saving and memory extraction method according to claim 1, wherein: when the predicted transform rate is higher than a preset rate threshold, a first preset frame rate is used to collect the video data by frame rate to obtain a plurality of image data corresponding to the video scene type; when the predicted transform rate is lower than the preset rate threshold, a second preset frame rate is used to collect the video data by frame rate to obtain a plurality of image data corresponding to the video scene type; wherein the first preset frame rate is greater than the second preset frame rate. 3.The artificial intelligence-based memory saving and memory extraction method according to claim 1, characterized in that, The classification and identification of the image information and the audio information to obtain a memory original text comprises: classifying and identifying the image information to obtain an image text, classifying and identifying the audio information to obtain an audio text, and associating the image text and the audio text to obtain the memory original text.

4. The artificial intelligence-based memory storing and memory retrieving method according to claim 3, wherein, The semantic association of the image text and the audio text comprises: associating the image text and the audio text based on a scene, a time, a place or a theme to structure and combine the image text and the audio text.

5. The artificial intelligence-based memory storing and memory retrieving method according to claim 3, wherein, The method further comprises: classifying and compressing the image information corresponding to the video scene type; and storing the audio scene type and the corresponding audio information.

6. The artificial intelligence-based memory saving and memory extraction method according to claim 5, characterized in that, When the query question is a voice query question input by the user in a voice form, the querying in the database and outputting the memory summary corresponding to the query question comprises: performing voice recognition on the voice query question to obtain a text query question; querying in the database and outputting the memory summary corresponding to the text query question.

7. An electronic device, comprising: A computer program product comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the artificial intelligence-based memory saving and memory extracting method according to any one of claims 1 to 6 when executing the computer program.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the artificial intelligence-based memory saving and memory extracting method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Artificial intelligence learning method based on speech recognition

    CN109410911A

  • Memory aid method using audio / video data

    US20150338912A1