Cigarette retail spatio-temporal data cloud edge collaborative determination system, method, equipment and medium

By designing a cloud-edge collaborative determination system for cigarette retail space-time data, using edge-end and cloud-side technologies, video and audio data are collected and analyzed, and cigarette category information and retail prices are extracted, the problem of difficulty in collecting data in the cigarette retail market is solved, and efficient and accurate spatio-time data acquisition is achieved.

CN120126048AActive Publication Date: 2025-06-10贵州省烟草公司安顺市公司
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510193734.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

It is difficult to collect data in the cigarette retail market, lack of time and space information, and it is difficult to adapt to the development reality of the cigarette industry in the new era.

Method used

Design a cloud-edge collaborative determination system for cigarette retail space-time data. Through the edge-end acquisition module, analysis module, speech recognition module and cloud-side analysis module, video and audio data are collected and analyzed, cigarette category information and retail prices are extracted, time-time information is integrated, and space-time information is written into the information management system.

Benefits of technology

It realizes rapid and accurate acquisition of time and space data on cigarette retail, and improves the efficiency of market analysis and resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126048A_ABST
    Figure CN120126048A_ABST
Patent Text Reader

Abstract

The invention discloses a cigarette retail spatio-temporal data cloud edge collaborative determination system, method and device and a medium, and relates to the technical field of artificial intelligence. The system comprises an edge end acquisition module, an analysis module, a voice recognition module, a price extraction module and a cloud analysis module. The acquisition module can acquire video data and audio data of a settlement area where cigarettes are located during retail, and the video data and the audio data correspond to each other through timestamps; the analysis module can analyze frame images in the video data to obtain category information of each cigarette; the voice recognition module can intercept voice segments from the audio data according to the category information and timestamps of the cigarettes, and performs text recognition on the voice segments to obtain text information of the voice segments; the price extraction module can extract retail prices of the cigarettes in the text information; and the cloud analysis module can perform fusion to obtain the cigarette retail price containing the space-time information. By adopting the system, the cigarette retail spatio-temporal data can be quickly obtained, and the accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly to a cloud-edge collaborative determination system, method, device and medium for cigarette retail spatio-temporal data. Background Art

[0002] It is difficult to obtain cigarette retail market data, which is a huge challenge faced by the high-quality development of the current tobacco industry. The cigarette retail guiding price has a planned nature, while the actual cigarette retail price contains market demand information. The low degree of coincidence between the two will make it difficult for the existing cigarette retail market analysis model to adapt to the actual development of the cigarette industry in the new era. The difficulties in current cigarette retail market data collection at least include the following aspects: First, the informatization level of retail terminals is not high, and the cost of collecting market information such as retail prices by traditional market survey methods is too high; Second, the lack of key market factors such as the profit status of cigarette retailers, cigarette brands, and cigarette price ranges limits the means of analyzing and regulating the market; Third, the existing data collection lacks time series and spatial elements, and cannot well reflect the time and space dynamic information of the market, etc., which is not conducive to more effective resource allocation. Summary of the Invention

[0003] In view of the above-mentioned defects or deficiencies in the related art, it is desirable to provide a cloud-edge collaborative determination system, method, device and medium for cigarette retail spatio-temporal data, which can quickly obtain cigarette retail prices containing spatio-temporal information to obtain accurate cigarette retail spatio-temporal data.

[0004] In a first aspect, the present application provides a cloud-edge collaborative determination system for cigarette retail spatio-temporal data, and the cloud-edge collaborative determination system for cigarette retail spatio-temporal data includes:

[0005] An edge-side acquisition module configured to acquire video data and audio data in the terminal settlement area where cigarettes are retailed, and the video data and the audio data are corresponding to spatial information through time stamps;

[0006] An edge-side analysis module electrically connected to the edge-side acquisition module, configured to analyze the frame images in the video data to obtain the category information of each cigarette and record spatio-temporal information; where the video data is each represents a video frame at time t i and the category data of the cigarette is obtained by frame-by-frame analysis where represents the cigarette category information corresponding to the video frame at time t i ;

[0007] An edge - end voice recognition module electrically connected to the edge - end analysis module, configured to adaptively intercept voice segments from the audio data according to the category information of the cigarette, the timestamp, and the spatial information; where the voice segment data is each represents an audio frame at time t i Align the video data of the edge - end acquisition module and the audio data of the voice recognition module through the timestamp. The above description can be expressed as:

[0008] Category data of cigarettes:

[0009] Voice segment data:

[0010] Timestamp: T = {t 1 , t 2 , t 3 , …, t n}

[0011] Fusion data: Here represents the combination of the category data of the i - th video frame and the voice segment data of the audio frame;

[0012] The edge - end voice recognition module compresses the bimodal data (W, A) containing the cigarette spatio - temporal information, which is the category data W and the voice segment data A, into (R), and then uploads R to the cloud deployed in the information center. At the cloud, decompress (R) → (W, A), and perform text recognition on the voice segment data A to obtain the text information of the voice segment data A, that is, (A) → (P), where (P) is the price text data P = {p 1 , p 2 , p 3 , …, p n}, and p i represents the i - th retail price in multiple groups of prices;

[0013] Finally, obtain the cigarette category C z = max{h(W)}, where h(W) refers to the frequency of the cigarette category in the category data, and max{h(W)} represents obtaining the cigarette category with the highest frequency;

[0014] An edge - end price extraction module electrically connected to the edge - end voice recognition module, configured to determine the final price P from the price text data P = {p i through the exponent P i_ind of p 1 , p 2 , p 3 , …, p n}z , where the calculation formula is:

[0015]

[0016] In the above formula, K f represents the price weight of the voice recognition retail price, p 0 represents the guiding price of the cigarette category C z obtained; represents the time weight of the voice recognition retail price, and Δt i represents the difference between the moment of audio extraction price and the moment of cigarette category recognition,

[0017] P z ={max(p i_ind )→p i )}

[0018] In the above formula, max(p i_ind )→p i represents the price corresponding to the maximum value;

[0019] The cloud analysis module is configured to fuse the category information of the cigarette with the text information of the voice segment to obtain a cigarette retail price including the spatio-temporal information, and write it into the information management system to obtain cigarette retail spatio-temporal data.

[0020] Optionally, in some embodiments of the present application, the edge analysis module includes a positioning unit and an output unit connected to each other;

[0021] The positioning unit is configured to sequentially locate the positions of the cigarettes in the frame image using the target detection neural network model according to the order of the timestamps, and generate the full-course movement trajectory of the cigarettes; and,

[0022] The output unit is configured to output the category information of the cigarette if the end point of the full-course movement trajectory of the cigarette is a customer.

[0023] Optionally, in some embodiments of the present application, the edge analysis module further includes a frame division unit and a preprocessing unit connected to each other;

[0024] The frame division unit is configured to perform frame division operations on the video data; the preprocessing unit is configured to extract the frame images from the frame division results according to a first preset time interval and perform noise reduction operations.

[0025] Optionally, in some embodiments of the present application, the edge voice recognition module includes a first determination unit and a first splicing unit connected to each other;

[0026] The first determination unit is configured to determine the first moment when the category information is recognized, and the time stamp includes the first moment; and,

[0027] The first splicing unit is configured to, based on the first moment, obtain voice data within a second preset time interval before the first moment, and splice the voice data with recording data within a third preset time interval after the first moment to obtain the voice segment.

[0028] Optionally, in some embodiments of the present application, the edge - side price extraction module includes a matching unit, and the matching unit is configured to traverse each character of the text information and sequentially match each character with a preset regular expression to obtain the retail price of the cigarette.

[0029] Optionally, in some embodiments of the present application, the edge - side price extraction module further includes a second determination unit and a correction unit that are connected to each other;

[0030] The second determination unit is configured to reversely determine the second moment corresponding to the retail price in the audio data, and the time stamp includes the second moment; and,

[0031] The correction unit is configured to correct the retail price according to the guiding price corresponding to the category information, the first moment when the category information is recognized, and the second moment to obtain the final retail price of the cigarette.

[0032] Optionally, in some embodiments of the present application, the edge - side acquisition module includes a detection unit, a recording unit, and a second splicing unit that are connected to each other;

[0033] The detection unit is configured to detect frame images in the video data to determine whether there is a cigarette sales behavior in the settlement area;

[0034] The recording unit is configured to, if there is a cigarette sales behavior in the settlement area, save historical recording data and restart recording for a preset duration to obtain current recording data; and,

[0035] The second splicing unit is configured to splice the historical recording data and the current recording data, and perform a decomposition and overwrite operation on the current recording data in the spliced recording data to obtain the audio data.

[0036] In a second aspect, the present application provides a method for cloud - edge collaborative determination of cigarette retail spatio - temporal data, and the method for cloud - edge collaborative determination of cigarette retail spatio - temporal data includes:

[0037] Collect video data and audio data in the terminal settlement area where cigarettes are retailed, and the video data and the audio data are corresponding to spatial information through timestamps;

[0038] Analyze the frame images in the video data to obtain the category information of each cigarette, and record the spatio-temporal information;

[0039] According to the category information of the cigarette and the timestamp and spatial information, adaptively intercept voice segments from the audio data, and perform text recognition on the voice segments to obtain the text information of the voice segments;

[0040] Extract the retail price of the cigarette in the text information;

[0041] Fuse the category information of the cigarette with the text information of the voice segment to obtain a cigarette retail price including the spatio-temporal information, and write it into the information management system to obtain cigarette retail spatio-temporal data.

[0042] In a third aspect, the present application provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the instruction, the program, the code set or the instruction set is loaded and executed by the processor to implement the steps of the cigarette retail spatio-temporal data cloud-edge collaboration method described in the second aspect.

[0043] In a fourth aspect, the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the cigarette retail spatio-temporal data cloud-edge collaboration method described in the second aspect.

[0044] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:

[0045] The embodiments of the present application provide a cigarette retail spatio-temporal data cloud-edge collaboration determination system, method, device and medium. When retailing cigarettes, the original video data and audio data in the terminal settlement area are directly collected. That is to say, the data is real-time, true and reliable. Then, the category information of each cigarette is obtained by analyzing the frame images in the video data. Based on the category information of the cigarette and the timestamp, voice segments are intercepted from the audio data and recognized to obtain text information, and the retail price of the cigarette in the text information is extracted. Furthermore, the category information of the cigarette is fused with the text information of the voice segment to obtain a cigarette retail price including the spatio-temporal information, and it is written into the information management system, thereby being able to quickly obtain cigarette retail spatio-temporal data and improving the accuracy. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0047] Figure 1 It is a structural block diagram of a cloud-edge collaborative determination system for cigarette retail spatio-temporal data provided by an embodiment of the present application;

[0048] Figure 2 It is another structural block diagram of a cloud-edge collaborative determination system for cigarette retail spatio-temporal data provided by an embodiment of the present application;

[0049] Figure 3 It is a schematic diagram of the process of processing recording data provided by an embodiment of the present application;

[0050] Figure 4 It is another schematic diagram of the process of processing recording data provided by an embodiment of the present application;

[0051] Figure 5 It is still another structural block diagram of a cloud-edge collaborative determination system for cigarette retail spatio-temporal data provided by an embodiment of the present application;

[0052] Figure 6 It is a schematic diagram of a scenario where a customer picks up cigarettes first and then pays for them in an embodiment of the present application, where Figure 6 (a) represents the customer picking up cigarettes, Figure 6 (b) represents the customer paying;

[0053] Figure 7 It is a schematic diagram of a scenario where a customer pays first and then picks up cigarettes in an embodiment of the present application, where Figure 7 (a) represents the customer paying, Figure 7 (b) represents the customer picking up cigarettes;

[0054] Figure 8 It is yet another structural block diagram of a cloud-edge collaborative determination system for cigarette retail spatio-temporal data provided by an embodiment of the present application;

[0055] Figure 9 It is a structural block diagram of a cloud-edge collaborative determination system for cigarette retail spatio-temporal data provided by another embodiment of the present application;

[0056] Figure 10 It is still another structural block diagram of a cloud-edge collaborative determination system for cigarette retail spatio-temporal data provided by another embodiment of the present application;

[0057] Figure 11 It is a schematic diagram of a scenario where two or more customers buy cigarettes provided by an embodiment of the present application;

[0058] Figure 12 A schematic diagram of a voice recognition process provided by an embodiment of the present application;

[0059] Figure 13 A schematic diagram of the basic process of a method for jointly determining cloud-edge collaborative spatio-temporal data of cigarette retail provided by an embodiment of the present application;

[0060] Figure 14 A block diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0062] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0063] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The following will Figures 1 to 14 elaborate in detail on the cloud-edge collaborative determination system, method, device, and medium for cigarette retail spatio-temporal data provided by the embodiments of the present application.

[0064] Please refer to Figure 1, which is a structural block diagram of a cloud-edge collaborative determination system for cigarette retail spatio-temporal data provided by an embodiment of the present application. The cloud-edge collaborative determination system 100 for cigarette retail spatio-temporal data includes an edge-side acquisition module 101, an edge-side analysis module 102 electrically connected to the edge-side acquisition module 101, an edge-side speech recognition module 103 electrically connected to the edge-side analysis module 102, an edge-side price extraction module 104 electrically connected to the edge-side speech recognition module 103, and a cloud-side analysis module 105. The electrical connection method can include any one of wired physical connection and wireless communication connection. Among them, the edge-side acquisition module 101 can acquire video data and audio data in the terminal settlement area during cigarette retail, and the video data and the audio data are corresponding to spatial information through timestamps; the edge-side analysis module 102 can analyze the frame images in the video data, obtain the category information of each cigarette, and record spatio-temporal information; the edge-side speech recognition module 103 can adaptively intercept voice segments from the audio data according to the category information of the cigarette and the timestamp and spatial information, and perform text recognition on the voice segments to obtain the text information of the voice segments; the edge-side price extraction module 104 can extract the retail price of the cigarette in the text information; and the cloud-side analysis module 105 can fuse the category information of the cigarette with the text information of the voice segments to obtain a cigarette retail price including spatio-temporal information, and write it into the information management system to obtain cigarette retail spatio-temporal data.

[0065] Exemplarily, the following combines Figures 2 to 12 as shown to detail some component modules of the cloud-edge collaborative determination system 100 for cigarette retail spatio-temporal data in the embodiments of the present application. For example Figure 2 as shown, the edge-side acquisition module 101 includes, but is not limited to, a detection unit 1011, a recording unit 1012, and a second splicing unit 1013 connected to each other, etc. Among them, the detection unit 1011 can continuously detect the frame images in the video data to determine whether there is a cigarette sales behavior in the settlement area. For example, the video data is collected by a camera, and when the frame image shows people walking outside the counter, it can be determined that there is a cigarette sales behavior in the settlement area. If there is a cigarette sales behavior in the settlement area, the recording unit 1012 can save the historical recording data and start a recording with a preset duration again to obtain the current recording data. For example, the recording data is collected by a pick-up microphone, the historical recording data is saved as wav_1, the preset duration is 1 minute and 30 seconds, the current recording data is wav_2. For another example, when a new cigarette sales behavior appears in the settlement area during the recording of wav_2, stop the recording of wav_2 and save the data, and at the same time start a 1 minute and 30 seconds recording again, named wav_3. For another example, when a new cigarette sales behavior appears in the settlement area during the recording of wav_3, stop the recording of wav_3 and save the data, and at the same time start a 1 minute and 30 seconds recording again, named wav_4, and so on, until the complete wav_n is recorded, that is, the recording ends.

[0066] Further, as Figure 3 shown, the second splicing unit 1013 can splice historical recording data and current recording data, and perform decomposition and overwrite operations on the current recording data in the spliced recording data to obtain audio data. For example, Figure 4 shown, divide 1 minute and 30 seconds into 9 time periods such as t1, t2,..., t9, each period is 10 seconds long, and the corresponding current recording data is decomposed into segments f 1 , segment f 2 ,..., segment f 9 , and then obtain the 10th recording segment as segment f 10 , that is, the duration of the 10th recording segment is also 10 seconds, and then perform segment f 2 overwrite segment f 1 , segment f 3 overwrite segment f 2 ,..., segment f 9 overwrite segment f 8 and segment f 10 overwrite segment f 9 etc. The advantage of such a setting is that it can reduce the calculation amount, only retain the recordings before and after the transaction time node, and at the same time protect the privacy of the merchant. In addition, after recording the complete wav_n data, the embodiment of the present application can extract the last 20 seconds of the current recording data as the memory voice, then record 1 minute and 10 seconds starting from this memory voice, and continue to perform the decomposition and overwrite operations, which will not be elaborated here.

[0067] Again, as Figure 5 shown, the edge - side analysis module 102 includes, but is not limited to, a positioning unit 1021 and an output unit 1022 connected to each other, etc. Among them, the positioning unit 1021 can, in the order of time stamps, use the target - detection neural network model to locate the positions of each cigarette in the frame image in turn, and generate the full - course movement trajectory of the cigarette. For example, the target - detection neural network model can be the YOLO model, and this YOLO model can locate and frame the positions of each cigarette in the frame image, and give the category information of the cigarette around the frame. Again, for example, the full - course movement trajectory of the cigarette can be that the merchant places and retrieves the cigarette, or it can be that the customer picks up and puts back the cigarette, etc. And if the end point of the full - course movement trajectory of the cigarette is the customer, it means that the cigarette is in the settlement state. At this time, the output unit 1022 can output the category information of the cigarette. For example, Figure 6 shown is a schematic diagram of the scenario where the customer picks up the cigarette first and then pays. Among them, Figure 6 (a) represents the customer picking up the cigarette, Figure 6 (b) represents the customer paying, and 1 represents the camera, 2 represents the server, 3 represents the pickup, 4 represents the cigarette, 5 represents the table, and 6 represents the payment device. Again, for exampleFigure 7 The figure shows a schematic diagram of the scenario where the customer pays first and then picks up cigarettes. Among them Figure 7 (a) represents the customer making a payment, Figure 7 (b) represents the customer picking up cigarettes.

[0068] Furthermore, as Figure 8 shown, in the embodiment of the present application, the edge - end analysis module 102 may further include a frame - splitting unit 1023 and a pre - processing unit 1024 that are connected to each other. Among them, the frame - splitting unit 1023 can perform frame - splitting operations on video data, and the pre - processing unit 1024 can extract frame images from the frame - splitting result and perform noise reduction operations according to a first preset time interval. For example, the first preset time interval is 2 seconds, and the noise reduction method includes but is not limited to the median filtering method, thereby improving the detection accuracy of the subsequent target - detection neural network model.

[0069] Also, as Figure 9 shown, the edge - end speech recognition module 103 includes but is not limited to a first determination unit 1031 and a first splicing unit 1032 that are connected to each other, etc. Among them, the first determination unit 1031 can determine the first moment when the category information is recognized, and the timestamp includes the first moment. Thus, the first splicing unit 1032 can, based on the first moment, obtain the speech data within a second preset time interval before the first moment, and splice the speech data with the recording data within a third preset time interval after the first moment to obtain a speech segment. The second preset time interval and the third preset time interval can be equal or unequal. Further, the edge - end speech recognition module 103 may further include a text conversion unit 1033. The text conversion unit 1033 can use a speech recognition model to perform text recognition on the speech segment to obtain the normalized text information of the speech segment. For example, the speech recognition model is a fine - tuned whisper model, and the sampling frequency is 16000Hz. Optionally, before text conversion, the text conversion unit 1033 of the embodiment of the present application can also perform noise reduction operations on the speech segment, thereby reducing noise interference and improving the recognition accuracy of the speech recognition model.

[0070] Another example is Figure 10 shown, the edge - end price extraction module 104 may include a matching unit 1041. The matching unit 1041 can traverse each character of the text information and match each character with a preset regular expression in turn to obtain the retail price of cigarettes. For example, the preset regular expression can be "Alipay received *** yuan". In addition, after obtaining the retail price of cigarettes in the embodiment of the present application, the retail quantity of cigarettes can be obtained by counting the number of retail prices. Further, the edge - end price extraction module 104 may further include a second determination unit 1042 and a correction unit 1043 that are connected to each other. As Figure 11As shown, when more than two customers (such as customer A and customer B) purchase cigarettes within a certain period of time, the matching unit 1041 will extract multiple groups of prices. Since the voice segment is associated with the timestamp, at this time, as Figure 12 shown, the second determination unit 1042 can reversely determine the second moment corresponding to the retail price in the audio data, and the timestamp includes the second moment. Furthermore, the calibration unit 1043 can calibrate the retail price according to the guiding price corresponding to the category information, the first moment for identifying the category information, and the second moment, and obtain the final retail price of the cigarette. For example, the final retail price of the cigarette can be calculated by the following formula, that is

[0071]

[0072] In formula (1), K f represents the price weight of the voice recognition retail price, p 0 represents the guiding price of the cigarette category C z obtained; represents the time weight of the voice recognition retail price, and Δt i represents the difference between the moment of extracting the price from the audio and the moment of identifying the cigarette category,

[0073] P z ={max(p i_ind )→p i} (2)

[0074] In formula (2), max(p i_ind )→p i represents the price corresponding to taking the maximum value.

[0075] The cigarette retail spatio-temporal data cloud-edge collaborative determination system provided by the embodiment of the present application directly collects the original video data and audio data in the terminal settlement area when retailing cigarettes. That is to say, the data is real-time, true and reliable. Then, by analyzing the frame images in the video data, the category information of each cigarette is obtained. Based on the category information of the cigarette and the timestamp, the voice segment is intercepted from the audio data and the text information is recognized, and the retail price of the cigarette in the text information is extracted. Furthermore, the category information of the cigarette is fused with the text information of the voice segment to obtain a cigarette retail price containing spatio-temporal information and written into the information management system. Thus, the cigarette retail spatio-temporal data can be obtained quickly, improving the accuracy.

[0076] Based on the foregoing embodiments, the embodiment of the present application provides a method for cloud-edge collaborative determination of cigarette retail spatio-temporal data. Please refer to Figure 13 , which is a schematic diagram of the basic process of a method for cloud-edge collaborative determination of cigarette retail spatio-temporal data provided by the embodiment of the present application. The method specifically includes the following steps:

[0077] S101, Collect video data and audio data in the terminal settlement area during cigarette retail. The video data and audio data are corresponding to each other through timestamps and spatial information.

[0078] Exemplarily, in the embodiments of the present application, the video data can be collected by a camera, and the audio data can be collected by a pick-up. Further, when collecting the video data and audio data, the embodiments of the present application can continuously detect the frame images in the video data to determine whether there is a cigarette sales behavior in the settlement area; then, if there is a cigarette sales behavior in the settlement area, save the historical recording data, and restart the recording for a preset duration to obtain the current recording data; furthermore, splice the historical recording data and the current recording data, and perform decomposition and covering operations on the current recording data in the spliced recording data to obtain the audio data.

[0079] S102, Analyze the frame images in the video data to obtain the category information of each cigarette and record the spatio-temporal information.

[0080] Exemplarily, in the embodiments of the present application, in the order of the timestamps, the target detection neural network model can be used to locate the positions of each cigarette in the frame images in turn and generate the whole-course movement trajectory of the cigarette. If the end point of the whole-course movement trajectory of the cigarette is the customer, output the category information of the cigarette. Optionally, before analyzing the frame images in the video data, the embodiments of the present application can also perform frame splitting on the video data, and extract frame images from the frame splitting result at a first preset time interval and perform noise reduction operations.

[0081] S103, Adaptively intercept voice segments from the audio data according to the category information of the cigarette and the timestamps and spatial information, and perform text recognition on the voice segments to obtain the text information of the voice segments.

[0082] Exemplarily, the embodiments of the present application first determine the first moment when the category information is recognized, and the timestamp includes the first moment. Then, based on the first moment, obtain the voice data within a second preset time interval before the first moment, and splice the voice data with the recording data within a third preset time interval after the first moment to obtain the voice segment. Further, the embodiments of the present application can use a speech recognition model to perform text recognition on the voice segment to obtain the normalized text information of the voice segment. For example, the speech recognition model is a fine-tuned whisper model with a sampling frequency of 16000Hz. Optionally, before text conversion, the embodiments of the present application can also perform noise reduction operations on the voice segment to reduce noise interference and improve the model recognition accuracy.

[0083] S104, Extract the retail price of the cigarette in the text information.

[0084] Exemplarily, embodiments of the present application can traverse each character of the text information, and sequentially match each character with a preset regular expression to obtain the retail price of cigarettes. Further, if more than two retail prices of cigarettes are extracted, embodiments of the present application can also reversely determine the second moment corresponding to the retail price in the audio data, the timestamp includes the second moment, and correct the retail price according to the guiding price corresponding to the category information, the first moment when the category information is recognized, and the second moment to obtain the final retail price of cigarettes.

[0085] S105, fuse the category information of the cigarettes with the text information of the voice segment to obtain a cigarette retail price including spatio-temporal information, and write it into the information management system to obtain cigarette retail spatio-temporal data.

[0086] It should be noted that the descriptions of the same steps and the same content in this embodiment and other embodiments can be referred to the descriptions in other embodiments, and will not be repeated here.

[0087] The cigarette retail spatio-temporal data cloud-edge collaborative determination method provided by embodiments of the present application directly collects the original video data and audio data in the terminal settlement area when retailing cigarettes. That is to say, the data is real-time, true and reliable. Then, by analyzing the frame images in the video data, the category information of each cigarette is obtained. Based on the category information of the cigarette and the timestamp, a voice segment is intercepted from the audio data and the text information is recognized, and the retail price of the cigarette in the text information is extracted. Furthermore, the category information of the cigarette is fused with the text information of the voice segment to obtain a cigarette retail price including spatio-temporal information, and it is written into the information management system. Thus, cigarette retail spatio-temporal data can be quickly obtained, improving the accuracy.

[0088] Based on the foregoing embodiments, embodiments of the present application provide an electronic device. Please refer to Figure 14 , the electronic device 200 may include a processor 201 and a memory 202. At least one instruction, at least one program, a code set or an instruction set is stored in the memory 202, and the instruction, program, code set or instruction set is loaded and executed by the processor 201 to implement Figure 13 the steps of the cigarette retail spatio-temporal data cloud-edge collaborative determination method corresponding to the embodiment.

[0089] On the other hand, embodiments of the present application provide a computer-readable storage medium for storing program codes, and the program codes are used to execute any one of the implementation manners in the cigarette retail spatio-temporal data cloud-edge collaborative determination method corresponding to the foregoing Figure 13 embodiments.

[0090] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0091] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0092] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more units can be integrated in one module. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0093] Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method for determining the cloud-edge collaboration of cigarette retail spatio-temporal data in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0094] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0095] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A cloud-edge collaborative determination system for cigarette retail spatiotemporal data, characterized in that: The cigarette retail spatiotemporal data cloud-edge collaborative determination system includes: The edge acquisition module is configured to collect video data and audio data of the terminal settlement area where cigarettes are retailed, wherein the video data and the audio data correspond to the spatial information through a timestamp; An edge end analysis module electrically connected to the edge end acquisition module, configured to analyze frame images in the video data, obtain category information of each of the cigarettes, and record spatiotemporal information; an edge-end speech recognition module electrically connected to the edge-end analysis module, configured to adaptively intercept speech segments from the audio data according to the category information of the cigarette and the timestamp and spatial information, and perform text recognition on the speech segments to obtain text information of the speech segments; An edge-end price extraction module electrically connected to the edge-end speech recognition module, configured to extract the retail price of the cigarette in the text information; The cloud analysis module is configured to merge the category information of the cigarette with the text information of the voice segment to obtain a cigarette retail price containing the spatiotemporal information, and write it into the information management system to obtain cigarette retail spatiotemporal data.

2. The cloud-edge collaborative determination system for cigarette retail spatiotemporal data according to claim 1 is characterized in that: The edge end analysis module includes a positioning unit and an output unit connected to each other; The positioning unit is configured to locate the position of each cigarette in the frame image in sequence according to the sequence of the timestamps using a target detection neural network model, and generate a full movement trajectory of the cigarette; and, The output unit is configured to output the category information of the cigarette if the end point of the entire moving trajectory of the cigarette is a customer.

3. The cloud-edge collaborative determination system for cigarette retail spatiotemporal data according to claim 2 is characterized in that: The edge end analysis module also includes a framing unit and a pre-processing unit connected to each other; The framing unit is configured to perform a framing operation on the video data; the preprocessing unit is configured to extract the frame image from the framing result and perform a noise reduction operation according to a first preset time interval.

4. The cloud-edge collaborative determination system for cigarette retail spatiotemporal data according to claim 1, characterized in that: The edge end speech recognition module includes a first determination unit and a first splicing unit connected to each other; The first determining unit is configured to determine a first moment at which the category information is identified, the timestamp including the first moment; and, The first splicing unit is configured to obtain voice data within a second preset time interval before the first moment based on the first moment, and splice the voice data with recording data within a third preset time interval after the first moment to obtain the voice segment.

5. The cloud-edge collaborative determination system for cigarette retail spatiotemporal data according to claim 1, characterized in that: The edge price extraction module includes a matching unit, which is configured to traverse each character of the text information and match each character with a preset regular expression in sequence to obtain the retail price of the cigarette.

6. The cloud-edge collaborative determination system for cigarette retail spatiotemporal data according to claim 5, characterized in that: The edge price extraction module further includes a second determination unit and a correction unit connected to each other; The second determination unit is configured to reversely determine a second time instant corresponding to the retail price in the audio data, the timestamp including the second time instant; and, The correction unit is configured to correct the retail price according to the guide price corresponding to the category information, the first moment when the category information is identified, and the second moment to obtain the final retail price of the cigarette.

7. The cigarette retail spatiotemporal data cloud-edge collaborative determination system according to any one of claims 1 to 6, characterized in that: The edge acquisition module includes a detection unit, a recording unit and a second splicing unit that are connected to each other; The detection unit is configured to detect the frame image in the video data to determine whether there is cigarette sales behavior in the settlement area; The recording unit is configured to save historical recording data and restart recording for a preset time to obtain current recording data if there is cigarette sales in the settlement area; as well as, The second splicing unit is configured to splice the historical recording data with the current recording data, and to decompose and overwrite the current recording data in the spliced ​​recording data to obtain the audio data.

8. A cloud-edge collaborative determination method for cigarette retail spatiotemporal data, characterized in that: The cloud-edge collaborative determination method for cigarette retail spatiotemporal data includes: Collecting video data and audio data of the terminal settlement area where cigarettes are sold at retail, wherein the video data and the audio data correspond to the spatial information through a timestamp; Analyze the frame images in the video data to obtain the category information of each cigarette and record the time and space information; Adaptively extracting a voice segment from the audio data according to the cigarette category information and the timestamp and space information, and performing text recognition on the voice segment to obtain text information of the voice segment; Extracting the retail price of the cigarettes in the text information; The category information of the cigarettes is integrated with the text information of the voice segment to obtain a cigarette retail price containing the spatiotemporal information, and the price is written into the information management system to obtain cigarette retail spatiotemporal data.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the instruction, the program, the code set or the instruction set is loaded and executed by the processor to implement the steps of the cloud-edge collaborative determination method for cigarette retail spatiotemporal data as described in claim 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the cloud-edge collaborative determination method for cigarette retail spatiotemporal data as described in claim 8.

Citation Information

Patent Citations

  • Method for acquisition and distribution of product price information

    CN103839166A

  • Settlement method and system based on visual identification

    CN111222388A

  • Image recognition device and method, equipment and storage medium

    CN115311791A

  • Task processing method, commodity classification method and commodity classification method for e-commerce live broadcast

    CN118097490A

  • Multi-algorithm fusion risk prediction method for tobacco monopoly retailers

    CN118396393A