Skip to content

qubic.run

AbstractCircuitRunner implementation which runs locally on the ZCU216. Used by soc_rpc_server to load and run circuits. Can also be used locally, if user environment is on the ZCU216.

CircuitRunner

Bases: AbstractCircuitRunner

Class for taking a program in binary/ASM form and running it on the FPGA. Currently, this class is meant to be run on the QubiC FPGA PS + pynq system. It will load and configure the specified PL bitfile, and can then be used to configure PL memory and registers, and read back data from experiments.

Attributes:

Name Type Description
_pl_driver PLInterface

used for low level access to memory and registers

loaded_channels list

channels with a program currently loaded

Source code in qubic/run.py
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
class CircuitRunner(AbstractCircuitRunner):
    """
    Class for taking a program in binary/ASM form and running it on 
    the FPGA. Currently, this class is meant to be run on the QubiC FPGA 
    PS + pynq system. It will load and configure the specified PL bitfile,
    and can then be used to configure PL memory and registers, and read 
    back data from experiments.

    Attributes
    ----------
    _pl_driver: pl.PLInterface 
        used for low level access to memory and registers
    loaded_channels: list 
        channels with a program currently loaded
    """

    def __init__(self, pl_driver):
        self._pl_driver = pl_driver

        self.result_channels = []
        self._cur_nshots = None


    def load_and_run(self, executable: Executable, n_total_shots: int): 
        """
        Load circuit described by rawasm "binary", then run for n_total_shots. 

        Parameters
        ----------
        executable: Executable
        n_total_shots: int
            number of shots to run. Program is restarted from the beginning 
            for each new shot

        Returns
        -------
        dict:
            Complex IQ shots for each accbuf in chanlist; each array has 
            shape `(n_total_shots, reads_per_shot)`
        """
        self.load_executable(Executable)
        return self.run_circuit(n_total_shots)

    def load_executable(self, executable: Executable | dict, load_commands: bool = True, 
                        load_freqs: bool = True, load_envs: bool = True, zero: bool = True):
        """
        Load the circuit described by `executable`, which is the output of 
        the final distributed proc assembler stage. Loads command memory, env memory
        and freq buffer memory, according to specified input parameters. Before circuit is loaded, 
        if zero=True, all channels are zeroed out using zero_command_buf()

        Parameters
        ----------
        executable: Executable | dict
            Compiled program binary. Can be an Executable object or its dictionary
            representation (for xmlrpc support).
        zero: bool
            if True, (default), zero out all cmd buffers before loading circuit
        load_commands: bool
            if True, (default), load command buffers
        load_freqs: bool
            if True, (default), load freq buffers
        load_envs: bool
            if True, (default), load env buffers
        """
        if isinstance(executable, dict):
            executable = executable_from_dict(executable)

        if zero:
            self.zero_command_buf()

        self._cur_nshots = None
        for memname, data in executable.get_binaries_fromboard().items():
            if load_commands and 'command' in memname:
                self._pl_driver.write_mem_buf(memname, data)
            elif load_envs and 'env' in memname:
                self._pl_driver.write_mem_buf(memname, data)
            elif load_freqs and 'freq' in memname:
                self._pl_driver.write_mem_buf(memname, data)

        for regname, value in executable.get_registers_fromboard().items():
            self._pl_driver.write_reg(regname, value)

        self.loaded_result_channels = executable.result_channels

    def start_program(self, nshots, clock_count=None):
        self._cur_nshots = nshots
        self._pl_driver.start_program(nshots, clock_count)

    def ddr_start_circuit(self, raw_asm, n_total_shots, circuit_index,
                          reload_cmd=True, reload_freq=True, reload_env=True,
                          zero_between_reload=True):
        """Load and start a single circuit for DDR readout.

        Called by the client when it connects directly to the DMA server,
        bypassing XML-RPC for bulk data transfer.
        """
        t0 = time.time()
        if circuit_index == 0:
            self._pl_driver.set_default_regs()
            self.load_executable(raw_asm, True, True, True, True)
        else:
            self.load_executable(raw_asm, zero=zero_between_reload,
                                 load_commands=reload_cmd, load_freqs=reload_freq,
                                 load_envs=reload_env)
        t_load = time.time()
        # Clear stale SHOT_DONE / CUR_ADDR / FINAL_ADDR from previous circuits
        # or BRAM mode before starting this circuit. SHOT_DONE is clear-on-read.
        # See Desktop/ddr_producer_stale_state_bug.md.
        axil = self._pl_driver.overlay.axi_lite_ctrl_0.mmio
        _ = axil.read(0x1C)            # SHOT_DONE  (clear-on-read)
        axil.write(0x18, 0)            # FINAL_ADDR
        axil.write(0x20, 0)            # CUR_ADDR
        t_clear = time.time()
        self.start_program(n_total_shots)
        t_start = time.time()
        print(f'[DDR timing server] circuit {circuit_index}: '
              f'load={t_load-t0:.3f}s  clear={t_clear-t_load:.3f}s  '
              f'start={t_start-t_clear:.3f}s  total={t_start-t0:.3f}s')
        return True

    def ddr_batch_prepare(self, executable, load_freqs=True, load_envs=True):
        """Prepare the board for a DDR command-streaming batch (633e80df lineage).

        Python keeps its side of the split: dsp resets + env/freq/register BRAM
        loads. Commands do NOT go through here — the BRAM command path is
        physically removed in this gateware; the client streams K x 256 KiB
        command images to batch_server (port 8081), which feeds DDR bank c1
        via MM2S and issues global_start.

        A hardware-sequenced batch shares the BRAM env/freq tables across its
        circuits, so the client verifies uniformity and passes ONE representative
        executable per hardware batch. When per-circuit tables differ and the
        user allowed reload_freq/reload_env, the client instead calls this once
        per circuit (software-sequenced; circuits are then NOT continuous).
        Board-verified prep order from ddr_verify.c: resets first, then tables.
        """
        if isinstance(executable, dict):
            executable = executable_from_dict(executable)
        self._pl_driver.set_default_regs()
        for reg in ('dspreset', 'resetacc', 'reset_bram_read'):
            self._pl_driver.write_reg(reg, 1)
            self._pl_driver.write_reg(reg, 0)
        self.load_executable(executable, load_commands=False, load_freqs=load_freqs,
                             load_envs=load_envs, zero=False)
        return True

    def wait_and_readback(self, reads_per_shot=None, from_server=False) -> Dict[str, np.ndarray]:
        """
        Returns
        -------
        Dict[str, bytes]
            results for each channel
        """
        if isinstance(reads_per_shot, int):
            for _, chan in self.loaded_result_channels.items():
                chan.reads_per_shot = reads_per_shot
        elif isinstance(reads_per_shot, dict):
            for channame, n_reads in reads_per_shot.items():
                if channame in self.loaded_result_channels:
                    self.loaded_result_channels[channame].reads_per_shot = n_reads
        else:
            if reads_per_shot is not None:
                raise Exception(f'reads per shot: {reads_per_shot} invalid type')

        # s11 = {channame: np.nan*np.zeros((1, self._cur_nshots, chan.reads_per_shot), dtype=np.complex128) 
        #        for channame, chan in self.loaded_result_channels.items()}
        return self._pl_driver.wait_and_readback(self.loaded_result_channels, self._cur_nshots)

    def zero_command_buf(self, core_inds: List[str | int] = None):
        """
        Loads command memory with dummy asm program: reset phase, 
        output done signal, then idle. This is useful/necessary if 
        a new program is loaded on a subset of cores such that the 
        previous program is not completely overwritten (e.g. you 
        are loading a program that runs only on core 2, and the 
        previous program used cores 2 and 3).

        Parameters
        ----------
        core_inds: list
            list of channels (proc cores) to load. Defaults to
            all channels in currently loaded gateware.
        """
        zero_prog = Executable({':' + mem: bytes(16) for mem in self._pl_driver.get_program_memories(core_inds)}) 
        mems = zero_prog.get_binaries_fromboard()
        for memname, data in mems.items():
            self._pl_driver.write_mem_buf(memname, data)

    def run_circuit_batch(self, 
                          executables: List[Executable | Dict], 
                          n_total_shots: int, 
                          reads_per_shot: int = None,
                          timeout_per_shot: float = 20,
                          reload_cmd: bool = True,
                          reload_freq: bool = True,
                          reload_env: bool = True,
                          zero_between_reload: bool = True,
                          ddr: bool = False,
                          from_server: bool = False):
        """
        Runs a batch of circuits given by a list of compiled executables. Each circuit is run n_total_shots
        times. `reads_per_shot` and `n_total_shots` are passed directly into `run_circuit`, and must
        be the same for all circuits in the batch. The parameters `reload_cmd`, `reload_freq`, `reload_env`, and 
        `zero_between_reload` control which of these fields is rewritten circuit-to-circuit (everything is 
        rewritten initially). Leave these all at `True` (default) for maximum safety, to ensure that QubiC 
        is in a clean state before each run. Depending on the circuits, some of these can be turned off 
        to save time.

        TODO: consider throwing some version of all the args here into a BatchedCircuitRun or somesuch
        object

        Parameters
        ----------
        executables: List[Executable, Dict]
            List of compiled program binaries. Each element can be an Executable object or its dictionary
            representation (for xmlrpc support).
        n_total_shots: int
            number of shots per circuit
        reads_per_shot: int | dict
            number of values per shot per channel to read back from accbuf. If dict, indexed
            by channel name (e.g. `Q0.rdlo`). If int, assumed to be the same across channels. 
            Unless multiple circuits were rastered pre-compilation or there is mid-circuit 
            measurement involved this is typically 1
        timeout_per_shot: float
            job will time out if time to take a single shot exceeds this value in seconds 
            (this likely means the job is hanging due to timing issues in the program or gateware)
        reload_cmd: bool
            if True, reload command buffer between circuits
        reload_freq: bool
            if True, reload freq buffer between circuits
        reload_env: bool
            if True, reload env buffer between circuits
        from_server: bool
            set to true if calling over RPC. If True, pack returned s11 arrays into
            byte objects
        ddr: bool
            if True, use DDR memory readout instead of BRAM. Data is fetched
            automatically via socket after each circuit completes. Default is False.
        Returns
        -------
        list:
            Complex IQ shots for each accbuf in chanlist; each array has
            shape `(len(executables), n_total_shots, reads_per_shot)`
        """
        results = []

        # Defer all timing prints until AFTER the loop so the print() latency
        # never pollutes the next circuit's load measurement. Also use 6
        # decimal places (microsecond precision).
        timing_log = []

        if ddr:
            # DDR mode: open one persistent RING session for all circuits.
            # The DMA server waits for SHOT_DONE after each "NEXT".
            t_batch_start = time.time()
            with _DdrRingSession(len(executables)) as ddr_session:
                for i, raw_asm in enumerate(tqdm(executables)):
                    logging.getLogger(__name__).info(f'starting circuit {i}/{len(executables)-1}')

                    t0 = time.time()
                    if i == 0:
                        self._pl_driver.set_default_regs()
                        self.load_executable(raw_asm, True, True, True, True)
                    else:
                        self.load_executable(raw_asm, zero=zero_between_reload, load_commands=reload_cmd,
                                             load_freqs=reload_freq, load_envs=reload_env)
                    t_load = time.time()

                    self.start_program(n_total_shots)
                    t_start = time.time()

                    raw_bytes = ddr_session.read_next()
                    t_read = time.time()

                    results.append(raw_bytes)
                    timing_log.append(
                        f'[DDR timing server] circuit {i}: '
                        f'load={t_load-t0:.6f}s  start={t_start-t_load:.6f}s  '
                        f'dma_read={t_read-t_start:.6f}s  data={len(raw_bytes)} bytes')

            t_batch_end = time.time()
            timing_log.append(
                f'[DDR timing server] total: {t_batch_end-t_batch_start:.6f}s '
                f'for {len(executables)} circuits')
        else:
            for i, raw_asm in enumerate(tqdm(executables)):
                logging.getLogger(__name__).info(f'starting circuit {i}/{len(executables)-1}')
                t0 = time.time()
                if i == 0:
                    self._pl_driver.set_default_regs()
                    self.load_executable(raw_asm, True, True, True, True)
                else:
                    self.load_executable(raw_asm, zero=zero_between_reload, load_commands=reload_cmd,
                                         load_freqs=reload_freq, load_envs=reload_env)
                t_load = time.time()

                results_i = self.run_circuit(n_total_shots, reads_per_shot, timeout_per_shot, from_server)
                t_run = time.time()
                results.append(results_i)
                timing_log.append(
                    f'[BRAM timing server] circuit {i}: '
                    f'load={t_load-t0:.6f}s  run={t_run-t_load:.6f}s  '
                    f'total={t_run-t0:.6f}s')

        # Emit deferred timing log AFTER all measurements are complete.
        for line in timing_log:
            print(line)

        logging.getLogger(__name__).info('batch finished')
        return results

    def load_and_run_acq(self, 
                         raw_asm_prog: Executable, 
                         n_total_shots: int = 1, 
                         nsamples: int = 8192, 
                         acq_chans: Dict[str, int] = {'0':0,'1':1}, 
                         trig_delay: float = 0, 
                         decimator: int = 0, 
                         return_acc: bool = False, 
                         from_server: bool = False):
        """
        Load the program given by `raw_asm_prog` and acquire raw (or downconverted) adc traces.

        Parameters
        ----------
        raw_asm_prog: Executable | Dict
            Compiled program binary to run. See `load_executable` for details.
        n_total_shots: int
            number of shots to run. Program is restarted from the beginning 
            for each new shot
        nsamples: int
            number of samples to read from the acq buffer
        acq_chans: dict
            current channel mapping is:

                '0': ADC_237_2 (main readout ADC)
                '1': ADC_237_0 (other ADC connected in gateware)
                TODO: figure out DLO channels, etc and what they mean
        trig_delay: float
            time to delay acquisition, relative to circuit start.
            NOTE: this value, when converted to units of clock cycles, is a 
            16-bit value. So, it maxes out at CLK_PERIOD*(2**16) = 131.072e-6
        decimator: int
            decimation interval when sampling. e.g. 0 means full sample rate, 1
            means capture every other sample, 2 means capture every third sample, etc
        return_acc: bool
            if True, return a single acc (integrated + accumulated readout) value per shot,
            on each loaded channel. Default is False.
        from_server: bool
            set to true if calling over RPC. If True, pack returned acq arrays into
            byte objects

        Returns
        -------
        tuple | Dict
            - if `return_acc` is `False`:

                - dict:
                    array of acq samples for each channel in acq_chans with shape (n_total_shots, nsamples)

            - if `return_acc` is `True`:

                - tuple:
                    - dict:
                        array of acq samples for each channel in acq_chans with shape `(n_total_shots, nsamples)`
                    - dict:
                        array of acc values for each loaded channel with length `n_total_shots`

        """
        self.load_executable(raw_asm_prog)
        return self.run_circuit_acq(n_total_shots, nsamples, acq_chans, trig_delay, decimator, return_acc, from_server)

    def run_circuit(self, 
                    n_total_shots: int, 
                    reads_per_shot: int | Dict[str, int] = None, 
                    timeout_per_shot: float = 20,
                    from_server: bool = False):
        """
        Run the currently loaded program and acquire integrated IQ shots. Program is
        run `n_total_shots` times, in batches of size `shots_per_run` (i.e. `shots_per_run` runs of the program
        are executed in logic before each readback/restart cycle). The current gateware 
        is limited to ~1000 reads in its IQ buffer, which generally means 
        shots_per_run = 1000//reads_per_shot

        Parameters
        ----------
        n_total_shots: int
            number of shots to run. Program is restarted from the beginning 
            for each new shot
        reads_per_shot: int | dict
            number of values per shot per channel to read back from accbuf. If `dict`, indexed
            by str(channel_number) (same indices as `raw_asm_list`). If `int`, assumed to be 
            the same across channels. Unless multiple circuits were rastered pre-compilation or 
            there is mid-circuit measurement involved this is typically 1
        timeout_per_shot: float
            job will time out if time to take a single shot exceeds this value in seconds 
            (this likely means the job is hanging due to timing issues in the program or gateware)
        from_server: bool
            set to true if calling over RPC. If `True`, pack returned s11 arrays into
            byte objects

        Returns
        -------
        dict:
            Complex IQ shots for each accbuf in `chanlist`; each array has 
            shape `(n_total_shots, reads_per_shot)`
        """
        if isinstance(reads_per_shot, int):
            for _, chan in self.loaded_result_channels.items():
                chan.reads_per_shot = reads_per_shot
        elif isinstance(reads_per_shot, dict):
            for channame, n_reads in reads_per_shot.items():
                if channame in self.loaded_result_channels:
                    self.loaded_result_channels[channame].reads_per_shot = n_reads
        else:
            if reads_per_shot is not None:
                raise Exception(f'reads per shot: {reads_per_shot} invalid type')

        logging.getLogger(__name__).info(f'starting circuit with {n_total_shots} shots')

        if len(self.loaded_result_channels) == 0:
            shots_per_run = n_total_shots
        else:
            shots_per_run = min(self._pl_driver.get_mem_size(res_chan.mem_name)
                                // (res_chan.reads_per_shot * get_result_class(res_chan.dtype).word_size())
                                for res_chan in self.loaded_result_channels.values())

        logging.getLogger(__name__).info(f'shots_per_run: {shots_per_run}')
        n_runs = int(np.ceil(n_total_shots/shots_per_run))
        results = {ch: [] for ch in self.loaded_result_channels.keys()}

        _t_loop_start = time.time()
        _t_start_total = 0
        _t_append_total = 0
        self._pl_driver._total_poll_time = 0
        self._pl_driver._total_read_time = 0

        for i in range(n_runs):
            _t0 = time.time()
            self.start_program(shots_per_run if i < (n_runs - 1) else n_total_shots - i*shots_per_run)
            _t1 = time.time()
            cur_result = self.wait_and_readback()
            _t2 = time.time()
            for ch in self.loaded_result_channels.keys():
                results[ch].append(cur_result[ch])
            _t3 = time.time()

            _t_start_total += _t1 - _t0
            _t_append_total += _t3 - _t2

        _t_loop_end = time.time()
        _t_poll_total = self._pl_driver._total_poll_time
        _t_read_total = self._pl_driver._total_read_time
        logging.getLogger(__name__).info('done circuit')

        results = {ch: b''.join(chunks) for ch, chunks in results.items()}
        _t_join_end = time.time()

        if not from_server:
            for ch in results.keys():
                results[ch] = get_result_class(self.loaded_result_channels[ch].dtype)(results[ch], n_total_shots)
        _t_parse_end = time.time()

        print(f'[run_circuit timing] n_runs={n_runs}  shots_per_run={shots_per_run}')
        print(f'  start_program  = {_t_start_total:.3f}s  ({_t_start_total/n_runs*1000:.1f}ms/run)')
        print(f'  wait(poll)     = {_t_poll_total:.3f}s  ({_t_poll_total/n_runs*1000:.1f}ms/run)')
        print(f'  wait(readback) = {_t_read_total:.3f}s  ({_t_read_total/n_runs*1000:.1f}ms/run)')
        print(f'  append         = {_t_append_total:.3f}s')
        print(f'  loop total     = {_t_loop_end - _t_loop_start:.3f}s')
        print(f'  join           = {_t_join_end - _t_loop_end:.3f}s')
        print(f'  result parse   = {_t_parse_end - _t_join_end:.3f}s')
        print(f'  TOTAL          = {_t_parse_end - _t_loop_start:.3f}s')

        return results

    def run_circuit_acq(self,
                        n_total_shots: int = 1, 
                        nsamples: int = 8192, 
                        acq_chans: Dict[str, int] = {'0':0,'1':1}, 
                        trig_delay: float = 0, 
                        decimator: int = 0, 
                        return_acc: bool = False, 
                        from_server: bool = False):
        """
        Run the currently loaded program and acquire raw (or downconverted) adc traces.

        Parameters
        ----------
        n_total_shots: int
            number of shots to run. Program is restarted from the beginning 
            for each new shot
        nsamples: int
            number of samples to read from the acq buffer
        acq_chans: dict
            current channel mapping is:

                '0': ADC_237_2 (main readout ADC)
                '1': ADC_237_0 (other ADC connected in gateware)
                TODO: figure out DLO channels, etc and what they mean
        trig_delay: float
            time to delay acquisition, relative to circuit start.
            NOTE: this value, when converted to units of clock cycles, is a 
            16-bit value. So, it maxes out at CLK_PERIOD*(2**16) = 131.072e-6
        decimator: int
            decimation interval when sampling. e.g. 0 means full sample rate, 1
            means capture every other sample, 2 means capture every third sample, etc
        return_acc: bool
            if True, return a single acc (integrated + accumulated readout) value per shot,
            on each loaded channel. Default is False.
        from_server: bool
            set to true if calling over RPC. If True, pack returned acq arrays into
            byte objects

        Returns
        -------
        tuple | Dict
            - if return_acc is False:

                - dict:
                    array of acq samples for each channel in `acq_chans` with shape `(n_total_shots, nsamples)`

            - if return_acc is True:

                - tuple:
                    - dict:
                        array of acq samples for each channel in acq_chans with shape `(n_total_shots, nsamples)`
                    - dict:
                        array of acc values for each loaded channel with length `n_total_shots`

        """
        if nsamples > MAX_NSAMPLES:
            raise RuntimeError(f'{nsamples} exceeds max_nsamples length of {MAX_NSAMPLES}')

        if return_acc:
            acc_chans = self.loaded_result_channels
        else:
            acc_chans = {}
        acq_data, acc_data = self._pl_driver.run_prog_acq(n_total_shots, nsamples, acq_chans, acc_chans,
                                                int(trig_delay/CLK_PERIOD), decimator)

        if from_server:
            for ch in acq_data.keys():
                acq_data[ch] = acq_data[ch].tobytes()
            for ch in acc_data.keys():
                acc_data[ch] = acc_data[ch].tobytes()

        if return_acc:
            return acq_data, acc_data

        else: 
            return acq_data

ddr_batch_prepare(executable, load_freqs=True, load_envs=True)

Prepare the board for a DDR command-streaming batch (633e80df lineage).

Python keeps its side of the split: dsp resets + env/freq/register BRAM loads. Commands do NOT go through here — the BRAM command path is physically removed in this gateware; the client streams K x 256 KiB command images to batch_server (port 8081), which feeds DDR bank c1 via MM2S and issues global_start.

A hardware-sequenced batch shares the BRAM env/freq tables across its circuits, so the client verifies uniformity and passes ONE representative executable per hardware batch. When per-circuit tables differ and the user allowed reload_freq/reload_env, the client instead calls this once per circuit (software-sequenced; circuits are then NOT continuous). Board-verified prep order from ddr_verify.c: resets first, then tables.

Source code in qubic/run.py
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
def ddr_batch_prepare(self, executable, load_freqs=True, load_envs=True):
    """Prepare the board for a DDR command-streaming batch (633e80df lineage).

    Python keeps its side of the split: dsp resets + env/freq/register BRAM
    loads. Commands do NOT go through here — the BRAM command path is
    physically removed in this gateware; the client streams K x 256 KiB
    command images to batch_server (port 8081), which feeds DDR bank c1
    via MM2S and issues global_start.

    A hardware-sequenced batch shares the BRAM env/freq tables across its
    circuits, so the client verifies uniformity and passes ONE representative
    executable per hardware batch. When per-circuit tables differ and the
    user allowed reload_freq/reload_env, the client instead calls this once
    per circuit (software-sequenced; circuits are then NOT continuous).
    Board-verified prep order from ddr_verify.c: resets first, then tables.
    """
    if isinstance(executable, dict):
        executable = executable_from_dict(executable)
    self._pl_driver.set_default_regs()
    for reg in ('dspreset', 'resetacc', 'reset_bram_read'):
        self._pl_driver.write_reg(reg, 1)
        self._pl_driver.write_reg(reg, 0)
    self.load_executable(executable, load_commands=False, load_freqs=load_freqs,
                         load_envs=load_envs, zero=False)
    return True

ddr_start_circuit(raw_asm, n_total_shots, circuit_index, reload_cmd=True, reload_freq=True, reload_env=True, zero_between_reload=True)

Load and start a single circuit for DDR readout.

Called by the client when it connects directly to the DMA server, bypassing XML-RPC for bulk data transfer.

Source code in qubic/run.py
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
def ddr_start_circuit(self, raw_asm, n_total_shots, circuit_index,
                      reload_cmd=True, reload_freq=True, reload_env=True,
                      zero_between_reload=True):
    """Load and start a single circuit for DDR readout.

    Called by the client when it connects directly to the DMA server,
    bypassing XML-RPC for bulk data transfer.
    """
    t0 = time.time()
    if circuit_index == 0:
        self._pl_driver.set_default_regs()
        self.load_executable(raw_asm, True, True, True, True)
    else:
        self.load_executable(raw_asm, zero=zero_between_reload,
                             load_commands=reload_cmd, load_freqs=reload_freq,
                             load_envs=reload_env)
    t_load = time.time()
    # Clear stale SHOT_DONE / CUR_ADDR / FINAL_ADDR from previous circuits
    # or BRAM mode before starting this circuit. SHOT_DONE is clear-on-read.
    # See Desktop/ddr_producer_stale_state_bug.md.
    axil = self._pl_driver.overlay.axi_lite_ctrl_0.mmio
    _ = axil.read(0x1C)            # SHOT_DONE  (clear-on-read)
    axil.write(0x18, 0)            # FINAL_ADDR
    axil.write(0x20, 0)            # CUR_ADDR
    t_clear = time.time()
    self.start_program(n_total_shots)
    t_start = time.time()
    print(f'[DDR timing server] circuit {circuit_index}: '
          f'load={t_load-t0:.3f}s  clear={t_clear-t_load:.3f}s  '
          f'start={t_start-t_clear:.3f}s  total={t_start-t0:.3f}s')
    return True

load_and_run(executable, n_total_shots)

Load circuit described by rawasm "binary", then run for n_total_shots.

Parameters:

Name Type Description Default
executable Executable
required
n_total_shots int

number of shots to run. Program is restarted from the beginning for each new shot

required

Returns:

Name Type Description
dict

Complex IQ shots for each accbuf in chanlist; each array has shape (n_total_shots, reads_per_shot)

Source code in qubic/run.py
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
def load_and_run(self, executable: Executable, n_total_shots: int): 
    """
    Load circuit described by rawasm "binary", then run for n_total_shots. 

    Parameters
    ----------
    executable: Executable
    n_total_shots: int
        number of shots to run. Program is restarted from the beginning 
        for each new shot

    Returns
    -------
    dict:
        Complex IQ shots for each accbuf in chanlist; each array has 
        shape `(n_total_shots, reads_per_shot)`
    """
    self.load_executable(Executable)
    return self.run_circuit(n_total_shots)

load_and_run_acq(raw_asm_prog, n_total_shots=1, nsamples=8192, acq_chans={'0': 0, '1': 1}, trig_delay=0, decimator=0, return_acc=False, from_server=False)

Load the program given by raw_asm_prog and acquire raw (or downconverted) adc traces.

Parameters:

Name Type Description Default
raw_asm_prog Executable

Compiled program binary to run. See load_executable for details.

required
n_total_shots int

number of shots to run. Program is restarted from the beginning for each new shot

1
nsamples int

number of samples to read from the acq buffer

8192
acq_chans Dict[str, int]

current channel mapping is:

'0': ADC_237_2 (main readout ADC)
'1': ADC_237_0 (other ADC connected in gateware)
TODO: figure out DLO channels, etc and what they mean
{'0': 0, '1': 1}
trig_delay float

time to delay acquisition, relative to circuit start. NOTE: this value, when converted to units of clock cycles, is a 16-bit value. So, it maxes out at CLK_PERIOD(2*16) = 131.072e-6

0
decimator int

decimation interval when sampling. e.g. 0 means full sample rate, 1 means capture every other sample, 2 means capture every third sample, etc

0
return_acc bool

if True, return a single acc (integrated + accumulated readout) value per shot, on each loaded channel. Default is False.

False
from_server bool

set to true if calling over RPC. If True, pack returned acq arrays into byte objects

False

Returns:

Type Description
tuple | Dict
  • if return_acc is False:

    • dict: array of acq samples for each channel in acq_chans with shape (n_total_shots, nsamples)
  • if return_acc is True:

    • tuple:
      • dict: array of acq samples for each channel in acq_chans with shape (n_total_shots, nsamples)
      • dict: array of acc values for each loaded channel with length n_total_shots
Source code in qubic/run.py
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
def load_and_run_acq(self, 
                     raw_asm_prog: Executable, 
                     n_total_shots: int = 1, 
                     nsamples: int = 8192, 
                     acq_chans: Dict[str, int] = {'0':0,'1':1}, 
                     trig_delay: float = 0, 
                     decimator: int = 0, 
                     return_acc: bool = False, 
                     from_server: bool = False):
    """
    Load the program given by `raw_asm_prog` and acquire raw (or downconverted) adc traces.

    Parameters
    ----------
    raw_asm_prog: Executable | Dict
        Compiled program binary to run. See `load_executable` for details.
    n_total_shots: int
        number of shots to run. Program is restarted from the beginning 
        for each new shot
    nsamples: int
        number of samples to read from the acq buffer
    acq_chans: dict
        current channel mapping is:

            '0': ADC_237_2 (main readout ADC)
            '1': ADC_237_0 (other ADC connected in gateware)
            TODO: figure out DLO channels, etc and what they mean
    trig_delay: float
        time to delay acquisition, relative to circuit start.
        NOTE: this value, when converted to units of clock cycles, is a 
        16-bit value. So, it maxes out at CLK_PERIOD*(2**16) = 131.072e-6
    decimator: int
        decimation interval when sampling. e.g. 0 means full sample rate, 1
        means capture every other sample, 2 means capture every third sample, etc
    return_acc: bool
        if True, return a single acc (integrated + accumulated readout) value per shot,
        on each loaded channel. Default is False.
    from_server: bool
        set to true if calling over RPC. If True, pack returned acq arrays into
        byte objects

    Returns
    -------
    tuple | Dict
        - if `return_acc` is `False`:

            - dict:
                array of acq samples for each channel in acq_chans with shape (n_total_shots, nsamples)

        - if `return_acc` is `True`:

            - tuple:
                - dict:
                    array of acq samples for each channel in acq_chans with shape `(n_total_shots, nsamples)`
                - dict:
                    array of acc values for each loaded channel with length `n_total_shots`

    """
    self.load_executable(raw_asm_prog)
    return self.run_circuit_acq(n_total_shots, nsamples, acq_chans, trig_delay, decimator, return_acc, from_server)

load_executable(executable, load_commands=True, load_freqs=True, load_envs=True, zero=True)

Load the circuit described by executable, which is the output of the final distributed proc assembler stage. Loads command memory, env memory and freq buffer memory, according to specified input parameters. Before circuit is loaded, if zero=True, all channels are zeroed out using zero_command_buf()

Parameters:

Name Type Description Default
executable Executable | dict

Compiled program binary. Can be an Executable object or its dictionary representation (for xmlrpc support).

required
zero bool

if True, (default), zero out all cmd buffers before loading circuit

True
load_commands bool

if True, (default), load command buffers

True
load_freqs bool

if True, (default), load freq buffers

True
load_envs bool

if True, (default), load env buffers

True
Source code in qubic/run.py
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
def load_executable(self, executable: Executable | dict, load_commands: bool = True, 
                    load_freqs: bool = True, load_envs: bool = True, zero: bool = True):
    """
    Load the circuit described by `executable`, which is the output of 
    the final distributed proc assembler stage. Loads command memory, env memory
    and freq buffer memory, according to specified input parameters. Before circuit is loaded, 
    if zero=True, all channels are zeroed out using zero_command_buf()

    Parameters
    ----------
    executable: Executable | dict
        Compiled program binary. Can be an Executable object or its dictionary
        representation (for xmlrpc support).
    zero: bool
        if True, (default), zero out all cmd buffers before loading circuit
    load_commands: bool
        if True, (default), load command buffers
    load_freqs: bool
        if True, (default), load freq buffers
    load_envs: bool
        if True, (default), load env buffers
    """
    if isinstance(executable, dict):
        executable = executable_from_dict(executable)

    if zero:
        self.zero_command_buf()

    self._cur_nshots = None
    for memname, data in executable.get_binaries_fromboard().items():
        if load_commands and 'command' in memname:
            self._pl_driver.write_mem_buf(memname, data)
        elif load_envs and 'env' in memname:
            self._pl_driver.write_mem_buf(memname, data)
        elif load_freqs and 'freq' in memname:
            self._pl_driver.write_mem_buf(memname, data)

    for regname, value in executable.get_registers_fromboard().items():
        self._pl_driver.write_reg(regname, value)

    self.loaded_result_channels = executable.result_channels

run_circuit(n_total_shots, reads_per_shot=None, timeout_per_shot=20, from_server=False)

Run the currently loaded program and acquire integrated IQ shots. Program is run n_total_shots times, in batches of size shots_per_run (i.e. shots_per_run runs of the program are executed in logic before each readback/restart cycle). The current gateware is limited to ~1000 reads in its IQ buffer, which generally means shots_per_run = 1000//reads_per_shot

Parameters:

Name Type Description Default
n_total_shots int

number of shots to run. Program is restarted from the beginning for each new shot

required
reads_per_shot int | Dict[str, int]

number of values per shot per channel to read back from accbuf. If dict, indexed by str(channel_number) (same indices as raw_asm_list). If int, assumed to be the same across channels. Unless multiple circuits were rastered pre-compilation or there is mid-circuit measurement involved this is typically 1

None
timeout_per_shot float

job will time out if time to take a single shot exceeds this value in seconds (this likely means the job is hanging due to timing issues in the program or gateware)

20
from_server bool

set to true if calling over RPC. If True, pack returned s11 arrays into byte objects

False

Returns:

Name Type Description
dict

Complex IQ shots for each accbuf in chanlist; each array has shape (n_total_shots, reads_per_shot)

Source code in qubic/run.py
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
def run_circuit(self, 
                n_total_shots: int, 
                reads_per_shot: int | Dict[str, int] = None, 
                timeout_per_shot: float = 20,
                from_server: bool = False):
    """
    Run the currently loaded program and acquire integrated IQ shots. Program is
    run `n_total_shots` times, in batches of size `shots_per_run` (i.e. `shots_per_run` runs of the program
    are executed in logic before each readback/restart cycle). The current gateware 
    is limited to ~1000 reads in its IQ buffer, which generally means 
    shots_per_run = 1000//reads_per_shot

    Parameters
    ----------
    n_total_shots: int
        number of shots to run. Program is restarted from the beginning 
        for each new shot
    reads_per_shot: int | dict
        number of values per shot per channel to read back from accbuf. If `dict`, indexed
        by str(channel_number) (same indices as `raw_asm_list`). If `int`, assumed to be 
        the same across channels. Unless multiple circuits were rastered pre-compilation or 
        there is mid-circuit measurement involved this is typically 1
    timeout_per_shot: float
        job will time out if time to take a single shot exceeds this value in seconds 
        (this likely means the job is hanging due to timing issues in the program or gateware)
    from_server: bool
        set to true if calling over RPC. If `True`, pack returned s11 arrays into
        byte objects

    Returns
    -------
    dict:
        Complex IQ shots for each accbuf in `chanlist`; each array has 
        shape `(n_total_shots, reads_per_shot)`
    """
    if isinstance(reads_per_shot, int):
        for _, chan in self.loaded_result_channels.items():
            chan.reads_per_shot = reads_per_shot
    elif isinstance(reads_per_shot, dict):
        for channame, n_reads in reads_per_shot.items():
            if channame in self.loaded_result_channels:
                self.loaded_result_channels[channame].reads_per_shot = n_reads
    else:
        if reads_per_shot is not None:
            raise Exception(f'reads per shot: {reads_per_shot} invalid type')

    logging.getLogger(__name__).info(f'starting circuit with {n_total_shots} shots')

    if len(self.loaded_result_channels) == 0:
        shots_per_run = n_total_shots
    else:
        shots_per_run = min(self._pl_driver.get_mem_size(res_chan.mem_name)
                            // (res_chan.reads_per_shot * get_result_class(res_chan.dtype).word_size())
                            for res_chan in self.loaded_result_channels.values())

    logging.getLogger(__name__).info(f'shots_per_run: {shots_per_run}')
    n_runs = int(np.ceil(n_total_shots/shots_per_run))
    results = {ch: [] for ch in self.loaded_result_channels.keys()}

    _t_loop_start = time.time()
    _t_start_total = 0
    _t_append_total = 0
    self._pl_driver._total_poll_time = 0
    self._pl_driver._total_read_time = 0

    for i in range(n_runs):
        _t0 = time.time()
        self.start_program(shots_per_run if i < (n_runs - 1) else n_total_shots - i*shots_per_run)
        _t1 = time.time()
        cur_result = self.wait_and_readback()
        _t2 = time.time()
        for ch in self.loaded_result_channels.keys():
            results[ch].append(cur_result[ch])
        _t3 = time.time()

        _t_start_total += _t1 - _t0
        _t_append_total += _t3 - _t2

    _t_loop_end = time.time()
    _t_poll_total = self._pl_driver._total_poll_time
    _t_read_total = self._pl_driver._total_read_time
    logging.getLogger(__name__).info('done circuit')

    results = {ch: b''.join(chunks) for ch, chunks in results.items()}
    _t_join_end = time.time()

    if not from_server:
        for ch in results.keys():
            results[ch] = get_result_class(self.loaded_result_channels[ch].dtype)(results[ch], n_total_shots)
    _t_parse_end = time.time()

    print(f'[run_circuit timing] n_runs={n_runs}  shots_per_run={shots_per_run}')
    print(f'  start_program  = {_t_start_total:.3f}s  ({_t_start_total/n_runs*1000:.1f}ms/run)')
    print(f'  wait(poll)     = {_t_poll_total:.3f}s  ({_t_poll_total/n_runs*1000:.1f}ms/run)')
    print(f'  wait(readback) = {_t_read_total:.3f}s  ({_t_read_total/n_runs*1000:.1f}ms/run)')
    print(f'  append         = {_t_append_total:.3f}s')
    print(f'  loop total     = {_t_loop_end - _t_loop_start:.3f}s')
    print(f'  join           = {_t_join_end - _t_loop_end:.3f}s')
    print(f'  result parse   = {_t_parse_end - _t_join_end:.3f}s')
    print(f'  TOTAL          = {_t_parse_end - _t_loop_start:.3f}s')

    return results

run_circuit_acq(n_total_shots=1, nsamples=8192, acq_chans={'0': 0, '1': 1}, trig_delay=0, decimator=0, return_acc=False, from_server=False)

Run the currently loaded program and acquire raw (or downconverted) adc traces.

Parameters:

Name Type Description Default
n_total_shots int

number of shots to run. Program is restarted from the beginning for each new shot

1
nsamples int

number of samples to read from the acq buffer

8192
acq_chans Dict[str, int]

current channel mapping is:

'0': ADC_237_2 (main readout ADC)
'1': ADC_237_0 (other ADC connected in gateware)
TODO: figure out DLO channels, etc and what they mean
{'0': 0, '1': 1}
trig_delay float

time to delay acquisition, relative to circuit start. NOTE: this value, when converted to units of clock cycles, is a 16-bit value. So, it maxes out at CLK_PERIOD(2*16) = 131.072e-6

0
decimator int

decimation interval when sampling. e.g. 0 means full sample rate, 1 means capture every other sample, 2 means capture every third sample, etc

0
return_acc bool

if True, return a single acc (integrated + accumulated readout) value per shot, on each loaded channel. Default is False.

False
from_server bool

set to true if calling over RPC. If True, pack returned acq arrays into byte objects

False

Returns:

Type Description
tuple | Dict
  • if return_acc is False:

    • dict: array of acq samples for each channel in acq_chans with shape (n_total_shots, nsamples)
  • if return_acc is True:

    • tuple:
      • dict: array of acq samples for each channel in acq_chans with shape (n_total_shots, nsamples)
      • dict: array of acc values for each loaded channel with length n_total_shots
Source code in qubic/run.py
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
def run_circuit_acq(self,
                    n_total_shots: int = 1, 
                    nsamples: int = 8192, 
                    acq_chans: Dict[str, int] = {'0':0,'1':1}, 
                    trig_delay: float = 0, 
                    decimator: int = 0, 
                    return_acc: bool = False, 
                    from_server: bool = False):
    """
    Run the currently loaded program and acquire raw (or downconverted) adc traces.

    Parameters
    ----------
    n_total_shots: int
        number of shots to run. Program is restarted from the beginning 
        for each new shot
    nsamples: int
        number of samples to read from the acq buffer
    acq_chans: dict
        current channel mapping is:

            '0': ADC_237_2 (main readout ADC)
            '1': ADC_237_0 (other ADC connected in gateware)
            TODO: figure out DLO channels, etc and what they mean
    trig_delay: float
        time to delay acquisition, relative to circuit start.
        NOTE: this value, when converted to units of clock cycles, is a 
        16-bit value. So, it maxes out at CLK_PERIOD*(2**16) = 131.072e-6
    decimator: int
        decimation interval when sampling. e.g. 0 means full sample rate, 1
        means capture every other sample, 2 means capture every third sample, etc
    return_acc: bool
        if True, return a single acc (integrated + accumulated readout) value per shot,
        on each loaded channel. Default is False.
    from_server: bool
        set to true if calling over RPC. If True, pack returned acq arrays into
        byte objects

    Returns
    -------
    tuple | Dict
        - if return_acc is False:

            - dict:
                array of acq samples for each channel in `acq_chans` with shape `(n_total_shots, nsamples)`

        - if return_acc is True:

            - tuple:
                - dict:
                    array of acq samples for each channel in acq_chans with shape `(n_total_shots, nsamples)`
                - dict:
                    array of acc values for each loaded channel with length `n_total_shots`

    """
    if nsamples > MAX_NSAMPLES:
        raise RuntimeError(f'{nsamples} exceeds max_nsamples length of {MAX_NSAMPLES}')

    if return_acc:
        acc_chans = self.loaded_result_channels
    else:
        acc_chans = {}
    acq_data, acc_data = self._pl_driver.run_prog_acq(n_total_shots, nsamples, acq_chans, acc_chans,
                                            int(trig_delay/CLK_PERIOD), decimator)

    if from_server:
        for ch in acq_data.keys():
            acq_data[ch] = acq_data[ch].tobytes()
        for ch in acc_data.keys():
            acc_data[ch] = acc_data[ch].tobytes()

    if return_acc:
        return acq_data, acc_data

    else: 
        return acq_data

run_circuit_batch(executables, n_total_shots, reads_per_shot=None, timeout_per_shot=20, reload_cmd=True, reload_freq=True, reload_env=True, zero_between_reload=True, ddr=False, from_server=False)

Runs a batch of circuits given by a list of compiled executables. Each circuit is run n_total_shots times. reads_per_shot and n_total_shots are passed directly into run_circuit, and must be the same for all circuits in the batch. The parameters reload_cmd, reload_freq, reload_env, and zero_between_reload control which of these fields is rewritten circuit-to-circuit (everything is rewritten initially). Leave these all at True (default) for maximum safety, to ensure that QubiC is in a clean state before each run. Depending on the circuits, some of these can be turned off to save time.

TODO: consider throwing some version of all the args here into a BatchedCircuitRun or somesuch object

Parameters:

Name Type Description Default
executables List[Executable | Dict]

List of compiled program binaries. Each element can be an Executable object or its dictionary representation (for xmlrpc support).

required
n_total_shots int

number of shots per circuit

required
reads_per_shot int

number of values per shot per channel to read back from accbuf. If dict, indexed by channel name (e.g. Q0.rdlo). If int, assumed to be the same across channels. Unless multiple circuits were rastered pre-compilation or there is mid-circuit measurement involved this is typically 1

None
timeout_per_shot float

job will time out if time to take a single shot exceeds this value in seconds (this likely means the job is hanging due to timing issues in the program or gateware)

20
reload_cmd bool

if True, reload command buffer between circuits

True
reload_freq bool

if True, reload freq buffer between circuits

True
reload_env bool

if True, reload env buffer between circuits

True
from_server bool

set to true if calling over RPC. If True, pack returned s11 arrays into byte objects

False
ddr bool

if True, use DDR memory readout instead of BRAM. Data is fetched automatically via socket after each circuit completes. Default is False.

False

Returns:

Name Type Description
list

Complex IQ shots for each accbuf in chanlist; each array has shape (len(executables), n_total_shots, reads_per_shot)

Source code in qubic/run.py
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
def run_circuit_batch(self, 
                      executables: List[Executable | Dict], 
                      n_total_shots: int, 
                      reads_per_shot: int = None,
                      timeout_per_shot: float = 20,
                      reload_cmd: bool = True,
                      reload_freq: bool = True,
                      reload_env: bool = True,
                      zero_between_reload: bool = True,
                      ddr: bool = False,
                      from_server: bool = False):
    """
    Runs a batch of circuits given by a list of compiled executables. Each circuit is run n_total_shots
    times. `reads_per_shot` and `n_total_shots` are passed directly into `run_circuit`, and must
    be the same for all circuits in the batch. The parameters `reload_cmd`, `reload_freq`, `reload_env`, and 
    `zero_between_reload` control which of these fields is rewritten circuit-to-circuit (everything is 
    rewritten initially). Leave these all at `True` (default) for maximum safety, to ensure that QubiC 
    is in a clean state before each run. Depending on the circuits, some of these can be turned off 
    to save time.

    TODO: consider throwing some version of all the args here into a BatchedCircuitRun or somesuch
    object

    Parameters
    ----------
    executables: List[Executable, Dict]
        List of compiled program binaries. Each element can be an Executable object or its dictionary
        representation (for xmlrpc support).
    n_total_shots: int
        number of shots per circuit
    reads_per_shot: int | dict
        number of values per shot per channel to read back from accbuf. If dict, indexed
        by channel name (e.g. `Q0.rdlo`). If int, assumed to be the same across channels. 
        Unless multiple circuits were rastered pre-compilation or there is mid-circuit 
        measurement involved this is typically 1
    timeout_per_shot: float
        job will time out if time to take a single shot exceeds this value in seconds 
        (this likely means the job is hanging due to timing issues in the program or gateware)
    reload_cmd: bool
        if True, reload command buffer between circuits
    reload_freq: bool
        if True, reload freq buffer between circuits
    reload_env: bool
        if True, reload env buffer between circuits
    from_server: bool
        set to true if calling over RPC. If True, pack returned s11 arrays into
        byte objects
    ddr: bool
        if True, use DDR memory readout instead of BRAM. Data is fetched
        automatically via socket after each circuit completes. Default is False.
    Returns
    -------
    list:
        Complex IQ shots for each accbuf in chanlist; each array has
        shape `(len(executables), n_total_shots, reads_per_shot)`
    """
    results = []

    # Defer all timing prints until AFTER the loop so the print() latency
    # never pollutes the next circuit's load measurement. Also use 6
    # decimal places (microsecond precision).
    timing_log = []

    if ddr:
        # DDR mode: open one persistent RING session for all circuits.
        # The DMA server waits for SHOT_DONE after each "NEXT".
        t_batch_start = time.time()
        with _DdrRingSession(len(executables)) as ddr_session:
            for i, raw_asm in enumerate(tqdm(executables)):
                logging.getLogger(__name__).info(f'starting circuit {i}/{len(executables)-1}')

                t0 = time.time()
                if i == 0:
                    self._pl_driver.set_default_regs()
                    self.load_executable(raw_asm, True, True, True, True)
                else:
                    self.load_executable(raw_asm, zero=zero_between_reload, load_commands=reload_cmd,
                                         load_freqs=reload_freq, load_envs=reload_env)
                t_load = time.time()

                self.start_program(n_total_shots)
                t_start = time.time()

                raw_bytes = ddr_session.read_next()
                t_read = time.time()

                results.append(raw_bytes)
                timing_log.append(
                    f'[DDR timing server] circuit {i}: '
                    f'load={t_load-t0:.6f}s  start={t_start-t_load:.6f}s  '
                    f'dma_read={t_read-t_start:.6f}s  data={len(raw_bytes)} bytes')

        t_batch_end = time.time()
        timing_log.append(
            f'[DDR timing server] total: {t_batch_end-t_batch_start:.6f}s '
            f'for {len(executables)} circuits')
    else:
        for i, raw_asm in enumerate(tqdm(executables)):
            logging.getLogger(__name__).info(f'starting circuit {i}/{len(executables)-1}')
            t0 = time.time()
            if i == 0:
                self._pl_driver.set_default_regs()
                self.load_executable(raw_asm, True, True, True, True)
            else:
                self.load_executable(raw_asm, zero=zero_between_reload, load_commands=reload_cmd,
                                     load_freqs=reload_freq, load_envs=reload_env)
            t_load = time.time()

            results_i = self.run_circuit(n_total_shots, reads_per_shot, timeout_per_shot, from_server)
            t_run = time.time()
            results.append(results_i)
            timing_log.append(
                f'[BRAM timing server] circuit {i}: '
                f'load={t_load-t0:.6f}s  run={t_run-t_load:.6f}s  '
                f'total={t_run-t0:.6f}s')

    # Emit deferred timing log AFTER all measurements are complete.
    for line in timing_log:
        print(line)

    logging.getLogger(__name__).info('batch finished')
    return results

wait_and_readback(reads_per_shot=None, from_server=False)

Returns:

Type Description
Dict[str, bytes]

results for each channel

Source code in qubic/run.py
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
def wait_and_readback(self, reads_per_shot=None, from_server=False) -> Dict[str, np.ndarray]:
    """
    Returns
    -------
    Dict[str, bytes]
        results for each channel
    """
    if isinstance(reads_per_shot, int):
        for _, chan in self.loaded_result_channels.items():
            chan.reads_per_shot = reads_per_shot
    elif isinstance(reads_per_shot, dict):
        for channame, n_reads in reads_per_shot.items():
            if channame in self.loaded_result_channels:
                self.loaded_result_channels[channame].reads_per_shot = n_reads
    else:
        if reads_per_shot is not None:
            raise Exception(f'reads per shot: {reads_per_shot} invalid type')

    # s11 = {channame: np.nan*np.zeros((1, self._cur_nshots, chan.reads_per_shot), dtype=np.complex128) 
    #        for channame, chan in self.loaded_result_channels.items()}
    return self._pl_driver.wait_and_readback(self.loaded_result_channels, self._cur_nshots)

zero_command_buf(core_inds=None)

Loads command memory with dummy asm program: reset phase, output done signal, then idle. This is useful/necessary if a new program is loaded on a subset of cores such that the previous program is not completely overwritten (e.g. you are loading a program that runs only on core 2, and the previous program used cores 2 and 3).

Parameters:

Name Type Description Default
core_inds List[str | int]

list of channels (proc cores) to load. Defaults to all channels in currently loaded gateware.

None
Source code in qubic/run.py
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
def zero_command_buf(self, core_inds: List[str | int] = None):
    """
    Loads command memory with dummy asm program: reset phase, 
    output done signal, then idle. This is useful/necessary if 
    a new program is loaded on a subset of cores such that the 
    previous program is not completely overwritten (e.g. you 
    are loading a program that runs only on core 2, and the 
    previous program used cores 2 and 3).

    Parameters
    ----------
    core_inds: list
        list of channels (proc cores) to load. Defaults to
        all channels in currently loaded gateware.
    """
    zero_prog = Executable({':' + mem: bytes(16) for mem in self._pl_driver.get_program_memories(core_inds)}) 
    mems = zero_prog.get_binaries_fromboard()
    for memname, data in mems.items():
        self._pl_driver.write_mem_buf(memname, data)