| ID |
Date |
Author |
Topic |
Subject |
|
3242
|
08 Jul 2026 |
Konstantin Olchanski | Suggestion | mlogger MySQL improved connection | you can also copy the reconnect code from history_schema.cxx.
if I remember right, the test case for correct reconnect is to start the mlogger,
have it write something to mysql, then restart mysqld (mlogger connection is now
broken), then have mlogger write some more to mysql. it should get an error
"connection broken", reconnect and retry the write.
(to avoid it getting stuck, there should be a limit of 10 attempts to reconnect
plus retry, in case each retry crashes mysqld or throws an unrelated error).
the code in question I think is mlogger mysql supoprt for begin and end of run,
so you start a run, restart mysqld, stop the run, start a new run and should see
a successful reconnect.
K.O. |
|
3241
|
07 Jul 2026 |
Derek Fujimoto | Info | Automated Testing | I've been working on updating MIDAS' testing suite. It is yet incomplete, but the basic structure is there. Once the code coverage reaches a more acceptable level and it is decided that the new tests are good replacements to the existing tests, it will be merged into develop.
Tests are located in $MIDASSYS/tests directory in the automated_testing branch.
ENABLE
For now changes are relegated to the automated_testing branch:
git pull
git checkout automated_testing
make cmake
STRUCTURE
Because MIDAS is implemented in three main languages (cxx, python, javascript), tests must be written in three languages. So each language has its own set of tests and frameworks in which they run. This introduces some new dependencies needed to run the tests:
| Toolchain |
New required deps |
New optional deps |
| C++ |
GoogleTest 1.15.2, CMake ≥ 3.24 |
|
| Python |
pytest, pytest-cov |
numpy, lz4 |
| JS |
Node ≥ 22, jsdom 25, @playwright/test 1, Chromium |
|
These dependencies are ONLY needed to run the test suites and do not affect the main build of MIDAS. Compiling the tests is controlled via the flag MIDAS_BUILD_TESTS, which is off by default, unless make test is called. Tests are located in the $MIDASSYS/tests directory, which should share roughly the same directory structure as the MDIAS source code.
USAGE
1. Make commands:
- "make test": run all tests
- "make coverage": run all tests and generate coverage reports. Reports are located as indicated in the below table
- "make coverage-recapture": regenerate coverage report without re-running tests.
Coverage reports locations:
| Suite |
Location |
Format |
| C++ |
cov_html/index.html |
HTML (lcov/genhtml) |
| Python |
cov_html_python/index.html |
HTML (pytest-cov), plus a term summary printed to stdout |
| JS |
tests/resources/coverage/lcov.info |
lcov info (V8 via npm run coverage) |
2. Run executables directly:
$MIDASSYS/tests/README.md has the full instructions on how to do this, as well as detailed descriptions on the tests and frameworks used. In general each cxx subdirectory has a binary executable to run the tests for that subdirectory group, whereas python and javascript have their own executables related to the frameworks running the tests. |
|
3240
|
07 Jul 2026 |
Marius Koeppel | Suggestion | Multithreaded PySequencer | Dear all,
I was wondering if one can use multiple pysequencers at the same time and if not
if this feature is planned in the future? Our experiment (mu3e) has one frontend
per detector and each sequencer would only access a subset of ODB entries.
Best,
Marius |
|
3239
|
06 Jul 2026 |
Joel Sander | Suggestion | mlogger MySQL improved connection | We recently had an instance where mlogger lost its connection to the database resulting in 317 write errors on mysql_insert and (Good) did report connection errors to the Midas log but (Bad) didn't attempt to restore the connection. With some help from Gemini, I've drafted a potential solution to reheal. Our copy uses mysql_query_debug line 1903 in our copy that inserts rows, creates tables, etc:
/* execut sql query */
status = mysql_query(db, query);
if (status)
cm_msg(MERROR, "mysql_query_debug", "SQL error: %s", mysql_error(db));
return status;
to add a reconnection attempt like the following (untested) block:
/* execute sql query */
status = mysql_query(db, query);
if (status != 0) {
unsigned int err_num = mysql_errno(db);
cm_msg(MERROR, "mysql_query_debug", "SQL query execution failed (Error %u: %s).", err_num, mysql_error(db));
cm_msg(MTRCE, "mysql_query_debug", "Network drop suspected. Backing off 5 seconds to heal handle socket...");
#ifdef OS_WIN64
Sleep(5000);
#else
sleep(5);
#endif
cm_msg(MTRCE, "mysql_query_debug", "Initiating handle re-authentication to server...");
// Setting MYSQL_OPT_PROTOCOL to MYSQL_PROTOCOL_TCP ensures that even when passing NULL for host configuration
// fallbacks, the client library strictly forces a network TCP connection instead of a local Unix socket.
unsigned int protocol = MYSQL_PROTOCOL_TCP;
mysql_options(db, MYSQL_OPT_PROTOCOL, &protocol);
if (!mysql_real_connect(db, NULL, NULL, NULL, NULL, 0, NULL, 0)) {
cm_msg(MERROR, "mysql_query_debug", "Automatic handle reconciliation failed: %s. Logging halted.", mysql_error(db));
} else {
cm_msg(MINFO, "mysql_query_debug", "Database connection handle successfully restored.");
// Retry the blocked execution string
status = mysql_query(db, query);
if (status != 0) {
cm_msg(MERROR, "mysql_query_debug", "Query retry failed after reconnection (Error %u: %s).", mysql_errno(db), mysql_error(db));
} else {
cm_msg(MINFO, "mysql_query_debug", "SQL query successfully recovered and executed on retry.");
}
}
}
return status;
|
|
3238
|
26 Jun 2026 |
Derek Fujimoto | Info | MySQL Add Column | Hello,
The UCN group at TRIUMF has used the history system extensively with a MySQL backend. They see long wait times (up to 20 mins) to add new logged variables to MIDAS equipment. I did some testing and it seems to be a MySQL issue that is fixed in recent versions (UCN uses an older version of MySQL).
Likely no update is needed to MIDAS source.
Derek |
| Attachment 1: mysql_260622.pdf
|
|
|
3237
|
25 Jun 2026 |
Yiwen Yang | Suggestion | Multithreaded deferred transitions | > I recommend against using them. Maybe you can explain what you do and I can suggest a way
> to avoid using the
deferred transition.
I've only recently picked up the code and I'm not sure how it was envisioned when
initially designed, so here's my guess at how it's being used.
Deferred transition is registered by the FPN00
(main clock) process to make sure the logger has finished logging all events before continuing to stop readout
of other frontends. There is also a timeout, in case sometimes an event goes missing, to proceed with the run
stop without getting stuck waiting for the logger.
If there is a better way to implement this then I'm happy to
give it a go.
> P.S. I do not remember any use of deferred transition in the T2K/ND280 FGD, TPC and GSC
frontends,
> maybe it is in some other subsystem or was introduced after my time.
This is all from the global
DAQ code, so indeed none of the other subsystems' frontends use the feature directly. But if the clock module
frontend is started then run stops should be deferred.
Interestingly, I just checked and this feature is
disabled in the FGD ODB, but it is enabled in the TPC ODB.
Regards,
Yiwen. |
|
3236
|
25 Jun 2026 |
Konstantin Olchanski | Suggestion | Multithreaded deferred transitions | > Multithreaded transitions were introduced by KO in 2019. Please ask him to make deferred transitions work
> again or simply use non-multithreaded transitions.
Deferred transition is the bane of MIDAS, I personally can never understand how they work
and what they do, and I studied them and understood them many times now.
I recommend against using them. Maybe you can explain what you do and I can suggest a way
to avoid using the deferred transition.
Some other people use deferred transitions, and it works for them,
in conjunction with normal transitions, which have been multithreaded
for years.
So unlikely this is a new bug.
P.S. I do not remember any use of deferred transition in the T2K/ND280 FGD, TPC and GSC frontends,
maybe it is in some other subsystem or was introduced after my time.
K.O. |
|
3235
|
25 Jun 2026 |
Konstantin Olchanski | Bug Report | incompatible ODB XML dumps | > I fixed that by not requiring the handle explicitly ...
>
> [it was...] an uncaught exception. If you do not want the abort, catch the exception.
Hi, Stefan! Thank you for fixing this!
The issue was not the exception, but the failure to load the XML ODB dump from an (immutable) data file.
This should now be fixed (TBC), so all good now.
K.O. |
|
3234
|
25 Jun 2026 |
Konstantin Olchanski | Forum | midas forum elog updated | I updated the midas forum elog to the latest version from git: 083448f7
Also investigated elogd failure to start on reboot,
it turned out to be a crasher bug, see
https://elog.psi.ch/elogs/Forum/69919
K.O. |
|
3233
|
11 Jun 2026 |
Stefan Ritt | Suggestion | Multithreaded deferred transitions | Multithreaded transitions were introduced by KO in 2019. Please ask him to make deferred transitions work
again or simply use non-multithreaded transitions.
Stefan |
|
3232
|
09 Jun 2026 |
Stefan Ritt | Bug Report | incompatible ODB XML dumps | I fixed that by not requiring the handle explicitly:
if (mxml_get_attribute(node, "handle") != nullptr)
o->set_hkey(std::stoi(std::string(mxml_get_attribute(node, "handle"))));
The reason for the handle is the following: If you attach an midas::odb object to a very large subtree of the ODB, this can take very long since each
ODB element requires a few RPC roundtrips. In Mu3e, this took up to one minute.
To overcome the problem, we can initialize an midas::odb object remotely via an XML tree. The server creates a huge XML object, sends it over
a single RPC command, and the client re-creates the midas::odb tree from the XML object. Since each online midas::odb need the ODB handle for
watch functions etc. the handle got added to the XML file.
For a XML file, this makes no sense of course, so now it's optional with the code change above.
P.S.: The odbxx code does not core dump, it just produces an error which looks like core dump. This is an uncaught exception. Normal exceptions
just abort the program without much information. The odbxx exceptions add a stack dump if available on that OS. That makes it easier to debug.
If you do not want the abort, catch the exception.
Stefan |
|
3231
|
01 Jun 2026 |
Yiwen Yang | Suggestion | Multithreaded deferred transitions | Hi,
On the DAQ system for T2K's ND280 near detector, we use deferred
transitions to make sure all triggered events were logged before issuing run
stops to frontends.
I've recently managed to update the frontends to use a
relatively modern version of MIDAS. I then noticed that run transitions are now
by default multithreaded, when issued from e.g. mhttpd, but deferred transitions
called by cm_check_deferred_transition are still performed synchronously.
It
would be nice to make run stops use multithreaded transitions as well. A naive
patch of adding the TR_MTHREAD flag does not work, since the client handling the
deferred transition attempts to communicate with itself instead of calling
cm_transition_call_direct.
After looking into the code a bit further, I noticed
that there is an intentional check against multithreaded transitions in the
logic for determining whether the client is the one calling the transition:
https://bitbucket.org/tmidas/midas/src/fd71f63c023b7e2d4a5c91e3121651b14bd9d27b/
src/midas.cxx#lines-5009
Was there a particular concern that lead to this
particular check?
Regards,
Yiwen. |
|
3230
|
29 May 2026 |
Stefan Ritt | Info | ODBvalue timeout | > > > >
> > > > How can the MSL code figure out if the wait succeeded or timed out?
> > > >
> > > > Stefan
> > >
> > > You get a message, something like:
> > > 17:52:12.293 2026/05/29 [Sequencer,INFO] WAIT ODBValue timeout after 10.0 seconds: /Equipment/Test/Variables/V < 1 not satisfied
> > >
> > > Do we need something else?
> > >
> > > Zaher
> >
> > I mean how can the following code determine the timeout?
>
> My intention with this was dealing with something like setting a cryostat temperature or any non-critical parameter. If it is not reached within a given timeout we give up and move on with the plan rather than sitting and wasting a whole night of beam. If your ODBvalue is "mission critical" then the wait command should not be used with a timeout. If you do use the timeout option then you will have to check in the following lines what is the state of your ODBvalue (very easy). To me this is the simplest and most useful way for our use case.
I was more thinking like a return value 0/1 if the wait function. If you change the condition, you only have to change it in one location. More like normal C functions work.
Stefan |
|
3229
|
29 May 2026 |
Zaher Salman | Info | ODBvalue timeout | > > >
> > > How can the MSL code figure out if the wait succeeded or timed out?
> > >
> > > Stefan
> >
> > You get a message, something like:
> > 17:52:12.293 2026/05/29 [Sequencer,INFO] WAIT ODBValue timeout after 10.0 seconds: /Equipment/Test/Variables/V < 1 not satisfied
> >
> > Do we need something else?
> >
> > Zaher
>
> I mean how can the following code determine the timeout?
My intention with this was dealing with something like setting a cryostat temperature or any non-critical parameter. If it is not reached within a given timeout we give up and move on with the plan rather than sitting and wasting a whole night of beam. If your ODBvalue is "mission critical" then the wait command should not be used with a timeout. If you do use the timeout option then you will have to check in the following lines what is the state of your ODBvalue (very easy). To me this is the simplest and most useful way for our use case. |
|
3228
|
29 May 2026 |
Stefan Ritt | Info | ODBvalue timeout | > >
> > How can the MSL code figure out if the wait succeeded or timed out?
> >
> > Stefan
>
> You get a message, something like:
> 17:52:12.293 2026/05/29 [Sequencer,INFO] WAIT ODBValue timeout after 10.0 seconds: /Equipment/Test/Variables/V < 1 not satisfied
>
> Do we need something else?
>
> Zaher
I mean how can the following code determine the timeout? |
|
3227
|
29 May 2026 |
Zaher Salman | Info | ODBvalue timeout | >
> How can the MSL code figure out if the wait succeeded or timed out?
>
> Stefan
You get a message, something like:
17:52:12.293 2026/05/29 [Sequencer,INFO] WAIT ODBValue timeout after 10.0 seconds: /Equipment/Test/Variables/V < 1 not satisfied
Do we need something else?
Zaher |
|
3226
|
29 May 2026 |
Stefan Ritt | Info | ODBvalue timeout | > Dear all, I implemented an optional timeout for the wait ODBvalue command. The way it works is similar to the standard wait command:
>
> WAIT ODBvalue, /Equipment/HV/Variables/Measured[3], <, 100, timeout, 60
>
> where the "timeout" keyword start a countdown in seconds. If the ODB condition is not met after 60 seconds the sequencer moves on to the next line.
>
> To use this feature you must recompile the msequencer, delete /Sequencer/State and start the freshly compiled msequencer. This will add two ODBs to the /Sequencer/State: "Timeout value" (the countdown) and "Timeout limit" (the limit given in the wait command).
>
> I suggest that we add something similar to the pysequencer using the same ODBs.
How can the MSL code figure out if the wait succeeded or timed out?
Stefan |
|
3225
|
29 May 2026 |
Zaher Salman | Info | ODBvalue timeout | Dear all, I implemented an optional timeout for the wait ODBvalue command. The way it works is similar to the standard wait command:
WAIT ODBvalue, /Equipment/HV/Variables/Measured[3], <, 100, timeout, 60
where the "timeout" keyword start a countdown in seconds. If the ODB condition is not met after 60 seconds the sequencer moves on to the next line.
To use this feature you must recompile the msequencer, delete /Sequencer/State and start the freshly compiled msequencer. This will add two ODBs to the /Sequencer/State: "Timeout value" (the countdown) and "Timeout limit" (the limit given in the wait command).
I suggest that we add something similar to the pysequencer using the same ODBs. |
|
3224
|
21 May 2026 |
Konstantin Olchanski | Bug Report | incompatible ODB XML dumps | While testing manalyzer, I found that it dies from an exception on odbxx, error message is "/home/olchansk/packages/midas/include/odbxx.h:1231: No "handle"
attribute found in XML data".
Indeed, my data file is very old and it's XML ODB dump does not have the "handle" attribute:
daq00:midas$ more ~/git/midas/manalyzer/run9402bor.xml
<?xml version="1.0" encoding="ISO-8859-1"?>
<!-- created by MXML on Tue Aug 11 14:47:16 2020 -->
<odb root="/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="http://midas.psi.ch/odb.xsd">
<dir name="Experiment">
<key name="ODB timeout" type="INT32">10000</key>
While current MIDAS XML ODB dumps have it:
<?xml version="1.0" encoding="ISO-8859-1"?>
<!-- created by MXML on Thu May 21 20:37:06 2026 -->
<odb root="/" filename="odb.xml" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:noNamespaceSchemaLocation="/home/olchansk/packages/midas/odb.xsd">
<dir name="System" handle="135320">
<dir name="Flush" handle="135408">
<key name="Flush period" type="UINT32" handle="135496">60</key>
And odbxx requires this attribute unconditionally:
if (mxml_get_attribute(node, "handle") == nullptr)
mthrow("No \"handle\" attribute found in XML data");
o->set_hkey(std::stoi(std::string(mxml_get_attribute(node, "handle"))));
The "handle" attribute was added to XML ODB dumps in September 2024 (not sure to what purpose, JSON ODB dumps do not have a "handle" attribute):
git blame src/odb.cxx
...
dd23558fbd src/odb.cxx (Stefan Ritt 2024-09-20 15:30:00 +0200 9387) mxml_write_attribute(writer, "handle", std::to_string(hKey).c_str());
This change makes MIDAS data files written before this date un-analyzable (unless odbxx is turned off).
I can prevent manalyzer from crashing by catching the exception, but I think it is better if odbxx code is updated to accept the pre-Sep-2024 ODB XML data
format (which were valid XML ODB dumps when they were made and users are stuck with them inside compress binary MIDAS data files).
P.S. also please check the odbxx code for other crashes on malformed XML ODB dumps, it should complain, fail to load the dump, but not core dump or abort.
Malformed ODB dumps is not a theoretical situation, I am currently looking at MIDAS data files (mid.lz4) that have invalid JSON ODB dumps created from
corrupted ODB. Luckily the JSON parser handles this gracefully, does not crash manalyzer and I can look at the data. I did have to go 10 runs into the past
to find an uncorrupted ODB dump to reload a good ODB. Fixes to the JSON encoder and fixes for corrupt ODB are in progress.
K.O. |
|
3223
|
21 May 2026 |
Konstantin Olchanski | Info | manalyzer --save-odb | Due my oversight, the code for extracting ODB dumps from MIDAS data files from rootana/old_analyzer/event_dump.cxx was missing in
manalyzer.cxx.
This is now corrected, the new manalyzer command line flag is "--save-odb", to use it:
daq00:manalyzer$ ./manalyzer_test.exe --save-odb ~/git/midas/testexpt/run00002.mid.lz4
...
Saving begin of run ODB dump for run 2 from "/home/olchansk/git/midas/testexpt/run00002.mid.lz4" to "run2bor.json"
...
Saving end of run ODB dump for run 2 from "/home/olchansk/git/midas/testexpt/run00002.mid.lz4" to "run2eor.json"
...
manalyzer commit f4cbcb7426083edc9f74298965c90a3a91f461ab
K.O. |
|