• Portada
    • Recent
    • Users
    • Register
    • Login

    Hardlimit Museum

    Scheduled Pinned Locked Moved General
    92 Posts 11 Posters 34.3k Views 3 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • _Neptunno__ Offline
      _Neptunno_ MODERADOR @cobito
      last edited by

      @cobito I really love the hard work you're putting into the Museum and the new disk browser interface. As a digital preservation professional, I can only take my hat off to you.

      Let me tell you, because you'll like this: my company is dedicated precisely to digital preservation on a global scale. We work with clients ranging from the National Library of Spain (BNE) to other international ones like Harvard University, the United States Holocaust Memorial Museum, or HILA (Stanford University), among many other top-tier museums and universities, setting up systems that cost hundreds of thousands of euros. And let me tell you something: seeing what you're achieving with limited resources is incredibly impressive; the structure with MD5 and SHA-256 checksums and metadata cataloging is at a level of rigor that many institutions would envy.

      That's why I'm so excited about it. In my humble opinion (bear in mind I'm just a technician, but after so many projects, especially the BNE one where I saw how millions of pages were published thanks in part to my years of work to allow people to access them from home), your work strikes me as professional and necessary. I know how much it takes to digitize and bring visibility to these archives, and what you're doing is truly something to take your hat off for.

      It's crucial to make this content visible and accessible before it's lost forever. In fact, if you'd like, I can reach out to my company so you can negotiate a safebox deal with them, and we can take the Museum to the big leagues, haha.

      All jokes aside, you're doing a great job and the community will always be grateful. Once I get around to tinkering with the 486 or Pentium 166 I have lying around, I'll definitely be making full use of this material!

      Cheers!!

      cobitoC 1 Reply Last reply Reply Quote 5
      • cobitoC Offline
        cobito Administrador @_Neptunno_
        last edited by

        @_Neptunno_ You're going to make me blush ?

        You have no idea how much your words make me happy and motivated. On the technical side of things, I don't have that many doubts. It's the kind of thing that can be done well in many different ways, and I'm probably not doing it perfectly, but not badly either. However, when it comes to structure and organization, I rely more on intuition than on "technical" benchmarks. My references are Archive.org and WinWorldPC. And well, this being the third time I'm trying it out (let's see if it's true that the third time's the charm).

        @_Neptunno_ said in Museo Hardlimit:

        We work with clients ranging from the Biblioteca Nacional de España (BNE) to other international ones like Harvard University, the Holocaust Museum in the U.S., or HILA (Stanford University), among many other top museums and universities, building systems that cost hundreds of thousands of euros.

        Having someone with this kind of track record tell you that you're on the right track is highly motivating. Since you're likely the biggest Hardlimit expert (at least, you're the foremost expert on the subject that I know), please share any criticism or suggestions you might have, whatever they may be.

        This is something I want to do well, and since the financial cost (as opposed to time) of development is zero, and hardware resources can be scavenged here and there, it would be great if the formal presentation were handled with rigor.

        It's still very much in its early stages, both in terms of functionality and content. So as it continues to evolve, you'll likely spot things that could be improved (if you haven't already).

        Thanks!

        Toda la actualidad en la portada de Hardlimit
        Mis cacharros

        hlbm signature

        _Neptunno__ 1 Reply Last reply Reply Quote 4
        • _Neptunno__ Offline
          _Neptunno_ MODERADOR @cobito
          last edited by _Neptunno_

          @cobito you know it's always a pleasure to help you with anything, but don't get the wrong idea, I'm not at the level of a preservation engineer! ?

          I think you'd really enjoy learning from my development colleagues; to program the software that manages not just terabytes, but petabytes of information, there's a tremendous amount of hard work behind it (not to mention a thousand things I didn't know or hadn't heard of, like the "Transfer Connector": (transfer connector) a critical function that acts as a bridge for secure and structured data ingestion). On that front, I have less to say, since my role is more on the system support side, but I've also worked on many digitization projects of all scales and I can't help but view your museum through professional eyes. I find it admirable and, above all, incredibly useful for the community!!

          To give you an idea of the scale, the systems we set up are designed to preserve "digital knowledge" for the next 200 years (or so I heard in a meeting a while ago hahaha). We always say that if this had existed in the time of the Library of Alexandria, nothing would have been lost by now. The data is stored in redundant systems that constantly audit files to ensure they're healthy and that the disks maintain their integrity. If something fails, there are several extra copies in the pools ready to step in while the damaged or "heavily used" disks are replaced.

          Measures are even taken in case of catastrophes or wars. With the Ukraine conflict, for example, critical copies have already started being moved to secure locations (for example, from the UK to Ireland) to prevent information loss in case of direct confrontation.

          That said, just so you know, I'm just the "lowest-ranking employee" at my company! But these things are cool, and I'm sharing them as some industry insider scoop?

          Cheers!!

          cobitoC 1 Reply Last reply Reply Quote 5
          • cobitoC Offline
            cobito Administrador @_Neptunno_
            last edited by

            @_Neptunno_ That's incredible. A while back, I read about a new Rosetta Stone (in the sense that it was recently created) where information was stored in a spiral, reducing the character size from visible to the naked eye down to microscopic, with a goal similar to the original Rosetta Stone (having a translation table). And then, reproduce and distribute it worldwide to ensure that kind of redundancy.

            After a quick search, I found it. It's called The Rosetta Project, and what they've done is create a disk that stores 13,000 pages of information in 1,500 languages so that the text can be read with a microscope (no digital information involved).

            You probably already know about it, since, looking at the website, it's run by Stanford University, which your company has worked with. You might have even been involved in some related development!

            Anyway, my point was that this was the only example I knew of an attempt to preserve knowledge on that scale. I never would have imagined there were so many resources dedicated to preserving knowledge at the level you describe.

            Really fascinating, to be honest.

            Toda la actualidad en la portada de Hardlimit
            Mis cacharros

            hlbm signature

            _Neptunno__ 1 Reply Last reply Reply Quote 4
            • _Neptunno__ Offline
              _Neptunno_ MODERADOR @cobito
              last edited by

              @cobito The most important part of my company is preservation, which is where the most resources are invested and is the "soul" of the company. Then there's the digitization department, which was the one that started it all.

              At first, we were fully involved in mass digitization, a quite "mechanical" process (scanning material and generating metadata and bibliographic records for the images). We worked with specific scanners to obtain TIF files (300dpi, though over time they increased to 400 and 600 using very specific cameras for this) and then generated the derivatives (JPG, PDF). Depending on the project, some were straightforward while others required mass renaming and complex structuring, such as periodicals (newspapers and magazines). Just to give you an idea, we've digitized everything from the AS newspaper to early 20th-century press for the BNE, including projects with the Prado Museum or the University of Granada (some projects for the Real Hospital de Granada), among others.
              We also carried out smaller projects for town councils and urban planning files, just to give another example. But anyway, I want you to see that it's been many years of doing a bit of everything.

              For a long time, my role was precisely that image processing, although I also handled system support. In parallel, the company grew exponentially by developing digital preservation software. That software is what manages everything now. We don't just digitize documents anymore; we ensure they remain intact and accessible.

              If you visit the BNE Digital Library, you can see millions of pages to which I've contributed in a small way through my work. As a curious anecdote: I worked with collections of Civil War photos that weren't open to the public due to copyright, but which the BNE had to preserve to release them in the future. Seeing those images was quite impactful ?

              I'm sorry to disappoint you, but that Stanford Rosetta Project is something they develop on their own. That said, as a client, we work with their infrastructure, and well... you'd laugh if I told you we've had to complain about them being so stingy. They gave us virtual machines with 127GB system disks (the default size for Hyper-V), and we had to protest to get them to add "more juice," because they constantly filled up with logs ?

              Best regards!!

              1 Reply Last reply Reply Quote 5
              • cobitoC Offline
                cobito Administrador
                last edited by cobito

                I've finally managed to organize all the content in the museum. I dug out a 2TB drive from the storage room that should give me plenty of room to keep going with this. I've also managed to free up some space on the backup drive, which solves the part that bothered me most about using a drive in this way.

                For now, all the magazine discs are already available. @vreyes1981, your Micromanía discs are already up, specifically from 2002 (full year) and 2003 (up to April), which is exactly what you had uploaded for me. If you have any more material lying around, feel free to share it (just let me know so I can keep adding it).

                I've also finished defining the database structure to move on to the next step: the file explorer. The initial phase will allow browsing the directory tree of all available discs, and it should be fully operational by the next Museum update. I'll decide next time whether to dedicate my time to the test bench or to the museum.

                Toda la actualidad en la portada de Hardlimit
                Mis cacharros

                hlbm signature

                1 Reply Last reply Reply Quote 2
                • cobitoC Offline
                  cobito Administrador
                  last edited by cobito

                  This week, we released the first phase of the file browser. All files from almost all discs are now indexed and browsable. We had to leave a couple of MDF/MDS files from Micromanía aside, as there's no way to extract them, even though most MDF/MDS files can be extracted successfully.

                  Here is a disc that, in addition to the previously announced information, now includes a browsable list of files and folders.

                  In addition to the file browser, you can view detailed information for each file, including duplicates—files that may have different names, dates, and other attributes but share the same content. The information displayed is as follows:

                  • Current file name on the medium and path.
                  • Original creation date on the medium.
                  • Size in ISO/IEC 80000-13 binary format.
                  • Descriptive content type (for most files, this information is not yet available; it will be added over time).
                  • MIME type (not 100% accurate, but close).
                  • A more detailed description of the file's contents (e.g., if it's a self-extracting executable, details about the encapsulated content are provided).
                  • An MD5 hash
                  • A SHA256 hash

                  Here is an example of pkunzip.exe, which is a fairly popular file.

                  In this first phase, determining the file system's character set has been a challenge. Sometimes UTF-8 is used, other times CP850, and there are a couple of disc images that appear to have been corrupted at the source due to bugs in the creation software (which apparently wasn't uncommon in the 90s). In any case, file names are displayed correctly, complete with ñs, accents, and other special characters, regardless of the original format.

                  We currently have 570,000 files. These come from the media of the five publications we have so far, which are being used as a reference for all development before we add more.

                  We have also begun work on the second phase of the browser, which involves extracting all extractable files. For those extracted files, the process is repeated recursively. We currently support over 70 archive formats from various eras. These are identified using heuristics rather than file extensions or magic numbers, ensuring that no compatible files are missed. Once this first version of the extractor is polished, it will be deployed to production. More formats will be added over time, but for now we're sticking with these 70-80 to prioritize other aspects.

                  On another note, we've updated the v86 configuration, making the virtual machines respond much faster now (example). I hadn't noticed this while running it locally, but now that all traffic routes through a VPS, I've become aware of latency-dependent performance bottlenecks that can be improved.

                  Finally, we've fixed the algorithm behind the new page translation system, allowing pages to render much faster now (this applies to both the museum and the test bank). This change also fixed certain broken features, such as the magnifying glass in magazines and mouse pointer capture in virtual machines.

                  Toda la actualidad en la portada de Hardlimit
                  Mis cacharros

                  hlbm signature

                  1 Reply Last reply Reply Quote 2
                  • cobitoC Offline
                    cobito Administrador
                    last edited by

                    It is now possible to view the contents of extractable files. These can include any file type, and their contents may consist of other files or binary sections. While these sections often contain irrelevant binary data, they sometimes embed relevant content such as images, sounds, animations, cursors, plain text, and more (primarily in DLLs).

                    Here is an example of a .zip file that, in turn, contains another zip. The description shows the format and algorithm used for each one.

                    70% has been indexed. I assume the process will be completed over the weekend. At this point, that adds 1.3 million files/sections to what we already had.

                    With this, phase 2 is essentially complete (pending the finalization of indexing, which is an automated process), and the most tedious part of the project is wrapped up until new formats are added (a list is already prepared for the next iteration).

                    Phase 3 will likely begin soon, which involves enabling file viewing directly from the browser. This covers images, videos, audio, MIDI files, MOD files, documents, etc., etc., etc. It’s one of the coolest features of the explorer, but it required all the groundwork we’ve done so far. Thanks to this, the project can truly become a powerful digital archaeology tool.

                    Additionally, the frontend has been deployed across the entire museum (except for the hardware pages): the new version is now active in all sections, improving the desktop appearance and fixing numerous issues that were broken in the mobile version.

                    Furthermore, the videos covering each hardware and software item that I originally uploaded to Peertube have been added: example.

                    Toda la actualidad en la portada de Hardlimit
                    Mis cacharros

                    hlbm signature

                    1 Reply Last reply Reply Quote 2
                    • cobitoC Offline
                      cobito Administrador
                      last edited by cobito

                      Phase 3A is in progress and has already entered production (there's still a lot left to process, but that part is now automated). Standardized, open formats are being used for playback. The goal is to ensure that, regardless of the original format, files can be viewed in any browser. This addresses a common issue with older files, which often become unplayable due to codec, format, or algorithm incompatibilities. The selected formats are: VP9 for video, OPUS for audio, and WebP for images.

                      The following file types can be viewed directly in the browser:

                      Images, videos, and audio
                      These three types of media are being extracted using heuristics. This means a large number of images, videos, and audio files are being pulled out, even from files that aren't explicitly identifiable as such. For images, a good amount of metadata is also displayed. Over time, metadata will be added to videos and audio files, and color histograms may also be displayed for images (this data is already being captured; it just needs to be rendered).

                      Images are also processed with OCR. As long as the text is clear enough, the results are quite good. It will be used in the future for searching text within images, but for now, the OCR output is already displayed when you view an image.

                      Image Example 1
                      Image Example 2 (representative OCR example)

                      MIDIs
                      MIDIs are available in six variations (no more, no less). They have been rendered using:

                      • OPL2 (synthesizer).
                      • OPL3 (synthesizer).
                      • Gravis UltraSound (official PATs).
                      • Roland MT-32 (official ROMs).
                      • FluidR3 (modern soundfont).
                      • ToH (modern soundfont).

                      In some cases, generating MT-32, GUS, and/or ToH outputs wasn't possible. Many older MIDIs are malformed, non-compliant with standards, etc.

                      MIDI Example
                      MIDI Example 2

                      MODs
                      These are being rendered to match the Amiga's Paula chip as closely as possible, thanks to OpenMPT. Initially, multiple versions were planned, but the situation here is more challenging, and it seems all efforts are currently focused on this single implementation. File scanning here also relies on heuristics, and honestly, some really interesting finds are surfacing. For example, PSM files, which were a type of MOD used by Epic in games like their Pinball or Jazz Jackrabbit (I actually had to look these up because I had no idea they existed).

                      Jazz JackRabbit PSM Example
                      Another Silver Pinball PSM (precursor to Epic Megagames' Pinball)
                      Standard MOD
                      Another MOD

                      File Browser
                      On a separate note, when you open any folder, a selection of these files is displayed, covering the current folder and all parent directories (showing a maximum of 6 files per type). As you navigate through the directories, the displayed media becomes progressively filtered. If you click on a viewable file, all its details are displayed alongside the preview. Items are sorted by "importance," which is determined by pixel count for images and duration for everything else.

                      Processing is slow (we're only at 2%). It's still working through the earlier entries in the queue. Here's an example:
                      PCMania 21 Root Directory

                      Additionally, you can view all files of a specific type from the current directory. For example, here are all the images from PCMania 27.

                      Finally, icons are now being displayed next to files and folders to make them easier to identify. There are still many left to add, but they'll be rolled out gradually. I'm using Unicode characters for this, as I'm currently getting into the habit of normalizing formats and encodings.

                      One small thing: this is supposed to be a file search engine. I went to look up a few files to include in this thread as examples, only to realize I completely forgot to implement the search feature and hadn't noticed until now. I'll work on getting that up and running soon.

                      Phase 3B will involve applying the same process to documents: txt, rtf, wp5.1, pdfs, docs, etc., etc. But that will be pushed to a much later date.

                      Toda la actualidad en la portada de Hardlimit
                      Mis cacharros

                      hlbm signature

                      1 Reply Last reply Reply Quote 1
                      • cobitoC Offline
                        cobito Administrador
                        last edited by cobito

                        The proof of concept for the fourth (and final) phase of the file browser has been implemented. The idea here, rather than just adding pure functionality, was to demonstrate whether this was possible. And it is!

                        It consists of being able to run executables directly from the browser. The system automatically detects the executable type and launches it in MS-DOS 6.22 or Windows 95/98. Additionally, we wanted to test something else related to all this, which is currently limited to certain image and sound formats: the ability to view files in native software. For now, you can view GIFs/JPEGs in Internet Explorer 3/4, BMPs in Windows 95 Paint, and all three formats in Imaging. Also, you can listen to .wav files in the Windows Sound Recorder.

                        Some examples (to hear sound, you need to click on the emulation screen):

                        Example 0: Tomb Raider II Gold Demo
                        Example 1: Theme Hospital DOS Demo
                        Example 2: Theme Hospital Windows 95 Demo.
                        Example 3: International Rally Championship Demo (- to brake, key to the right of Ñ to accelerate, z/x to steer):
                        Example 4: Duken Nukem 3D Demo (do not select a sound card because the audio files were not included and it will crash)
                        Example 5: Epic MegaPinball Demo.

                        The Windows 95 emulations also come in two flavors: with 32-bit and 8-bit color depth, so you can experiment with image palettes and improve software compatibility. This way, you have Windows 95 at 8 and 32 bits, and Windows 98 SE at 32 bits, giving you several options in case you encounter incompatibilities.

                        Example 5: JPEG image from Internet Explorer 3/4
                        Example 6: BMP image from MS Paint
                        Example 7: Sound from Windows 95 Sound Recorder

                        In the previous version of the museum, it was possible to emulate some MS-DOS programs, but there was a very strong constraint: the program had to be contained in a .zip file, which meant preparing individual emulations, something that was extremely time-consuming. Now, thanks to HLFSv2 (Hardlimit File System!), flexibility is absolute, and each file can be handled individually, wherever it may be, at an amazing speed. And that's not counting the fact that the MS-DOS limit is broken and expanded, theoretically, to any x86 system (there's still a lot of work to be done here). With this, we have already far surpassed the functionality of the previous version, and as far as I know, we are the only ones able to run programs in this way.

                        This feature is still in its early stages and will be polished very gradually: for example, from MS-DOS and Windows 98 it is not possible to read from cylinder 1024 of the disk onwards (you will encounter read errors in very large folders: >500MB): if this happens to you with Windows 98, use 95. Additionally, the virtual drive only supports files in 8.3 format and other related issues.

                        On another note, from the file search tool, it is now possible to search for files. Searches are performed across the entire file system until a specific medium/disk is visited. From there, searches are scoped to that medium or directory recursively. The search, unlike the multimedia section, sorts by number of hits, meaning that more popular files appear first. A "media" selection filter has also been added to the search itself.

                        With this, I'm quite satisfied, and this intense development season for the museum is now closed. From now until the end of the month, changes will be consolidated and documented, and there will be no major new features (beyond minor fixes).

                        The museum, as a platform, is now fully defined.

                        Now, the big things will come from the content side, regardless of the fact that there's still plenty of room for improvement across the board.

                        PS: Media indexing is at 33%, so in another two weeks, practically everything will be ready.

                        Toda la actualidad en la portada de Hardlimit
                        Mis cacharros

                        hlbm signature

                        1 Reply Last reply Reply Quote 2

                        Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                        Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                        With your input, this post could be even better 💗

                        Register Login
                        • 1
                        • 2
                        • 3
                        • 4
                        • 5
                        • 3 / 5
                        • First post
                          Last post

                        Foreros conectados [Conectados hoy]

                        0 usuarios activos (0 miembros y 0 invitados).
                        febesin, pAtO,

                        Estadísticas de Hardlimit

                        Los hardlimitianos han creado un total de 543.8k posts en 62.9k hilos.
                        Somos un total de 37.5k miembros registrados.
                        libbyvanwagene ha sido nuestro último fichaje.
                        El récord de usuarios en linea fue de 222 y se produjo el Mon Jul 27 2026.